{
 "S1sh1::signage": {
  "fp": "2c20399da794ddb6",
  "inscriptions": [
   {
    "text_native": "LEGACY AUTHORITY DETECTED",
    "source": "scene_text_quoted",
    "reason_ko": "통신판 스크린 중앙에 떠오른 감지 경고 텍스트입니다.",
    "source_quote": "LEGACY AUTHORITY DETECTED"
   }
  ],
  "cues": [],
  "dropped": []
 },
 "S1sh1": {
  "input_fingerprint": "205c51cda76d8d09",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — Wyoming, United States, 2419; all people are English-speaking Americans unless stated. TIME OF DAY (lock): twilight.\n\nSHOT TEXT (authoritative, Korean): 붉은빛이 맥박치듯 가장 밝아진 순간의 통신판 스크린 중앙에 뜬 LEGACY AUTHORITY DETECTED 텍스트 클로즈업\n\nLOCATION (lock): Inside the mountain contact post's underground control room, directly at the communications console. The console is washed in pulsing red alert light. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: communication panel display (showing “LEGACY AUTHORITY DETECTED” at the brightest point of its red pulse) — The active front face is square to the camera, with the warning text centered and legible; used as Primary focal surface aligned directly with the lens axis.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Cool, muted low-key rendering is punctuated by the warning text at the peak of its red pulse.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The communications panel pulses at peak red brightness with “LEGACY AUTHORITY DETECTED” centered on its screen.\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 토니(앤서니 로저스) (미국인 남성, 30대 중반, 자연스러운 성인 남성 얼굴, 구체적으로 명시되지 않은 자연색 머리카락). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nWORDS TO RENDER (authoritative — the scene itself calls for these; render each as period-real physical lettering in the native script, exactly as written; add no other readable text anywhere):\n- \"LEGACY AUTHORITY DETECTED\"\n\nThe WORDS TO RENDER above are the only readable writing in this image: render those words exactly as given, in the place and era's own language and script, and nothing else legible. Invent no other wording a viewer could read. No caption, subtitle, watermark, logo or overlay. Surfaces that would carry writing may still be present — stage any wording they would carry out of legibility: a hand across, an oblique angle, shallow focus.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — Wyoming, United States, 2419; all people are English-speaking Americans unless stated. TIME OF DAY (lock): twilight.\n\nSHOT TEXT (authoritative, Korean): 붉은빛이 맥박치듯 가장 밝아진 순간의 통신판 스크린 중앙에 뜬 LEGACY AUTHORITY DETECTED 텍스트 클로즈업\n\nLOCATION (lock): Inside the mountain contact post's underground control room, directly at the communications console. The console is washed in pulsing red alert light. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: communication panel display (showing “LEGACY AUTHORITY DETECTED” at the brightest point of its red pulse) — The active front face is square to the camera, with the warning text centered and legible; used as Primary focal surface aligned directly with the lens axis.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Cool, muted low-key rendering is punctuated by the warning text at the peak of its red pulse.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The communications panel pulses at peak red brightness with “LEGACY AUTHORITY DETECTED” centered on its screen.\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 토니(앤서니 로저스) (미국인 남성, 30대 중반, 자연스러운 성인 남성 얼굴, 구체적으로 명시되지 않은 자연색 머리카락). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nWORDS TO RENDER (authoritative — the scene itself calls for these; render each as period-real physical lettering in the native script, exactly as written; add no other readable text anywhere):\n- \"LEGACY AUTHORITY DETECTED\"\n\nThe WORDS TO RENDER above are the only readable writing in this image: render those words exactly as given, in the place and era's own language and script, and nothing else legible. Invent no other wording a viewer could read. No caption, subtitle, watermark, logo or overlay. Surfaces that would carry writing may still be present — stage any wording they would carry out of legibility: a hand across, an oblique angle, shallow focus.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — Wyoming, United States, 2419; all people are English-speaking Americans unless stated. TIME OF DAY (lock): twilight.\n\nSHOT TEXT (authoritative, Korean): 붉은빛이 맥박치듯 가장 밝아진 순간의 통신판 스크린 중앙에 뜬 LEGACY AUTHORITY DETECTED 텍스트 클로즈업\n\nLOCATION (lock): Inside the mountain contact post's underground control room, directly at the communications console. The console is washed in pulsing red alert light. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: communication panel display (showing “LEGACY AUTHORITY DETECTED” at the brightest point of its red pulse) — The active front face is square to the camera, with the warning text centered and legible; used as Primary focal surface aligned directly with the lens axis.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Cool, muted low-key rendering is punctuated by the warning text at the peak of its red pulse.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The communications panel pulses at peak red brightness with “LEGACY AUTHORITY DETECTED” centered on its screen.\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 토니(앤서니 로저스) (미국인 남성, 30대 중반, 자연스러운 성인 남성 얼굴, 구체적으로 명시되지 않은 자연색 머리카락). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nWORDS TO RENDER (authoritative — the scene itself calls for these; render each as period-real physical lettering in the native script, exactly as written; add no other readable text anywhere):\n- \"LEGACY AUTHORITY DETECTED\"\n\nThe WORDS TO RENDER above are the only readable writing in this image: render those words exactly as given, in the place and era's own language and script, and nothing else legible. Invent no other wording a viewer could read. No caption, subtitle, watermark, logo or overlay. Surfaces that would carry writing may still be present — stage any wording they would carry out of legibility: a hand across, an oblique angle, shallow focus.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "카메라는 통신 패널의 스크린을 비스듬한 각도로 내려다보고 있다.",
    "built_space": "금속 재질의 제어판 표면에 사각형 스크린과 여러 개의 토글 스위치, 버튼들이 배치되어 있다.",
    "entities": "붉은색으로 빛나는 스크린 중앙에 'LEGACY AUTHORITY DETECTED'라는 텍스트가 표시되어 있으며, 주변 버튼에 'ON', 'OFF' 등의 추가 텍스트가 식별된다.",
    "hard_violations": [
     "[gemini-pro] 지정된 카메라 축(스크린 전면이 카메라와 직각을 이루어야 함)을 위반하고 비스듬한 시점으로 렌더링됨.",
     "[gpt] 콘솔에 “ON”과 “OFF”라는 허용되지 않은 추가 가독성 문자가 노출되어 있다."
    ],
    "physics": "물리적으로 안정적인 구조의 제어판 및 표면에 부착된 버튼들."
   },
   {
    "label": "B",
    "direction": "카메라는 통신 패널의 스크린을 정확히 정면으로 바라보고 있다.",
    "built_space": "금속 재질의 수직 제어판 중앙에 브라운관 형태의 사각형 스크린이 있고, 주변에 다이얼과 스위치들이 배치되어 있다.",
    "entities": "붉은 조명이 감도는 스크린 중앙에 'LEGACY AUTHORITY DETECTED' 텍스트가 뚜렷하게 표시되어 있으며, 패널 좌측에 'COMMUNICATIONS' 등의 금지된 추가 텍스트가 식별된다.",
    "hard_violations": [
     "[gpt] 주변 패널에 “COMMUNICATIONS” 등 프롬프트가 허용하지 않은 추가 가독성 문자가 노출되어 있다."
    ],
    "physics": "물리적으로 안정적인 구조의 제어판 및 견고하게 결합된 부품들."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "스크린이 카메라 렌즈 축과 완벽하게 정면으로 직각을 이루도록 구도를 맞추었으며 붉은 조명 묘사도 우수하지만, 지정되지 않은 추가 텍스트(COMMUNICATIONS 등)가 패널에 다소 포함된 점이 아쉽습니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "요구된 텍스트는 정확히 렌더링되었으나, 가장 중요한 프레이밍 지시사항인 '카메라와 직각을 이루는 정면 구도(square to the camera)'를 무시하고 비스듬한 각도로 렌더링하여 큰 감점을 받았습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "카메라는 통신 패널의 스크린을 비스듬한 각도로 내려다보고 있다.",
        "built_space": "금속 재질의 제어판 표면에 사각형 스크린과 여러 개의 토글 스위치, 버튼들이 배치되어 있다.",
        "entities": "붉은색으로 빛나는 스크린 중앙에 'LEGACY AUTHORITY DETECTED'라는 텍스트가 표시되어 있으며, 주변 버튼에 'ON', 'OFF' 등의 추가 텍스트가 식별된다.",
        "hard_violations": [
         "지정된 카메라 축(스크린 전면이 카메라와 직각을 이루어야 함)을 위반하고 비스듬한 시점으로 렌더링됨."
        ],
        "physics": "물리적으로 안정적인 구조의 제어판 및 표면에 부착된 버튼들."
       },
       {
        "label": "B",
        "direction": "카메라는 통신 패널의 스크린을 정확히 정면으로 바라보고 있다.",
        "built_space": "금속 재질의 수직 제어판 중앙에 브라운관 형태의 사각형 스크린이 있고, 주변에 다이얼과 스위치들이 배치되어 있다.",
        "entities": "붉은 조명이 감도는 스크린 중앙에 'LEGACY AUTHORITY DETECTED' 텍스트가 뚜렷하게 표시되어 있으며, 패널 좌측에 'COMMUNICATIONS' 등의 금지된 추가 텍스트가 식별된다.",
        "hard_violations": [],
        "physics": "물리적으로 안정적인 구조의 제어판 및 견고하게 결합된 부품들."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "스크린이 카메라 렌즈 축과 완벽하게 정면으로 직각을 이루도록 구도를 맞추었으며 붉은 조명 묘사도 우수하지만, 지정되지 않은 추가 텍스트(COMMUNICATIONS 등)가 패널에 다소 포함된 점이 아쉽습니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "요구된 텍스트는 정확히 렌더링되었으나, 가장 중요한 프레이밍 지시사항인 '카메라와 직각을 이루는 정면 구도(square to the camera)'를 무시하고 비스듬한 각도로 렌더링하여 큰 감점을 받았습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "카메라는 통신 패널의 스크린을 비스듬한 각도로 내려다보고 있다.",
        "built_space": "금속 재질의 제어판 표면에 사각형 스크린과 여러 개의 토글 스위치, 버튼들이 배치되어 있다.",
        "entities": "붉은색으로 빛나는 스크린 중앙에 'LEGACY AUTHORITY DETECTED'라는 텍스트가 표시되어 있으며, 주변 버튼에 'ON', 'OFF' 등의 추가 텍스트가 식별된다.",
        "hard_violations": [
         "지정된 카메라 축(스크린 전면이 카메라와 직각을 이루어야 함)을 위반하고 비스듬한 시점으로 렌더링됨."
        ],
        "physics": "물리적으로 안정적인 구조의 제어판 및 표면에 부착된 버튼들."
       },
       {
        "label": "B",
        "direction": "카메라는 통신 패널의 스크린을 정확히 정면으로 바라보고 있다.",
        "built_space": "금속 재질의 수직 제어판 중앙에 브라운관 형태의 사각형 스크린이 있고, 주변에 다이얼과 스위치들이 배치되어 있다.",
        "entities": "붉은 조명이 감도는 스크린 중앙에 'LEGACY AUTHORITY DETECTED' 텍스트가 뚜렷하게 표시되어 있으며, 패널 좌측에 'COMMUNICATIONS' 등의 금지된 추가 텍스트가 식별된다.",
        "hard_violations": [],
        "physics": "물리적으로 안정적인 구조의 제어판 및 견고하게 결합된 부품들."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "스크린 정면이 렌즈 축과 거의 일치하고 경고 문구도 중앙에 정확히 표시되어 핵심 클로즈업은 충실하지만, 주변 패널에 허용되지 않은 추가 영문 표기가 읽힌다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "정확한 경고 문구와 사실적인 콘솔 재질은 갖췄으나 화면을 비스듬히 내려다봐 ‘활성 전면이 카메라와 정면’이라는 핵심 구도를 어기며 ON/OFF 추가 글자도 보인다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "사람, 시선, 무기 또는 이동체는 없다. 통신판의 활성 화면은 거의 정면으로 카메라를 향하며 렌즈 축이 화면 중앙의 경고 문구에 도달한다.",
        "built_space": "지하 통제실의 낡은 금속 벽면 통신 설비가 보이고, 중앙에 고정된 적색 CRT형 화면 1개와 그 주변의 조절기·계기 패널들이 있다. 화면은 카메라에 거의 평행하고 중앙을 크게 차지해 요구된 정면 클로즈업에 부합한다. 반사광은 화면 유리 표면에서 가능한 형태다.",
        "entities": "사람은 없으며 샷 텍스트와 일치한다. 통신판 화면 1개가 실제 금속 하우징에 장착되어 있고, “LEGACY AUTHORITY DETECTED”가 정확한 철자와 순서로 화면 중앙에 두 줄로 표시된다. 다만 주변에 “COMMUNICATIONS” 등 금지된 추가 영문이 읽힌다.",
        "hard_violations": [
         "주변 패널에 “COMMUNICATIONS” 등 프롬프트가 허용하지 않은 추가 가독성 문자가 노출되어 있다."
        ],
        "physics": "화면과 조절 패널은 금속 벽체 및 하우징에 볼트로 고정되어 있으며 떠 있는 물체가 없다. 화면의 광택과 적색 반사는 유리와 주변 금속 표면에서 물리적으로 성립한다."
       },
       {
        "label": "B",
        "direction": "사람, 시선, 무기 또는 이동체는 없다. 화면의 기능면은 카메라 쪽을 향하지만 카메라가 좌측 위에서 비스듬히 내려다보므로 화면 정면과 렌즈 축이 일치하지 않는다. 경고 문구 자체는 화면 중앙에 있다.",
        "built_space": "수평에 가까운 금속 콘솔 상판에 적색 화면 1개가 매립되어 있고, 주변에 여러 버튼과 토글 스위치가 배치되어 있다. 실제 콘솔 구조와 재질은 그럴듯하지만 화면이 강한 사선 원근으로 보이며, 요구된 정면 정렬을 충족하지 못한다. 화면 유리의 반사는 이 카메라 각도에서 가능하다.",
        "entities": "사람은 없다. 통신판 화면과 “LEGACY AUTHORITY DETECTED” 문구는 정확히 식별되며 중앙에 두 줄로 표시된다. 그러나 좌측 스위치 주변의 “ON”, “OFF”가 추가로 읽혀 유일 허용 문구 조건을 위반한다.",
        "hard_violations": [
         "콘솔에 “ON”과 “OFF”라는 허용되지 않은 추가 가독성 문자가 노출되어 있다."
        ],
        "physics": "화면은 나사로 체결된 금속 프레임 안에 지지되고 버튼과 스위치도 콘솔에 장착되어 있다. 떠 있거나 지지되지 않은 물체는 없으며 재질의 마모와 반사도 물리적으로 가능하다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "스크린 정면이 렌즈 축과 거의 일치하고 경고 문구도 중앙에 정확히 표시되어 핵심 클로즈업은 충실하지만, 주변 패널에 허용되지 않은 추가 영문 표기가 읽힌다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "정확한 경고 문구와 사실적인 콘솔 재질은 갖췄으나 화면을 비스듬히 내려다봐 ‘활성 전면이 카메라와 정면’이라는 핵심 구도를 어기며 ON/OFF 추가 글자도 보인다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "사람, 시선, 무기 또는 이동체는 없다. 통신판의 활성 화면은 거의 정면으로 카메라를 향하며 렌즈 축이 화면 중앙의 경고 문구에 도달한다.",
        "built_space": "지하 통제실의 낡은 금속 벽면 통신 설비가 보이고, 중앙에 고정된 적색 CRT형 화면 1개와 그 주변의 조절기·계기 패널들이 있다. 화면은 카메라에 거의 평행하고 중앙을 크게 차지해 요구된 정면 클로즈업에 부합한다. 반사광은 화면 유리 표면에서 가능한 형태다.",
        "entities": "사람은 없으며 샷 텍스트와 일치한다. 통신판 화면 1개가 실제 금속 하우징에 장착되어 있고, “LEGACY AUTHORITY DETECTED”가 정확한 철자와 순서로 화면 중앙에 두 줄로 표시된다. 다만 주변에 “COMMUNICATIONS” 등 금지된 추가 영문이 읽힌다.",
        "hard_violations": [
         "주변 패널에 “COMMUNICATIONS” 등 프롬프트가 허용하지 않은 추가 가독성 문자가 노출되어 있다."
        ],
        "physics": "화면과 조절 패널은 금속 벽체 및 하우징에 볼트로 고정되어 있으며 떠 있는 물체가 없다. 화면의 광택과 적색 반사는 유리와 주변 금속 표면에서 물리적으로 성립한다."
       },
       {
        "label": "A",
        "direction": "사람, 시선, 무기 또는 이동체는 없다. 화면의 기능면은 카메라 쪽을 향하지만 카메라가 좌측 위에서 비스듬히 내려다보므로 화면 정면과 렌즈 축이 일치하지 않는다. 경고 문구 자체는 화면 중앙에 있다.",
        "built_space": "수평에 가까운 금속 콘솔 상판에 적색 화면 1개가 매립되어 있고, 주변에 여러 버튼과 토글 스위치가 배치되어 있다. 실제 콘솔 구조와 재질은 그럴듯하지만 화면이 강한 사선 원근으로 보이며, 요구된 정면 정렬을 충족하지 못한다. 화면 유리의 반사는 이 카메라 각도에서 가능하다.",
        "entities": "사람은 없다. 통신판 화면과 “LEGACY AUTHORITY DETECTED” 문구는 정확히 식별되며 중앙에 두 줄로 표시된다. 그러나 좌측 스위치 주변의 “ON”, “OFF”가 추가로 읽혀 유일 허용 문구 조건을 위반한다.",
        "hard_violations": [
         "콘솔에 “ON”과 “OFF”라는 허용되지 않은 추가 가독성 문자가 노출되어 있다."
        ],
        "physics": "화면은 나사로 체결된 금속 프레임 안에 지지되고 버튼과 스위치도 콘솔에 장착되어 있다. 떠 있거나 지지되지 않은 물체는 없으며 재질의 마모와 반사도 물리적으로 가능하다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt"
   ],
   "normalized": {
    "A": 1.111,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.861,
    "B": 1.75
   },
   "violations": {
    "A": [
     "[gemini-pro] 지정된 카메라 축(스크린 전면이 카메라와 직각을 이루어야 함)을 위반하고 비스듬한 시점으로 렌더링됨.",
     "[gpt] 콘솔에 “ON”과 “OFF”라는 허용되지 않은 추가 가독성 문자가 노출되어 있다."
    ],
    "B": [
     "[gpt] 주변 패널에 “COMMUNICATIONS” 등 프롬프트가 허용하지 않은 추가 가독성 문자가 노출되어 있다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 1750,
   "A": 861
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 1750,
    "verdict_ko": "스크린이 카메라 렌즈 축과 완벽하게 정면으로 직각을 이루도록 구도를 맞추었으며 붉은 조명 묘사도 우수하지만, 지정되지 않은 추가 텍스트(COMMUNICATIONS 등)가 패널에 다소 포함된 점이 아쉽습니다.  ★위반: [gpt] 주변 패널에 “COMMUNICATIONS” 등 프롬프트가 허용하지 않은 추가 가독성 문자가 노출되어 있다."
   },
   {
    "label": "A",
    "score": 861,
    "verdict_ko": "요구된 텍스트는 정확히 렌더링되었으나, 가장 중요한 프레이밍 지시사항인 '카메라와 직각을 이루는 정면 구도(square to the camera)'를 무시하고 비스듬한 각도로 렌더링하여 큰 감점을 받았습니다.  ★위반: [gemini-pro] 지정된 카메라 축(스크린 전면이 카메라와 직각을 이루어야 함)을 위반하고 비스듬한 시점으로 렌더링됨. / [gpt] 콘솔에 “ON”과 “OFF”라는 허용되지 않은 추가 가독성 문자가 노출되어 있다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/37154490-a4ed-4b37-9092-489c054d8a72/episodes/cc1d3d0b-626e-4f20-bc33-f5f9e2c6ee7e/images/background_chain/L11B02.png",
    "asset_id": "eb7f72c2-429d-4eb2-880a-e3c8b93c34f5",
    "role": "location_plate"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": true,
  "shot_run_uid": "06a9bd7a-d9f4-7cc2-9146-c6b5848e48ae",
  "ref_mode": "플레이트만 (배경 전용)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S1sh1::cine": {
  "applied": true,
  "attempted_at": "2026-09-05T08:50:30.100794+00:00",
  "fingerprint": "392ca54f640d9c01a85712868a796b908d91999abdb9d8c20c175976550c7c8e",
  "fingerprint_version": 2,
  "provider": "grok",
  "endpoint": "openrouter/chat-completions",
  "model": "x-ai/grok-imagine-image-2.0",
  "pack": "24.202608252115",
  "source_file": "S1sh1_sel.png",
  "source_sha256": "76739e6d3b7134c06e4b1791b068e7f4b830587443a4e039439f3c2983fd073b",
  "file": "S1sh1_cine.png",
  "staged_sha256": "1ef1236869a78a15447a8ca29794485b1a1a0dec91a248e1226035b84b7ce537",
  "latency_ms": 13542
 },
 "S1sh9::signage": {
  "fp": "ec003374a838cc15",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "era_assess::79cf163614b07aec": {
  "subjects": [],
  "subject_text": "산악 접촉소 지하실 내부\n바위벽 뒤에 숨겨진 작은 지하실. 벽에는 둥근 통신판이 있고 천장에는 산맥 범위를 표시하는 지도가 설치되어 있다.",
  "identity": "canonical",
  "scope_id": "L11",
  "scope_role": "location_interior",
  "scope_sha": "7ff131d7d64f46d4"
 },
 "S1sh9::bgfirst_bg": {
  "input_fingerprint": "e7f264ed77c003fd",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 장총의 총구가 토니의 가슴 정중앙에 닿은 찰나\n\nLOCATION (lock): Inside the mountain contact post's underground control room, in the open floor area between the armed guard and the captive. Red emergency lighting fills the room.\n\nTIME OF DAY (lock): twilight.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: 토니(앤서니 로저스) in the middle-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: long rifle (muzzle touching the exact center of Tony's chest) — The barrel is seen obliquely from its side, entering from frame left and pointing into Tony's chest rather than toward the camera; used as Forms the converging threat line and the sharp foreground focus.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The contact-post interior remains under the stated red lighting, rendered with low-key contrast and restrained highlights.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 장총의 총구가 토니의 가슴 정중앙에 닿은 찰나\n\nLOCATION (lock): Inside the mountain contact post's underground control room, in the open floor area between the armed guard and the captive. Red emergency lighting fills the room.\n\nTIME OF DAY (lock): twilight.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: 토니(앤서니 로저스) in the middle-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: long rifle (muzzle touching the exact center of Tony's chest) — The barrel is seen obliquely from its side, entering from frame left and pointing into Tony's chest rather than toward the camera; used as Forms the converging threat line and the sharp foreground focus.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The contact-post interior remains under the stated red lighting, rendered with low-key contrast and restrained highlights.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/37154490-a4ed-4b37-9092-489c054d8a72/images/cc1d3d0b-626e-4f20-bc33-f5f9e2c6ee7e/scene/recipe/S1sh9__bgfirst_bg.png",
  "asset_id": "15774562-bc78-41d4-8cb3-19b7f4b41044",
  "input_asset_ids": [
   "4dfacf07-54c0-433f-b7b8-1d0a055d5325",
   "eb7f72c2-429d-4eb2-880a-e3c8b93c34f5"
  ]
 },
 "S1sh9": {
  "input_fingerprint": "18eb5c0c26d7f5f7",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — Wyoming, United States, 2419; all people are English-speaking Americans unless stated. TIME OF DAY (lock): twilight.\n\nSHOT TEXT (authoritative, Korean): 장총의 총구가 토니의 가슴 정중앙에 닿은 찰나\n\nLOCATION (lock): Inside the mountain contact post's underground control room, in the open floor area between the armed guard and the captive. Red emergency lighting fills the room. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: 토니(앤서니 로저스) in the middle-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: long rifle (muzzle touching the exact center of Tony's chest) — The barrel is seen obliquely from its side, entering from frame left and pointing into Tony's chest rather than toward the camera; used as Forms the converging threat line and the sharp foreground focus.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The contact-post interior remains under the stated red lighting, rendered with low-key contrast and restrained highlights.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The rock door is closed, the interior lighting is red, and the waterproof-pouched phone rests inside the metal locker. 토니(앤서니 로저스): He remains with the long-gun muzzle pressed to the center of his chest.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 토니(앤서니 로저스) (미국인 남성, 30대 중반, 자연스러운 성인 남성 얼굴, 구체적으로 명시되지 않은 자연색 머리카락) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — Wyoming, United States, 2419; all people are English-speaking Americans unless stated. TIME OF DAY (lock): twilight.\n\nSHOT TEXT (authoritative, Korean): 장총의 총구가 토니의 가슴 정중앙에 닿은 찰나\n\nLOCATION (lock): Inside the mountain contact post's underground control room, in the open floor area between the armed guard and the captive. Red emergency lighting fills the room. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: 토니(앤서니 로저스) in the middle-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: long rifle (muzzle touching the exact center of Tony's chest) — The barrel is seen obliquely from its side, entering from frame left and pointing into Tony's chest rather than toward the camera; used as Forms the converging threat line and the sharp foreground focus.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The contact-post interior remains under the stated red lighting, rendered with low-key contrast and restrained highlights.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The rock door is closed, the interior lighting is red, and the waterproof-pouched phone rests inside the metal locker. 토니(앤서니 로저스): He remains with the long-gun muzzle pressed to the center of his chest.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 토니(앤서니 로저스) (미국인 남성, 30대 중반, 자연스러운 성인 남성 얼굴, 구체적으로 명시되지 않은 자연색 머리카락) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — Wyoming, United States, 2419; all people are English-speaking Americans unless stated. TIME OF DAY (lock): twilight.\n\nSHOT TEXT (authoritative, Korean): 장총의 총구가 토니의 가슴 정중앙에 닿은 찰나\n\nLOCATION (lock): Inside the mountain contact post's underground control room, in the open floor area between the armed guard and the captive. Red emergency lighting fills the room. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: 토니(앤서니 로저스) in the middle-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: long rifle (muzzle touching the exact center of Tony's chest) — The barrel is seen obliquely from its side, entering from frame left and pointing into Tony's chest rather than toward the camera; used as Forms the converging threat line and the sharp foreground focus.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The contact-post interior remains under the stated red lighting, rendered with low-key contrast and restrained highlights.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The rock door is closed, the interior lighting is red, and the waterproof-pouched phone rests inside the metal locker. 토니(앤서니 로저스): He remains with the long-gun muzzle pressed to the center of his chest.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 토니(앤서니 로저스) (미국인 남성, 30대 중반, 자연스러운 성인 남성 얼굴, 구체적으로 명시되지 않은 자연색 머리카락) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/37154490-a4ed-4b37-9092-489c054d8a72/images/cc1d3d0b-626e-4f20-bc33-f5f9e2c6ee7e/scene/recipe/S1sh9__bgfirst_bg.png",
     "asset_id": "15774562-bc78-41d4-8cb3-19b7f4b41044",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/37154490-a4ed-4b37-9092-489c054d8a72/images/cc1d3d0b-626e-4f20-bc33-f5f9e2c6ee7e/conti/conti_S1sh9.png",
     "asset_id": "4dfacf07-54c0-433f-b7b8-1d0a055d5325",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 토니(앤서니 로저스): the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:929851>",
     "asset_id": "1d3c6fe6-553b-4c00-b68f-025be6d786c1",
     "role": "character_ref"
    },
    {
     "label": "PROP REFERENCE — 장총: the exact object appearing in this shot; match its look, material and wear exactly.",
     "path": "<bytes:785972>",
     "asset_id": "3c9dd37a-da89-4fa6-a8b0-83c72e84ff1f",
     "role": "prop_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/37154490-a4ed-4b37-9092-489c054d8a72/episodes/cc1d3d0b-626e-4f20-bc33-f5f9e2c6ee7e/images/background_chain/L11B02.png",
     "asset_id": "eb7f72c2-429d-4eb2-880a-e3c8b93c34f5",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 토니(앤서니 로저스): the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:929851>",
     "asset_id": "1d3c6fe6-553b-4c00-b68f-025be6d786c1",
     "role": "character_ref"
    },
    {
     "label": "PROP REFERENCE — 장총: the exact object appearing in this shot; match its look, material and wear exactly.",
     "path": "<bytes:785972>",
     "asset_id": "3c9dd37a-da89-4fa6-a8b0-83c72e84ff1f",
     "role": "prop_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "initial_roll_all_fail": true,
  "readings": [
   {
    "label": "A",
    "direction": "장총이 화면 좌측에서 진입해 토니의 가슴을 향하고 있으며 총구가 가슴 중심 부근에 닿아 있습니다. 토니는 프레임 우측 밖을 응시하고 있습니다.",
    "built_space": "참조 이미지의 지하 통제실 배경이 잘 적용되어 있으며, 붉은 조명, 좌측의 암석 문, 우측의 장비 및 책상 등이 제 위치에 구현되어 있습니다.",
    "entities": "토니의 인물 얼굴과 복장은 참조와 정확히 일치합니다. 장총 역시 참조 이미지의 조준경과 양각대를 포함한 외형을 띠고 있습니다.",
    "hard_violations": [
     "[gemini-pro] 지지하는 손이나 물체 없이 무거운 장총이 허공에 완벽하게 떠 있음",
     "[gemini-pro] 장총의 총구 끝부분이 토니의 가슴에 부착된 장비 패널과 물리적으로 융합되어 왜곡됨",
     "[gemini-pro] 지시된 클로즈업 샷(close-up)이 아닌 미디엄 샷 규모로 인물의 허벅지까지 노출되도록 프레임을 넓힘",
     "[gpt] 오른쪽 콘솔에 ‘MUNITIONS WARNING’으로 읽히는 문구가 선명하게 보여 ‘읽을 수 있는 글자 없음’ 조건을 위반한다."
    ],
    "physics": "장총을 잡고 있는 사람의 손이나 총을 받치는 어떤 지지대도 보이지 않아 무기가 공중에 떠 있는 물리적으로 불가능한 상태입니다."
   },
   {
    "label": "B",
    "direction": "장총의 총열이 프레임 좌측에서 들어와 토니의 가슴 정중앙을 향합니다. 토니는 시선을 아래로 떨어뜨려 자신의 가슴 쪽을 보고 있습니다.",
    "built_space": "지하 통제실의 붉은 조명과 배경 구조물(제어판, 암석 등)이 위치 참조에 맞게 구현되어 있습니다.",
    "entities": "토니의 외모와 의상은 참조와 일치합니다. 장총의 전면부 총열이 묘사되어 있습니다.",
    "hard_violations": [
     "[gemini-pro] 장총을 잡고 있는 손이나 지지 수단이 없어 무기가 허공에 떠 있음",
     "[gemini-pro] 총구 부분이 토니의 가슴에 달린 장비 안으로 완전히 파고들어 한 덩어리로 융합되는 물리적 불가능 상태",
     "[gpt] 왼쪽 문 옆 표지의 문구와 뒤쪽 화면의 ‘ALERT’ 등 읽을 수 있는 글자가 보여 ‘읽을 수 있는 글자 없음’ 조건을 위반한다."
    ],
    "physics": "프레임 안에 장총을 지탱하는 손이 전혀 묘사되지 않아 총기가 허공에 떠 있으며, 총구가 캐릭터의 가슴 부착물과 녹아들어 형태가 붕괴되었습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 0,
        "verdict_ko": "장총이 어떤 지지대나 손도 없이 허공에 떠 있고 총구가 캐릭터의 가슴 장비와 융합되는 치명적인 물리적 오류가 발생했으며, 프레이밍도 클로즈업이 아닌 미디엄 샷으로 넓게 렌더링되었습니다."
       },
       {
        "label": "B",
        "score": 0,
        "verdict_ko": "요청된 클로즈업 프레이밍에는 더 가깝지만, 후보 A와 마찬가지로 장총이 지지대 없이 허공에 떠 있고 총구 끝부분이 가슴 장비에 물리적으로 녹아드는 치명적 오류가 존재하여 사용할 수 없습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "장총이 화면 좌측에서 진입해 토니의 가슴을 향하고 있으며 총구가 가슴 중심 부근에 닿아 있습니다. 토니는 프레임 우측 밖을 응시하고 있습니다.",
        "built_space": "참조 이미지의 지하 통제실 배경이 잘 적용되어 있으며, 붉은 조명, 좌측의 암석 문, 우측의 장비 및 책상 등이 제 위치에 구현되어 있습니다.",
        "entities": "토니의 인물 얼굴과 복장은 참조와 정확히 일치합니다. 장총 역시 참조 이미지의 조준경과 양각대를 포함한 외형을 띠고 있습니다.",
        "hard_violations": [
         "지지하는 손이나 물체 없이 무거운 장총이 허공에 완벽하게 떠 있음",
         "장총의 총구 끝부분이 토니의 가슴에 부착된 장비 패널과 물리적으로 융합되어 왜곡됨",
         "지시된 클로즈업 샷(close-up)이 아닌 미디엄 샷 규모로 인물의 허벅지까지 노출되도록 프레임을 넓힘"
        ],
        "physics": "장총을 잡고 있는 사람의 손이나 총을 받치는 어떤 지지대도 보이지 않아 무기가 공중에 떠 있는 물리적으로 불가능한 상태입니다."
       },
       {
        "label": "B",
        "direction": "장총의 총열이 프레임 좌측에서 들어와 토니의 가슴 정중앙을 향합니다. 토니는 시선을 아래로 떨어뜨려 자신의 가슴 쪽을 보고 있습니다.",
        "built_space": "지하 통제실의 붉은 조명과 배경 구조물(제어판, 암석 등)이 위치 참조에 맞게 구현되어 있습니다.",
        "entities": "토니의 외모와 의상은 참조와 일치합니다. 장총의 전면부 총열이 묘사되어 있습니다.",
        "hard_violations": [
         "장총을 잡고 있는 손이나 지지 수단이 없어 무기가 허공에 떠 있음",
         "총구 부분이 토니의 가슴에 달린 장비 안으로 완전히 파고들어 한 덩어리로 융합되는 물리적 불가능 상태"
        ],
        "physics": "프레임 안에 장총을 지탱하는 손이 전혀 묘사되지 않아 총기가 허공에 떠 있으며, 총구가 캐릭터의 가슴 부착물과 녹아들어 형태가 붕괴되었습니다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 0,
        "verdict_ko": "장총이 어떤 지지대나 손도 없이 허공에 떠 있고 총구가 캐릭터의 가슴 장비와 융합되는 치명적인 물리적 오류가 발생했으며, 프레이밍도 클로즈업이 아닌 미디엄 샷으로 넓게 렌더링되었습니다."
       },
       {
        "label": "B",
        "score": 0,
        "verdict_ko": "요청된 클로즈업 프레이밍에는 더 가깝지만, 후보 A와 마찬가지로 장총이 지지대 없이 허공에 떠 있고 총구 끝부분이 가슴 장비에 물리적으로 녹아드는 치명적 오류가 존재하여 사용할 수 없습니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "A",
        "direction": "장총이 화면 좌측에서 진입해 토니의 가슴을 향하고 있으며 총구가 가슴 중심 부근에 닿아 있습니다. 토니는 프레임 우측 밖을 응시하고 있습니다.",
        "built_space": "참조 이미지의 지하 통제실 배경이 잘 적용되어 있으며, 붉은 조명, 좌측의 암석 문, 우측의 장비 및 책상 등이 제 위치에 구현되어 있습니다.",
        "entities": "토니의 인물 얼굴과 복장은 참조와 정확히 일치합니다. 장총 역시 참조 이미지의 조준경과 양각대를 포함한 외형을 띠고 있습니다.",
        "hard_violations": [
         "지지하는 손이나 물체 없이 무거운 장총이 허공에 완벽하게 떠 있음",
         "장총의 총구 끝부분이 토니의 가슴에 부착된 장비 패널과 물리적으로 융합되어 왜곡됨",
         "지시된 클로즈업 샷(close-up)이 아닌 미디엄 샷 규모로 인물의 허벅지까지 노출되도록 프레임을 넓힘"
        ],
        "physics": "장총을 잡고 있는 사람의 손이나 총을 받치는 어떤 지지대도 보이지 않아 무기가 공중에 떠 있는 물리적으로 불가능한 상태입니다."
       },
       {
        "label": "B",
        "direction": "장총의 총열이 프레임 좌측에서 들어와 토니의 가슴 정중앙을 향합니다. 토니는 시선을 아래로 떨어뜨려 자신의 가슴 쪽을 보고 있습니다.",
        "built_space": "지하 통제실의 붉은 조명과 배경 구조물(제어판, 암석 등)이 위치 참조에 맞게 구현되어 있습니다.",
        "entities": "토니의 외모와 의상은 참조와 일치합니다. 장총의 전면부 총열이 묘사되어 있습니다.",
        "hard_violations": [
         "장총을 잡고 있는 손이나 지지 수단이 없어 무기가 허공에 떠 있음",
         "총구 부분이 토니의 가슴에 달린 장비 안으로 완전히 파고들어 한 덩어리로 융합되는 물리적 불가능 상태"
        ],
        "physics": "프레임 안에 장총을 지탱하는 손이 전혀 묘사되지 않아 총기가 허공에 떠 있으며, 총구가 캐릭터의 가슴 부착물과 녹아들어 형태가 붕괴되었습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "토니를 중간 오른쪽에 둔 근접 구도와 왼쪽에서 비스듬히 들어와 가슴 중앙에 닿는 총구가 핵심 순간을 가장 정확히 구현했지만, 배경의 읽을 수 있는 문구는 명시적 금지 사항을 위반한다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "총구 방향과 접촉점은 맞지만 허벅지까지 드러나는 중간 숏으로 지나치게 넓고 장총 전체를 과도하게 부각했으며, 읽을 수 있는 콘솔 문구도 남아 있다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "장총의 총열이 프레임 왼쪽에서 오른쪽으로 비스듬히 뻗어 토니의 흉골 중앙, 정확히는 가슴 장비 위 중심선에 총구가 닿는다. 토니는 고개와 시선을 아래쪽 총구 및 접촉점으로 향하고 있다.",
        "built_space": "닫힌 암석문 1개가 왼쪽에 있고, 그 옆으로 금속 장비 캐비닛과 제어 패널들이 보인다. 뒤쪽에는 제어용 화면과 작업대, 오른쪽에는 투명 장비 보관함과 콘솔이 있으며 위치 사진의 지하 통제실 구조와 대체로 일치한다. 토니는 고정 설비가 아닌 중앙의 열린 바닥에 서 있어 배치가 성립하며, 투명 덮개의 반사는 카메라와 적색 조명의 위치상 가능하다.",
        "entities": "토니는 자연색의 짧은 갈색 머리와 수염, 30대 중반의 백인계 미국 남성으로 보이며 얼굴과 체격은 인물 참조와 상당히 가깝다. 낡은 어두운 전술복과 하네스는 참조와 유사하지만 참조의 헬멧은 없다. 보이는 장총 전방부는 검은 금속 총열·통풍 핸드가드·총구 장치를 갖췄으나, 근접 크롭 때문에 참조 장총의 스코프와 몸통은 확인되지 않는다. 다른 사람은 보이지 않는다.",
        "hard_violations": [
         "왼쪽 문 옆 표지의 문구와 뒤쪽 화면의 ‘ALERT’ 등 읽을 수 있는 글자가 보여 ‘읽을 수 있는 글자 없음’ 조건을 위반한다."
        ],
        "physics": "토니의 하체는 프레임 밖이지만 수직 자세와 열린 바닥의 위치로 보아 바닥에 서 있는 상태가 자연스럽다. 장총은 왼쪽 프레임 밖으로 이어져 화면 밖 사수가 지지하는 배치로 읽히며, 총구는 가슴 장비에 실제로 접촉해 멈춰 있다."
       },
       {
        "label": "B",
        "direction": "장총은 왼쪽에서 오른쪽으로 수평에 가깝게 향하며 총구가 토니의 흉골 중앙 장비 위에 닿는다. 토니의 시선은 총구 접촉점보다 약간 높고 더 왼쪽인 화면 밖 사수 방향을 향한다.",
        "built_space": "왼쪽에 닫힌 암석문 1개, 중앙 왼쪽에 금속 캐비닛과 패널, 뒤쪽에 작업대와 대형 화면, 오른쪽에 투명 장비 보관함과 콘솔이 보인다. 토니는 중앙 원형 설비 앞의 열린 바닥에 서 있어 공간 배치는 가능하다. 다만 위치 사진의 핵심 대형 화면이 검게 꺼져 있어 통제실의 고정 배경 상태가 덜 충실하고, 넓은 프레이밍 때문에 원형 설비와 장총 몸통이 크게 노출된다. 투명 덮개의 반사는 광원과 카메라 방향상 가능하다.",
        "entities": "토니는 짧은 갈색 머리와 수염, 30대 중반의 백인계 미국 남성으로 보이고 얼굴·체격·낡은 전술복은 참조와 유사하지만 헬멧이 빠졌다. 장총은 검은 정밀소총으로서 스코프, 리시버, 통풍 핸드가드, 긴 총열과 양각대를 보여 참조 소품의 종류와 외형을 A보다 완전하게 드러낸다. 다른 사람은 보이지 않는다.",
        "hard_violations": [
         "오른쪽 콘솔에 ‘MUNITIONS WARNING’으로 읽히는 문구가 선명하게 보여 ‘읽을 수 있는 글자 없음’ 조건을 위반한다."
        ],
        "physics": "토니는 허벅지 중간에서 잘렸지만 자세상 바닥에 서 있는 것으로 읽힌다. 장총은 리시버가 왼쪽 프레임 밖으로 계속 이어져 화면 밖 사수가 지지할 수 있으나, 펼쳐진 양각대는 바닥이나 탁자에 닿지 않고 매달려 있어 운용 자세가 어색하다. 그래도 총 자체는 프레임 밖 지지점이 있어 완전히 무지지 상태로 단정되지는 않는다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "토니를 중간 오른쪽에 둔 근접 구도와 왼쪽에서 비스듬히 들어와 가슴 중앙에 닿는 총구가 핵심 순간을 가장 정확히 구현했지만, 배경의 읽을 수 있는 문구는 명시적 금지 사항을 위반한다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "총구 방향과 접촉점은 맞지만 허벅지까지 드러나는 중간 숏으로 지나치게 넓고 장총 전체를 과도하게 부각했으며, 읽을 수 있는 콘솔 문구도 남아 있다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "장총의 총열이 프레임 왼쪽에서 오른쪽으로 비스듬히 뻗어 토니의 흉골 중앙, 정확히는 가슴 장비 위 중심선에 총구가 닿는다. 토니는 고개와 시선을 아래쪽 총구 및 접촉점으로 향하고 있다.",
        "built_space": "닫힌 암석문 1개가 왼쪽에 있고, 그 옆으로 금속 장비 캐비닛과 제어 패널들이 보인다. 뒤쪽에는 제어용 화면과 작업대, 오른쪽에는 투명 장비 보관함과 콘솔이 있으며 위치 사진의 지하 통제실 구조와 대체로 일치한다. 토니는 고정 설비가 아닌 중앙의 열린 바닥에 서 있어 배치가 성립하며, 투명 덮개의 반사는 카메라와 적색 조명의 위치상 가능하다.",
        "entities": "토니는 자연색의 짧은 갈색 머리와 수염, 30대 중반의 백인계 미국 남성으로 보이며 얼굴과 체격은 인물 참조와 상당히 가깝다. 낡은 어두운 전술복과 하네스는 참조와 유사하지만 참조의 헬멧은 없다. 보이는 장총 전방부는 검은 금속 총열·통풍 핸드가드·총구 장치를 갖췄으나, 근접 크롭 때문에 참조 장총의 스코프와 몸통은 확인되지 않는다. 다른 사람은 보이지 않는다.",
        "hard_violations": [
         "왼쪽 문 옆 표지의 문구와 뒤쪽 화면의 ‘ALERT’ 등 읽을 수 있는 글자가 보여 ‘읽을 수 있는 글자 없음’ 조건을 위반한다."
        ],
        "physics": "토니의 하체는 프레임 밖이지만 수직 자세와 열린 바닥의 위치로 보아 바닥에 서 있는 상태가 자연스럽다. 장총은 왼쪽 프레임 밖으로 이어져 화면 밖 사수가 지지하는 배치로 읽히며, 총구는 가슴 장비에 실제로 접촉해 멈춰 있다."
       },
       {
        "label": "A",
        "direction": "장총은 왼쪽에서 오른쪽으로 수평에 가깝게 향하며 총구가 토니의 흉골 중앙 장비 위에 닿는다. 토니의 시선은 총구 접촉점보다 약간 높고 더 왼쪽인 화면 밖 사수 방향을 향한다.",
        "built_space": "왼쪽에 닫힌 암석문 1개, 중앙 왼쪽에 금속 캐비닛과 패널, 뒤쪽에 작업대와 대형 화면, 오른쪽에 투명 장비 보관함과 콘솔이 보인다. 토니는 중앙 원형 설비 앞의 열린 바닥에 서 있어 공간 배치는 가능하다. 다만 위치 사진의 핵심 대형 화면이 검게 꺼져 있어 통제실의 고정 배경 상태가 덜 충실하고, 넓은 프레이밍 때문에 원형 설비와 장총 몸통이 크게 노출된다. 투명 덮개의 반사는 광원과 카메라 방향상 가능하다.",
        "entities": "토니는 짧은 갈색 머리와 수염, 30대 중반의 백인계 미국 남성으로 보이고 얼굴·체격·낡은 전술복은 참조와 유사하지만 헬멧이 빠졌다. 장총은 검은 정밀소총으로서 스코프, 리시버, 통풍 핸드가드, 긴 총열과 양각대를 보여 참조 소품의 종류와 외형을 A보다 완전하게 드러낸다. 다른 사람은 보이지 않는다.",
        "hard_violations": [
         "오른쪽 콘솔에 ‘MUNITIONS WARNING’으로 읽히는 문구가 선명하게 보여 ‘읽을 수 있는 글자 없음’ 조건을 위반한다."
        ],
        "physics": "토니는 허벅지 중간에서 잘렸지만 자세상 바닥에 서 있는 것으로 읽힌다. 장총은 리시버가 왼쪽 프레임 밖으로 계속 이어져 화면 밖 사수가 지지할 수 있으나, 펼쳐진 양각대는 바닥이나 탁자에 닿지 않고 매달려 있어 운용 자세가 어색하다. 그래도 총 자체는 프레임 밖 지지점이 있어 완전히 무지지 상태로 단정되지는 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt"
   ],
   "normalized": {
    "A": 0.571,
    "B": 1.0
   },
   "adjusted": {
    "A": 0.321,
    "B": 0.75
   },
   "violations": {
    "A": [
     "[gemini-pro] 지지하는 손이나 물체 없이 무거운 장총이 허공에 완벽하게 떠 있음",
     "[gemini-pro] 장총의 총구 끝부분이 토니의 가슴에 부착된 장비 패널과 물리적으로 융합되어 왜곡됨",
     "[gemini-pro] 지시된 클로즈업 샷(close-up)이 아닌 미디엄 샷 규모로 인물의 허벅지까지 노출되도록 프레임을 넓힘",
     "[gpt] 오른쪽 콘솔에 ‘MUNITIONS WARNING’으로 읽히는 문구가 선명하게 보여 ‘읽을 수 있는 글자 없음’ 조건을 위반한다."
    ],
    "B": [
     "[gemini-pro] 장총을 잡고 있는 손이나 지지 수단이 없어 무기가 허공에 떠 있음",
     "[gemini-pro] 총구 부분이 토니의 가슴에 달린 장비 안으로 완전히 파고들어 한 덩어리로 융합되는 물리적 불가능 상태",
     "[gpt] 왼쪽 문 옆 표지의 문구와 뒤쪽 화면의 ‘ALERT’ 등 읽을 수 있는 글자가 보여 ‘읽을 수 있는 글자 없음’ 조건을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt": "B"
   },
   "agreed": true
  },
  "totals": {
   "A": 321,
   "B": 750
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 321,
    "verdict_ko": "장총이 어떤 지지대나 손도 없이 허공에 떠 있고 총구가 캐릭터의 가슴 장비와 융합되는 치명적인 물리적 오류가 발생했으며, 프레이밍도 클로즈업이 아닌 미디엄 샷으로 넓게 렌더링되었습니다.  ★위반: [gemini-pro] 지지하는 손이나 물체 없이 무거운 장총이 허공에 완벽하게 떠 있음 / [gemini-pro] 장총의 총구 끝부분이 토니의 가슴에 부착된 장비 패널과 물리적으로 융합되어 왜곡됨 / [gemini-pro] 지시된 클로즈업 샷(close-up)이 아닌 미디엄 샷 규모로 인물의 허벅지까지 노출되도록 프레임을 넓힘 / [gpt] 오른쪽 콘솔에 ‘MUNITIONS WARNING’으로 읽히는 문구가 선명하게 보여 ‘읽을 수 있는 글자 없음’ 조건을 위반한다."
   },
   {
    "label": "B",
    "score": 750,
    "verdict_ko": "요청된 클로즈업 프레이밍에는 더 가깝지만, 후보 A와 마찬가지로 장총이 지지대 없이 허공에 떠 있고 총구 끝부분이 가슴 장비에 물리적으로 녹아드는 치명적 오류가 존재하여 사용할 수 없습니다.  ★위반: [gemini-pro] 장총을 잡고 있는 손이나 지지 수단이 없어 무기가 허공에 떠 있음 / [gemini-pro] 총구 부분이 토니의 가슴에 달린 장비 안으로 완전히 파고들어 한 덩어리로 융합되는 물리적 불가능 상태 / [gpt] 왼쪽 문 옆 표지의 문구와 뒤쪽 화면의 ‘ALERT’ 등 읽을 수 있는 글자가 보여 ‘읽을 수 있는 글자 없음’ 조건을 위반한다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/37154490-a4ed-4b37-9092-489c054d8a72/episodes/cc1d3d0b-626e-4f20-bc33-f5f9e2c6ee7e/images/background_chain/L11B02.png",
    "asset_id": "eb7f72c2-429d-4eb2-880a-e3c8b93c34f5",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 토니(앤서니 로저스): the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:929851>",
    "asset_id": "1d3c6fe6-553b-4c00-b68f-025be6d786c1",
    "role": "character_ref"
   },
   {
    "label": "PROP REFERENCE — 장총: the exact object appearing in this shot; match its look, material and wear exactly.",
    "path": "<bytes:785972>",
    "asset_id": "3c9dd37a-da89-4fa6-a8b0-83c72e84ff1f",
    "role": "prop_ref"
   }
  ],
  "critique_skipped": true,
  "needs_reshoot": true,
  "shot_run_produced": true,
  "shot_run_uid": "06a9bd7e-3bc1-7509-87e2-e55d7ab87404",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/37154490-a4ed-4b37-9092-489c054d8a72/images/cc1d3d0b-626e-4f20-bc33-f5f9e2c6ee7e/scene/recipe/S1sh9__bgfirst_bg.png",
   "bg_asset_id": "15774562-bc78-41d4-8cb3-19b7f4b41044",
   "bg_record_key": "S1sh9::bgfirst_bg",
   "chain_winner": false,
   "authority": "plate"
  },
  "ref_mode": "플레이트+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S1sh9::cine": {
  "applied": true,
  "attempted_at": "2026-09-05T08:53:21.880140+00:00",
  "fingerprint": "c70dd502dcee4616521a7bb3106d669432c96196597a1463c2eb4a9ff47744cc",
  "fingerprint_version": 2,
  "provider": "grok",
  "endpoint": "openrouter/chat-completions",
  "model": "x-ai/grok-imagine-image-2.0",
  "pack": "24.202608252115",
  "source_file": "S1sh9_sel.png",
  "source_sha256": "6472f38770593bc476718e555b8946703dc43e52474478e63f6863e40184a877",
  "file": "S1sh9_cine.png",
  "staged_sha256": "c87d5b114453151c3f530577251ce6ad1414647c69deb5328f45d4b12750bcab",
  "latency_ms": 13274
 },
 "S1sh12::signage": {
  "fp": "83c16c215c0317b4",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S1sh12": {
  "input_fingerprint": "e2959842bec3e4ab",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — Wyoming, United States, 2419; all people are English-speaking Americans unless stated. TIME OF DAY (lock): twilight.\n\nSHOT TEXT (authoritative, Korean): 지도 스크린 위로 붉은 점들이 원형으로 밀집된 화면 클로즈업\n\nLOCATION (lock): Inside the mountain contact post's underground control room, at the tactical map display. Red alert lighting and the map screen illuminate the station. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: map screen (showing red points tightening into a circular encirclement) — The active map face is visible from a steep overhead angle, with the converging red-point circle centered in the displayed map; used as Terminal focal surface for the scene's encirclement reveal.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Low-key red interior illumination supports the concentrated warning-red points on the map without softening their graphic severity.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The rock door remains closed under red interior light, and the waterproof-pouched phone remains in the metal locker. Red points tighten into an encircling formation on the map display.\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 토니(앤서니 로저스) (미국인 남성, 30대 중반, 자연스러운 성인 남성 얼굴, 구체적으로 명시되지 않은 자연색 머리카락). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — Wyoming, United States, 2419; all people are English-speaking Americans unless stated. TIME OF DAY (lock): twilight.\n\nSHOT TEXT (authoritative, Korean): 지도 스크린 위로 붉은 점들이 원형으로 밀집된 화면 클로즈업\n\nLOCATION (lock): Inside the mountain contact post's underground control room, at the tactical map display. Red alert lighting and the map screen illuminate the station. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: map screen (showing red points tightening into a circular encirclement) — The active map face is visible from a steep overhead angle, with the converging red-point circle centered in the displayed map; used as Terminal focal surface for the scene's encirclement reveal.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Low-key red interior illumination supports the concentrated warning-red points on the map without softening their graphic severity.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The rock door remains closed under red interior light, and the waterproof-pouched phone remains in the metal locker. Red points tighten into an encircling formation on the map display.\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 토니(앤서니 로저스) (미국인 남성, 30대 중반, 자연스러운 성인 남성 얼굴, 구체적으로 명시되지 않은 자연색 머리카락). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — Wyoming, United States, 2419; all people are English-speaking Americans unless stated. TIME OF DAY (lock): twilight.\n\nSHOT TEXT (authoritative, Korean): 지도 스크린 위로 붉은 점들이 원형으로 밀집된 화면 클로즈업\n\nLOCATION (lock): Inside the mountain contact post's underground control room, at the tactical map display. Red alert lighting and the map screen illuminate the station. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: map screen (showing red points tightening into a circular encirclement) — The active map face is visible from a steep overhead angle, with the converging red-point circle centered in the displayed map; used as Terminal focal surface for the scene's encirclement reveal.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Low-key red interior illumination supports the concentrated warning-red points on the map without softening their graphic severity.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The rock door remains closed under red interior light, and the waterproof-pouched phone remains in the metal locker. Red points tighten into an encircling formation on the map display.\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 토니(앤서니 로저스) (미국인 남성, 30대 중반, 자연스러운 성인 남성 얼굴, 구체적으로 명시되지 않은 자연색 머리카락). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "initial_roll_all_fail": true,
  "readings": [
   {
    "label": "A",
    "direction": "카메라가 비스듬한 하향각으로 지도 스크린을 바라보고 있으며, 시선은 원형으로 배열된 붉은 점들에 집중됨.",
    "built_space": "금속 재질의 콘솔 표면에 평평하게 내장된 스크린이 프레임을 채우고 있음.",
    "entities": "붉은 점들이 방사형 원으로 배열된 지도 화면. 프롬프트에서 명시적으로 금지한 읽을 수 있는 영문 텍스트(FACILITY PERIMETER, BREACH POINTS 등)가 뚜렷하게 보임.",
    "hard_violations": [
     "[gemini-pro] 지시문에서 금지한 읽을 수 있는 텍스트가 화면에 명확하게 노출됨 (leaked markers/text 위반).",
     "[gpt] 지도 화면에 ‘FACILITY PERIMETER’, ‘BREACH POINTS’ 등 읽을 수 있는 영문 텍스트가 노출됨"
    ],
    "physics": "스크린은 콘솔에 단단히 고정되어 있으며 물리적으로 안정적임."
   },
   {
    "label": "B",
    "direction": "카메라가 비스듬한 하향각으로 지도 스크린을 바라보고 있으며, 화면 중앙의 빽빽한 붉은 점 원형을 향함.",
    "built_space": "다양한 버튼과 다이얼이 있는 금속 제어판에 내장된 모니터 공간.",
    "entities": "지형도 위로 붉은 점들이 빽빽하게 원형으로 밀집된 스크린. 샷 텍스트에 언급되지 않은 인물의 손과 팔이 화면 좌측 하단에 등장함.",
    "hard_violations": [
     "[gemini-pro] 샷 텍스트에 전혀 명시되지 않은 인물의 신체 일부(손과 팔)가 임의로 프레임에 추가됨 (invented people).",
     "[gpt] 쇼트 텍스트가 보여 주지 않은 인물의 손과 몸 일부가 추가됨"
    ],
    "physics": "추가된 인물의 손이 콘솔과 스크린 경계면에 자연스럽게 얹혀 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "지시문에 없는 인물을 추가하지 않아 B보다 낫지만, 철저히 금지된 가독성 높은 텍스트를 화면에 그대로 노출하여 실패한 컷입니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "붉은 점들이 밀집된 원형의 시각적 묘사는 더 우수하지만, 샷 텍스트에 없는 인물의 신체(손과 팔)를 임의로 추가하여 치명적인 규정 위반을 범했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "카메라가 비스듬한 하향각으로 지도 스크린을 바라보고 있으며, 시선은 원형으로 배열된 붉은 점들에 집중됨.",
        "built_space": "금속 재질의 콘솔 표면에 평평하게 내장된 스크린이 프레임을 채우고 있음.",
        "entities": "붉은 점들이 방사형 원으로 배열된 지도 화면. 프롬프트에서 명시적으로 금지한 읽을 수 있는 영문 텍스트(FACILITY PERIMETER, BREACH POINTS 등)가 뚜렷하게 보임.",
        "hard_violations": [
         "지시문에서 금지한 읽을 수 있는 텍스트가 화면에 명확하게 노출됨 (leaked markers/text 위반)."
        ],
        "physics": "스크린은 콘솔에 단단히 고정되어 있으며 물리적으로 안정적임."
       },
       {
        "label": "B",
        "direction": "카메라가 비스듬한 하향각으로 지도 스크린을 바라보고 있으며, 화면 중앙의 빽빽한 붉은 점 원형을 향함.",
        "built_space": "다양한 버튼과 다이얼이 있는 금속 제어판에 내장된 모니터 공간.",
        "entities": "지형도 위로 붉은 점들이 빽빽하게 원형으로 밀집된 스크린. 샷 텍스트에 언급되지 않은 인물의 손과 팔이 화면 좌측 하단에 등장함.",
        "hard_violations": [
         "샷 텍스트에 전혀 명시되지 않은 인물의 신체 일부(손과 팔)가 임의로 프레임에 추가됨 (invented people)."
        ],
        "physics": "추가된 인물의 손이 콘솔과 스크린 경계면에 자연스럽게 얹혀 있음."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "지시문에 없는 인물을 추가하지 않아 B보다 낫지만, 철저히 금지된 가독성 높은 텍스트를 화면에 그대로 노출하여 실패한 컷입니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "붉은 점들이 밀집된 원형의 시각적 묘사는 더 우수하지만, 샷 텍스트에 없는 인물의 신체(손과 팔)를 임의로 추가하여 치명적인 규정 위반을 범했습니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "A",
        "direction": "카메라가 비스듬한 하향각으로 지도 스크린을 바라보고 있으며, 시선은 원형으로 배열된 붉은 점들에 집중됨.",
        "built_space": "금속 재질의 콘솔 표면에 평평하게 내장된 스크린이 프레임을 채우고 있음.",
        "entities": "붉은 점들이 방사형 원으로 배열된 지도 화면. 프롬프트에서 명시적으로 금지한 읽을 수 있는 영문 텍스트(FACILITY PERIMETER, BREACH POINTS 등)가 뚜렷하게 보임.",
        "hard_violations": [
         "지시문에서 금지한 읽을 수 있는 텍스트가 화면에 명확하게 노출됨 (leaked markers/text 위반)."
        ],
        "physics": "스크린은 콘솔에 단단히 고정되어 있으며 물리적으로 안정적임."
       },
       {
        "label": "B",
        "direction": "카메라가 비스듬한 하향각으로 지도 스크린을 바라보고 있으며, 화면 중앙의 빽빽한 붉은 점 원형을 향함.",
        "built_space": "다양한 버튼과 다이얼이 있는 금속 제어판에 내장된 모니터 공간.",
        "entities": "지형도 위로 붉은 점들이 빽빽하게 원형으로 밀집된 스크린. 샷 텍스트에 언급되지 않은 인물의 손과 팔이 화면 좌측 하단에 등장함.",
        "hard_violations": [
         "샷 텍스트에 전혀 명시되지 않은 인물의 신체 일부(손과 팔)가 임의로 프레임에 추가됨 (invented people)."
        ],
        "physics": "추가된 인물의 손이 콘솔과 스크린 경계면에 자연스럽게 얹혀 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "가파른 오버헤드 클로즈업과 중앙의 붉은 포위 원은 맞지만, 인물을 요구하지 않은 쇼트에 손과 몸 일부를 추가해 사용할 수 없다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "단일 전술 지도 화면의 클로즈업, 중앙 포위 원, 붉은 경보 조명과 장소의 화면 양식은 가장 충실하지만, 명확히 읽히는 영문 표기가 금지 조건을 위반한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "무기, 시선, 뾰족한 물체는 없다. 붉은 점들은 지도 중앙의 빈 영역을 둘러싸는 원형 띠로 밀집해 있으며 포위의 목표는 화면 중앙 지역으로 읽힌다.",
        "built_space": "마모된 금속 콘솔에 매립된 활성 지도 화면 1개와 주변 버튼 패널 일부가 보인다. 카메라는 화면 위의 가파른 각도에 있고 화면이 프레임 대부분을 차지한다. 왼쪽 아래에는 사람의 상체 일부가, 아래쪽에는 손 1개가 화면 테두리에 놓여 있다. 반사나 중복 고정물 문제는 보이지 않는다.",
        "entities": "물리적인 전술 지도 화면과 다수의 붉은 점은 존재하며 원형 포위 대형도 중앙에 보인다. 그러나 쇼트 텍스트가 인물을 제시하지 않았는데 정체를 확인할 수 없는 사람의 손과 몸 일부가 추가되었다. 방수 파우치 전화기와 닫힌 암석문은 이 올바른 클로즈업 범위 밖이다.",
        "hard_violations": [
         "쇼트 텍스트가 보여 주지 않은 인물의 손과 몸 일부가 추가됨"
        ],
        "physics": "손은 아래쪽 화면 베젤에 직접 닿아 지지되고, 잘린 몸은 프레임 밖으로 이어져 있어 부유하지 않는다. 화면은 금속 콘솔에 고정되어 있다. 붉은 점은 화면에서 발광하는 표시이며 물리적으로 공중에 떠 있는 물체가 아니다."
       },
       {
        "label": "B",
        "direction": "시선이나 무기, 이동하는 몸은 없다. 방사형으로 배열된 붉은 점들이 지도 중앙부를 완전히 둘러싸며, 포위 목표는 원 안의 중앙 지도 지역으로 명확히 읽힌다.",
        "built_space": "마모된 붉은 금속 콘솔에 매립된 활성 지도 화면 1개만 클로즈업으로 보인다. 가파른 상부 각도이며 화면과 베젤이 프레임을 채운다. 인물, 의자, 문, 사물함은 보이지 않고 중복 고정물이나 불가능한 반사도 없다. 화면과 콘솔의 재질 및 붉은 조명은 참조 장소와 잘 이어진다.",
        "entities": "단일 전술 지도 화면, 지도 경계선, 중앙으로 조여진 붉은 포위 원이 모두 보이며 불필요한 사람이나 별도 소품은 없다. 다만 화면 좌측과 상단 및 우측에 영문 제목과 목록이 선명하게 읽혀 ‘읽을 수 있는 글자 금지’를 위반한다.",
        "hard_violations": [
         "지도 화면에 ‘FACILITY PERIMETER’, ‘BREACH POINTS’ 등 읽을 수 있는 영문 텍스트가 노출됨"
        ],
        "physics": "화면은 금속 콘솔에 정상적으로 매립·고정되어 있으며 지지되지 않은 물체나 신체는 없다. 붉은 점과 지도선은 활성 디스플레이 표면의 발광 정보로 표현되어 화면 위에 물리적으로 떠 있지 않는다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "가파른 오버헤드 클로즈업과 중앙의 붉은 포위 원은 맞지만, 인물을 요구하지 않은 쇼트에 손과 몸 일부를 추가해 사용할 수 없다."
       },
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "단일 전술 지도 화면의 클로즈업, 중앙 포위 원, 붉은 경보 조명과 장소의 화면 양식은 가장 충실하지만, 명확히 읽히는 영문 표기가 금지 조건을 위반한다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "무기, 시선, 뾰족한 물체는 없다. 붉은 점들은 지도 중앙의 빈 영역을 둘러싸는 원형 띠로 밀집해 있으며 포위의 목표는 화면 중앙 지역으로 읽힌다.",
        "built_space": "마모된 금속 콘솔에 매립된 활성 지도 화면 1개와 주변 버튼 패널 일부가 보인다. 카메라는 화면 위의 가파른 각도에 있고 화면이 프레임 대부분을 차지한다. 왼쪽 아래에는 사람의 상체 일부가, 아래쪽에는 손 1개가 화면 테두리에 놓여 있다. 반사나 중복 고정물 문제는 보이지 않는다.",
        "entities": "물리적인 전술 지도 화면과 다수의 붉은 점은 존재하며 원형 포위 대형도 중앙에 보인다. 그러나 쇼트 텍스트가 인물을 제시하지 않았는데 정체를 확인할 수 없는 사람의 손과 몸 일부가 추가되었다. 방수 파우치 전화기와 닫힌 암석문은 이 올바른 클로즈업 범위 밖이다.",
        "hard_violations": [
         "쇼트 텍스트가 보여 주지 않은 인물의 손과 몸 일부가 추가됨"
        ],
        "physics": "손은 아래쪽 화면 베젤에 직접 닿아 지지되고, 잘린 몸은 프레임 밖으로 이어져 있어 부유하지 않는다. 화면은 금속 콘솔에 고정되어 있다. 붉은 점은 화면에서 발광하는 표시이며 물리적으로 공중에 떠 있는 물체가 아니다."
       },
       {
        "label": "A",
        "direction": "시선이나 무기, 이동하는 몸은 없다. 방사형으로 배열된 붉은 점들이 지도 중앙부를 완전히 둘러싸며, 포위 목표는 원 안의 중앙 지도 지역으로 명확히 읽힌다.",
        "built_space": "마모된 붉은 금속 콘솔에 매립된 활성 지도 화면 1개만 클로즈업으로 보인다. 가파른 상부 각도이며 화면과 베젤이 프레임을 채운다. 인물, 의자, 문, 사물함은 보이지 않고 중복 고정물이나 불가능한 반사도 없다. 화면과 콘솔의 재질 및 붉은 조명은 참조 장소와 잘 이어진다.",
        "entities": "단일 전술 지도 화면, 지도 경계선, 중앙으로 조여진 붉은 포위 원이 모두 보이며 불필요한 사람이나 별도 소품은 없다. 다만 화면 좌측과 상단 및 우측에 영문 제목과 목록이 선명하게 읽혀 ‘읽을 수 있는 글자 금지’를 위반한다.",
        "hard_violations": [
         "지도 화면에 ‘FACILITY PERIMETER’, ‘BREACH POINTS’ 등 읽을 수 있는 영문 텍스트가 노출됨"
        ],
        "physics": "화면은 금속 콘솔에 정상적으로 매립·고정되어 있으며 지지되지 않은 물체나 신체는 없다. 붉은 점과 지도선은 활성 디스플레이 표면의 발광 정보로 표현되어 화면 위에 물리적으로 떠 있지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.35
   },
   "adjusted": {
    "A": 1.75,
    "B": 1.1
   },
   "violations": {
    "A": [
     "[gemini-pro] 지시문에서 금지한 읽을 수 있는 텍스트가 화면에 명확하게 노출됨 (leaked markers/text 위반).",
     "[gpt] 지도 화면에 ‘FACILITY PERIMETER’, ‘BREACH POINTS’ 등 읽을 수 있는 영문 텍스트가 노출됨"
    ],
    "B": [
     "[gemini-pro] 샷 텍스트에 전혀 명시되지 않은 인물의 신체 일부(손과 팔)가 임의로 프레임에 추가됨 (invented people).",
     "[gpt] 쇼트 텍스트가 보여 주지 않은 인물의 손과 몸 일부가 추가됨"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 1750,
   "B": 1100
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1750,
    "verdict_ko": "지시문에 없는 인물을 추가하지 않아 B보다 낫지만, 철저히 금지된 가독성 높은 텍스트를 화면에 그대로 노출하여 실패한 컷입니다.  ★위반: [gemini-pro] 지시문에서 금지한 읽을 수 있는 텍스트가 화면에 명확하게 노출됨 (leaked markers/text 위반). / [gpt] 지도 화면에 ‘FACILITY PERIMETER’, ‘BREACH POINTS’ 등 읽을 수 있는 영문 텍스트가 노출됨"
   },
   {
    "label": "B",
    "score": 1100,
    "verdict_ko": "붉은 점들이 밀집된 원형의 시각적 묘사는 더 우수하지만, 샷 텍스트에 없는 인물의 신체(손과 팔)를 임의로 추가하여 치명적인 규정 위반을 범했습니다.  ★위반: [gemini-pro] 샷 텍스트에 전혀 명시되지 않은 인물의 신체 일부(손과 팔)가 임의로 프레임에 추가됨 (invented people). / [gpt] 쇼트 텍스트가 보여 주지 않은 인물의 손과 몸 일부가 추가됨"
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/37154490-a4ed-4b37-9092-489c054d8a72/episodes/cc1d3d0b-626e-4f20-bc33-f5f9e2c6ee7e/images/background_chain/L11B02.png",
    "asset_id": "eb7f72c2-429d-4eb2-880a-e3c8b93c34f5",
    "role": "location_plate"
   }
  ],
  "critique_skipped": true,
  "needs_reshoot": true,
  "shot_run_produced": true,
  "shot_run_uid": "06a9bd88-f49c-7a9e-b6d2-2dc002f0b6f7",
  "ref_mode": "플레이트만 (배경 전용)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S1sh12::cine": {
  "applied": true,
  "attempted_at": "2026-09-05T08:54:48.139930+00:00",
  "fingerprint": "b0a902faf563f90c54add9697f672f1a2a708296f9f51ee78972af0cafc6d224",
  "fingerprint_version": 2,
  "provider": "grok",
  "endpoint": "openrouter/chat-completions",
  "model": "x-ai/grok-imagine-image-2.0",
  "pack": "24.202608252115",
  "source_file": "S1sh12_sel.png",
  "source_sha256": "8f173589dad342d11ea8fa146d34d954c385de1d7c28bd8b0474de9be8e4803b",
  "file": "S1sh12_cine.png",
  "staged_sha256": "85692b85da1d5d5e1b3d48f15d03064722104dd818c2a6f55a04db6dbceeb3ee",
  "latency_ms": 11282
 },
 "S2sh3::signage": {
  "fp": "7a5484fa57eb460c",
  "inscriptions": [
   {
    "text_native": "MASS DEFICIT / UNKNOWN RESIDUE",
    "source": "scene_text_quoted",
    "reason_ko": "스캐너 화면에 표시된 텍스트 클로즈업 샷에 직접 명시된 문구이다.",
    "source_quote": "MASS DEFICIT / UNKNOWN RESIDUE"
   }
  ],
  "cues": [],
  "dropped": []
 },
 "S2sh3": {
  "input_fingerprint": "a511a873a91e146d",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — Wyoming, United States, 2419; all people are English-speaking Americans unless stated. TIME OF DAY (lock): twilight.\n\nSHOT TEXT (authoritative, Korean): 스캐너 화면에 뜬 MASS DEFICIT / UNKNOWN RESIDUE 텍스트 클로즈업\n\nLOCATION (lock): Inside the mountain contact post's underground control room, beside the body-scanning station. The scanner screen provides a concentrated electronic glow. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: scanner display (showing “MASS DEFICIT / UNKNOWN RESIDUE”) — The active front face is angled slightly upward toward the camera, with both diagnostic lines fully visible; used as Isolated diagnostic focal surface at the end of the approach.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Neutral low-key ambient light appropriate to the enclosed work area keeps the austere display text crisp.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The metal locker containing the waterproof-pouched phone is enclosed in the transparent isolation case. The scanner display reads “MASS DEFICIT / UNKNOWN RESIDUE,” while the room remains sealed and red-lit.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nWORDS TO RENDER (authoritative — the scene itself calls for these; render each as period-real physical lettering in the native script, exactly as written; add no other readable text anywhere):\n- \"MASS DEFICIT / UNKNOWN RESIDUE\"\n\nThe WORDS TO RENDER above are the only readable writing in this image: render those words exactly as given, in the place and era's own language and script, and nothing else legible. Invent no other wording a viewer could read. No caption, subtitle, watermark, logo or overlay. Surfaces that would carry writing may still be present — stage any wording they would carry out of legibility: a hand across, an oblique angle, shallow focus.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — Wyoming, United States, 2419; all people are English-speaking Americans unless stated. TIME OF DAY (lock): twilight.\n\nSHOT TEXT (authoritative, Korean): 스캐너 화면에 뜬 MASS DEFICIT / UNKNOWN RESIDUE 텍스트 클로즈업\n\nLOCATION (lock): Inside the mountain contact post's underground control room, beside the body-scanning station. The scanner screen provides a concentrated electronic glow. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: scanner display (showing “MASS DEFICIT / UNKNOWN RESIDUE”) — The active front face is angled slightly upward toward the camera, with both diagnostic lines fully visible; used as Isolated diagnostic focal surface at the end of the approach.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Neutral low-key ambient light appropriate to the enclosed work area keeps the austere display text crisp.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The metal locker containing the waterproof-pouched phone is enclosed in the transparent isolation case. The scanner display reads “MASS DEFICIT / UNKNOWN RESIDUE,” while the room remains sealed and red-lit.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nWORDS TO RENDER (authoritative — the scene itself calls for these; render each as period-real physical lettering in the native script, exactly as written; add no other readable text anywhere):\n- \"MASS DEFICIT / UNKNOWN RESIDUE\"\n\nThe WORDS TO RENDER above are the only readable writing in this image: render those words exactly as given, in the place and era's own language and script, and nothing else legible. Invent no other wording a viewer could read. No caption, subtitle, watermark, logo or overlay. Surfaces that would carry writing may still be present — stage any wording they would carry out of legibility: a hand across, an oblique angle, shallow focus.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — Wyoming, United States, 2419; all people are English-speaking Americans unless stated. TIME OF DAY (lock): twilight.\n\nSHOT TEXT (authoritative, Korean): 스캐너 화면에 뜬 MASS DEFICIT / UNKNOWN RESIDUE 텍스트 클로즈업\n\nLOCATION (lock): Inside the mountain contact post's underground control room, beside the body-scanning station. The scanner screen provides a concentrated electronic glow. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: scanner display (showing “MASS DEFICIT / UNKNOWN RESIDUE”) — The active front face is angled slightly upward toward the camera, with both diagnostic lines fully visible; used as Isolated diagnostic focal surface at the end of the approach.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Neutral low-key ambient light appropriate to the enclosed work area keeps the austere display text crisp.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The metal locker containing the waterproof-pouched phone is enclosed in the transparent isolation case. The scanner display reads “MASS DEFICIT / UNKNOWN RESIDUE,” while the room remains sealed and red-lit.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nWORDS TO RENDER (authoritative — the scene itself calls for these; render each as period-real physical lettering in the native script, exactly as written; add no other readable text anywhere):\n- \"MASS DEFICIT / UNKNOWN RESIDUE\"\n\nThe WORDS TO RENDER above are the only readable writing in this image: render those words exactly as given, in the place and era's own language and script, and nothing else legible. Invent no other wording a viewer could read. No caption, subtitle, watermark, logo or overlay. Surfaces that would carry writing may still be present — stage any wording they would carry out of legibility: a hand across, an oblique angle, shallow focus.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "스캐너 화면이 카메라를 향해 약간 위로 기울어져 있음.",
    "built_space": "지하 통제실의 붉은 조명 아래 우측 책상 위에 스캐너 화면이 위치하며, 배경에 투명 격리 케이스가 보임.",
    "entities": "스캐너 화면에 'MASS DEFICIT UNKNOWN RESIDUE' 텍스트(슬래시 누락)와 빨간색, 노란색의 두 진단 선(파형)이 선명하게 표시됨. 투명 케이스 안에는 방수 파우치 형태의 기기가 있으나 금속 로커의 형태는 아님.",
    "hard_violations": [],
    "physics": "스캐너 화면과 투명 케이스가 책상 표면에 안정적으로 놓여 중력을 받고 있음."
   },
   {
    "label": "B",
    "direction": "스캐너 화면이 카메라를 향해 약간 위로 기울어져 있음.",
    "built_space": "통제실 우측 책상에 스캐너 화면이 놓여 있으며, 뒤편으로 투명 격리 케이스와 벽면 상황판, 의자가 보임.",
    "entities": "화면에 'MASS DEFICIT / UNKNOWN RESIDUE'가 표시되었으나 진단 선(파형)이 명확하지 않음. 투명 격리 케이스 내부에는 미니어처 크기로 축소된 스탠드형 체육관 사물함(locker)이 들어 있음.",
    "hard_violations": [
     "[gemini-pro] 물리적으로 불가능한 스케일/스테이징 (탁상용 소형 투명 케이스 안에 대형 스탠드형 철제 사물함이 미니어처 크기로 축소되어 들어감)"
    ],
    "physics": "모든 사물이 탁상 위에 놓여 있으나, 사물함의 크기와 비례가 물리적 현실성에 어긋남."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "텍스트의 슬래시(/)와 투명 케이스 안의 금속 로커가 누락되었으나, 클로즈업 프레이밍과 두 개의 진단 선(파형)을 지시대로 완벽하게 구현했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "텍스트는 정확하나 진단 선이 제대로 나타나지 않았고, 탁상용 투명 케이스 안에 사람 크기의 사물함이 미니어처로 축소되어 들어가는 심각한 스케일 오류가 발생했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "스캐너 화면이 카메라를 향해 약간 위로 기울어져 있음.",
        "built_space": "지하 통제실의 붉은 조명 아래 우측 책상 위에 스캐너 화면이 위치하며, 배경에 투명 격리 케이스가 보임.",
        "entities": "스캐너 화면에 'MASS DEFICIT UNKNOWN RESIDUE' 텍스트(슬래시 누락)와 빨간색, 노란색의 두 진단 선(파형)이 선명하게 표시됨. 투명 케이스 안에는 방수 파우치 형태의 기기가 있으나 금속 로커의 형태는 아님.",
        "hard_violations": [],
        "physics": "스캐너 화면과 투명 케이스가 책상 표면에 안정적으로 놓여 중력을 받고 있음."
       },
       {
        "label": "B",
        "direction": "스캐너 화면이 카메라를 향해 약간 위로 기울어져 있음.",
        "built_space": "통제실 우측 책상에 스캐너 화면이 놓여 있으며, 뒤편으로 투명 격리 케이스와 벽면 상황판, 의자가 보임.",
        "entities": "화면에 'MASS DEFICIT / UNKNOWN RESIDUE'가 표시되었으나 진단 선(파형)이 명확하지 않음. 투명 격리 케이스 내부에는 미니어처 크기로 축소된 스탠드형 체육관 사물함(locker)이 들어 있음.",
        "hard_violations": [
         "물리적으로 불가능한 스케일/스테이징 (탁상용 소형 투명 케이스 안에 대형 스탠드형 철제 사물함이 미니어처 크기로 축소되어 들어감)"
        ],
        "physics": "모든 사물이 탁상 위에 놓여 있으나, 사물함의 크기와 비례가 물리적 현실성에 어긋남."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "텍스트의 슬래시(/)와 투명 케이스 안의 금속 로커가 누락되었으나, 클로즈업 프레이밍과 두 개의 진단 선(파형)을 지시대로 완벽하게 구현했습니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "텍스트는 정확하나 진단 선이 제대로 나타나지 않았고, 탁상용 투명 케이스 안에 사람 크기의 사물함이 미니어처로 축소되어 들어가는 심각한 스케일 오류가 발생했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "스캐너 화면이 카메라를 향해 약간 위로 기울어져 있음.",
        "built_space": "지하 통제실의 붉은 조명 아래 우측 책상 위에 스캐너 화면이 위치하며, 배경에 투명 격리 케이스가 보임.",
        "entities": "스캐너 화면에 'MASS DEFICIT UNKNOWN RESIDUE' 텍스트(슬래시 누락)와 빨간색, 노란색의 두 진단 선(파형)이 선명하게 표시됨. 투명 케이스 안에는 방수 파우치 형태의 기기가 있으나 금속 로커의 형태는 아님.",
        "hard_violations": [],
        "physics": "스캐너 화면과 투명 케이스가 책상 표면에 안정적으로 놓여 중력을 받고 있음."
       },
       {
        "label": "B",
        "direction": "스캐너 화면이 카메라를 향해 약간 위로 기울어져 있음.",
        "built_space": "통제실 우측 책상에 스캐너 화면이 놓여 있으며, 뒤편으로 투명 격리 케이스와 벽면 상황판, 의자가 보임.",
        "entities": "화면에 'MASS DEFICIT / UNKNOWN RESIDUE'가 표시되었으나 진단 선(파형)이 명확하지 않음. 투명 격리 케이스 내부에는 미니어처 크기로 축소된 스탠드형 체육관 사물함(locker)이 들어 있음.",
        "hard_violations": [
         "물리적으로 불가능한 스케일/스테이징 (탁상용 소형 투명 케이스 안에 대형 스탠드형 철제 사물함이 미니어처 크기로 축소되어 들어감)"
        ],
        "physics": "모든 사물이 탁상 위에 놓여 있으나, 사물함의 크기와 비례가 물리적 현실성에 어긋남."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "화면이 다소 넓게 잡히고 미세한 추가 UI 문자가 있으나, 핵심 문구 “MASS DEFICIT / UNKNOWN RESIDUE”를 슬래시까지 정확히 표시하며 투명 격리 케이스 속 금속 로커와 붉은 밀폐실 상태도 충실하다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "더 정확한 클로즈업 구도이지만 핵심 문구에서 필수 슬래시가 빠졌고, 투명 격리 케이스 안에 금속 로커 대신 별도의 소형 장치가 보여 지속 상태를 어긴다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "스캐너 표시창의 활성 전면이 카메라 쪽을 향하면서 약간 위로 기울어져 있고, 두 진단 문구가 모두 정면에서 읽힌다. 시선·무기·이동체는 없다.",
        "built_space": "지하 통제실 작업대 위에 스캐너 표시창 1대가 놓여 있고, 뒤쪽에 투명 격리 케이스 1개와 그 안의 금속 로커 1개가 보인다. 배경에는 고정 콘솔과 의자 1개가 부분적으로 보이며 위치와 방향이 장소 사진의 우측 작업 구역에 부합한다. 불가능한 반사는 보이지 않는다.",
        "entities": "사람이나 신체는 없다. 주 피사체는 실제 재질의 스캐너 화면이며 “MASS DEFICIT / UNKNOWN RESIDUE”가 슬래시를 포함해 정확히 나타난다. 투명 격리 케이스와 내부 금속 로커도 식별된다. 다만 화면 상단과 가장자리에 희미하지만 일부 판독 가능한 추가 UI 문자가 있어 ‘다른 읽을 수 있는 글자 금지’ 조건에는 덜 충실하다.",
        "hard_violations": [],
        "physics": "스캐너는 작업대 표면에 안정적으로 놓여 있고, 금속 로커는 투명 격리 케이스의 바닥에 지지된다. 공중에 뜨거나 손 없이 떠 있는 물체가 없으며 모든 가시 물체에 물리적 지지가 있다."
       },
       {
        "label": "B",
        "direction": "스캐너 표시창의 활성 전면이 카메라를 향해 약간 위로 기울어져 있어 문구와 진단 파형이 정면에서 읽힌다. 시선·무기·이동체는 없다.",
        "built_space": "전경 작업대 위에 스캐너 표시창 1대가 있고, 뒤쪽에 투명 격리 케이스 1개가 보인다. 케이스 안에는 금속 로커가 아니라 화면이 달린 소형 장치가 직접 놓인 것으로 보인다. 배경 콘솔과 조명은 장소 사진의 통제실과 일치하며 불가능한 반사는 없다.",
        "entities": "사람이나 신체는 없다. 스캐너에는 “MASS DEFICIT”와 “UNKNOWN RESIDUE”가 선명하지만 두 문구 사이에 요구된 슬래시가 없다. 투명 격리 케이스는 있으나 그 안의 물체는 방수 파우치 휴대폰을 담은 금속 로커로 식별되지 않는다. 화면 파형은 물리적 디스플레이 그래픽으로 보이며 오버레이처럼 떠 있지는 않는다.",
        "hard_violations": [],
        "physics": "스캐너와 작업 도구는 작업대에 놓여 지지되고, 격리 케이스와 내부 소형 장치도 바닥면에 놓여 있다. 공중에 뜬 물체나 지지 없는 신체는 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "화면이 다소 넓게 잡히고 미세한 추가 UI 문자가 있으나, 핵심 문구 “MASS DEFICIT / UNKNOWN RESIDUE”를 슬래시까지 정확히 표시하며 투명 격리 케이스 속 금속 로커와 붉은 밀폐실 상태도 충실하다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "더 정확한 클로즈업 구도이지만 핵심 문구에서 필수 슬래시가 빠졌고, 투명 격리 케이스 안에 금속 로커 대신 별도의 소형 장치가 보여 지속 상태를 어긴다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "스캐너 표시창의 활성 전면이 카메라 쪽을 향하면서 약간 위로 기울어져 있고, 두 진단 문구가 모두 정면에서 읽힌다. 시선·무기·이동체는 없다.",
        "built_space": "지하 통제실 작업대 위에 스캐너 표시창 1대가 놓여 있고, 뒤쪽에 투명 격리 케이스 1개와 그 안의 금속 로커 1개가 보인다. 배경에는 고정 콘솔과 의자 1개가 부분적으로 보이며 위치와 방향이 장소 사진의 우측 작업 구역에 부합한다. 불가능한 반사는 보이지 않는다.",
        "entities": "사람이나 신체는 없다. 주 피사체는 실제 재질의 스캐너 화면이며 “MASS DEFICIT / UNKNOWN RESIDUE”가 슬래시를 포함해 정확히 나타난다. 투명 격리 케이스와 내부 금속 로커도 식별된다. 다만 화면 상단과 가장자리에 희미하지만 일부 판독 가능한 추가 UI 문자가 있어 ‘다른 읽을 수 있는 글자 금지’ 조건에는 덜 충실하다.",
        "hard_violations": [],
        "physics": "스캐너는 작업대 표면에 안정적으로 놓여 있고, 금속 로커는 투명 격리 케이스의 바닥에 지지된다. 공중에 뜨거나 손 없이 떠 있는 물체가 없으며 모든 가시 물체에 물리적 지지가 있다."
       },
       {
        "label": "A",
        "direction": "스캐너 표시창의 활성 전면이 카메라를 향해 약간 위로 기울어져 있어 문구와 진단 파형이 정면에서 읽힌다. 시선·무기·이동체는 없다.",
        "built_space": "전경 작업대 위에 스캐너 표시창 1대가 있고, 뒤쪽에 투명 격리 케이스 1개가 보인다. 케이스 안에는 금속 로커가 아니라 화면이 달린 소형 장치가 직접 놓인 것으로 보인다. 배경 콘솔과 조명은 장소 사진의 통제실과 일치하며 불가능한 반사는 없다.",
        "entities": "사람이나 신체는 없다. 스캐너에는 “MASS DEFICIT”와 “UNKNOWN RESIDUE”가 선명하지만 두 문구 사이에 요구된 슬래시가 없다. 투명 격리 케이스는 있으나 그 안의 물체는 방수 파우치 휴대폰을 담은 금속 로커로 식별되지 않는다. 화면 파형은 물리적 디스플레이 그래픽으로 보이며 오버레이처럼 떠 있지는 않는다.",
        "hard_violations": [],
        "physics": "스캐너와 작업 도구는 작업대에 놓여 지지되고, 격리 케이스와 내부 소형 장치도 바닥면에 놓여 있다. 공중에 뜬 물체나 지지 없는 신체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt"
   ],
   "normalized": {
    "A": 1.75,
    "B": 1.375
   },
   "adjusted": {
    "A": 1.75,
    "B": 1.125
   },
   "violations": {
    "B": [
     "[gemini-pro] 물리적으로 불가능한 스케일/스테이징 (탁상용 소형 투명 케이스 안에 대형 스탠드형 철제 사물함이 미니어처 크기로 축소되어 들어감)"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1750,
   "B": 1125
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1750,
    "verdict_ko": "텍스트의 슬래시(/)와 투명 케이스 안의 금속 로커가 누락되었으나, 클로즈업 프레이밍과 두 개의 진단 선(파형)을 지시대로 완벽하게 구현했습니다."
   },
   {
    "label": "B",
    "score": 1125,
    "verdict_ko": "텍스트는 정확하나 진단 선이 제대로 나타나지 않았고, 탁상용 투명 케이스 안에 사람 크기의 사물함이 미니어처로 축소되어 들어가는 심각한 스케일 오류가 발생했습니다.  ★위반: [gemini-pro] 물리적으로 불가능한 스케일/스테이징 (탁상용 소형 투명 케이스 안에 대형 스탠드형 철제 사물함이 미니어처 크기로 축소되어 들어감)"
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/37154490-a4ed-4b37-9092-489c054d8a72/episodes/cc1d3d0b-626e-4f20-bc33-f5f9e2c6ee7e/images/background_chain/L11B02.png",
    "asset_id": "eb7f72c2-429d-4eb2-880a-e3c8b93c34f5",
    "role": "location_plate"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": true,
  "shot_run_uid": "06a9bd8e-38d7-7b74-9683-7b0021c80d14",
  "ref_mode": "플레이트만 (배경 전용)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S2sh3::cine": {
  "applied": true,
  "attempted_at": "2026-09-05T08:55:46.668827+00:00",
  "fingerprint": "07ed7378e576b85cd03f70ef50d9fda30c1b7c47050d807780a2b500d7cfced6",
  "fingerprint_version": 2,
  "provider": "grok",
  "endpoint": "openrouter/chat-completions",
  "model": "x-ai/grok-imagine-image-2.0",
  "pack": "24.202608252115",
  "source_file": "S2sh3_sel.png",
  "source_sha256": "5799a49b1a814f651806b24d937a26833fb1fba7be6890468afc495833e0bf86",
  "file": "S2sh3_cine.png",
  "staged_sha256": "6ae71bcc76d8943526c2953b58c2da3116db0f5d60f4d36a98375abe06c2dcb6",
  "latency_ms": 16398
 },
 "S2sh8::signage": {
  "fp": "56aa26c80b6debf8",
  "inscriptions": [
   {
    "text_native": "BATTERY 2%",
    "source": "scene_text_quoted",
    "reason_ko": "외부 모니터 화면에 표시된 배터리 잔량 텍스트.",
    "source_quote": "BATTERY 2%"
   }
  ],
  "cues": [],
  "dropped": []
 },
 "S2sh8": {
  "input_fingerprint": "c39d30bce0ffc71f",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — Wyoming, United States, 2419; all people are English-speaking Americans unless stated. TIME OF DAY (lock): twilight.\n\nSHOT TEXT (authoritative, Korean): 외부 모니터에 휴대폰과 동일한 BATTERY 2% 화면이 뜬 상태\n\nLOCATION (lock): Inside the mountain contact post's underground control room, at the wired isolation box and its external monitor. The active monitor lights the immediate workstation. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: external monitor (displaying the duplicated phone screen with “BATTERY 2%”) — The active display face is seen slightly off-axis, with the duplicated phone interface and battery warning clearly visible; used as Primary focal surface held briefly for text legibility; transparent isolation box (holding the storage box and connected phone) — Its transparent boundary permits a limited view of the enclosed items behind the monitor-edge context; used as Provides peripheral evidence that the phone remains physically isolated; wired connection (connected between the isolation box terminal and the phone); used as Creates a visible physical link supporting the duplicated display.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Neutral, subdued ambient illumination preserves the monitor's austere interface and the urgency of the low-battery reading.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Inside the transparent isolation case, the wired terminal is plugged into the phone and wireless communication is blocked. The external monitor mirrors the phone’s lit “BATTERY 2%” screen.\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 윌마 디어링 (미국인 여성, 20대 후반, 자연스러운 성인 여성 얼굴, 구체적으로 명시되지 않은 자연색 머리카락). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nWORDS TO RENDER (authoritative — the scene itself calls for these; render each as period-real physical lettering in the native script, exactly as written; add no other readable text anywhere):\n- \"BATTERY 2%\"\n\nThe WORDS TO RENDER above are the only readable writing in this image: render those words exactly as given, in the place and era's own language and script, and nothing else legible. Invent no other wording a viewer could read. No caption, subtitle, watermark, logo or overlay. Surfaces that would carry writing may still be present — stage any wording they would carry out of legibility: a hand across, an oblique angle, shallow focus.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — Wyoming, United States, 2419; all people are English-speaking Americans unless stated. TIME OF DAY (lock): twilight.\n\nSHOT TEXT (authoritative, Korean): 외부 모니터에 휴대폰과 동일한 BATTERY 2% 화면이 뜬 상태\n\nLOCATION (lock): Inside the mountain contact post's underground control room, at the wired isolation box and its external monitor. The active monitor lights the immediate workstation. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: external monitor (displaying the duplicated phone screen with “BATTERY 2%”) — The active display face is seen slightly off-axis, with the duplicated phone interface and battery warning clearly visible; used as Primary focal surface held briefly for text legibility; transparent isolation box (holding the storage box and connected phone) — Its transparent boundary permits a limited view of the enclosed items behind the monitor-edge context; used as Provides peripheral evidence that the phone remains physically isolated; wired connection (connected between the isolation box terminal and the phone); used as Creates a visible physical link supporting the duplicated display.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Neutral, subdued ambient illumination preserves the monitor's austere interface and the urgency of the low-battery reading.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Inside the transparent isolation case, the wired terminal is plugged into the phone and wireless communication is blocked. The external monitor mirrors the phone’s lit “BATTERY 2%” screen.\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 윌마 디어링 (미국인 여성, 20대 후반, 자연스러운 성인 여성 얼굴, 구체적으로 명시되지 않은 자연색 머리카락). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nWORDS TO RENDER (authoritative — the scene itself calls for these; render each as period-real physical lettering in the native script, exactly as written; add no other readable text anywhere):\n- \"BATTERY 2%\"\n\nThe WORDS TO RENDER above are the only readable writing in this image: render those words exactly as given, in the place and era's own language and script, and nothing else legible. Invent no other wording a viewer could read. No caption, subtitle, watermark, logo or overlay. Surfaces that would carry writing may still be present — stage any wording they would carry out of legibility: a hand across, an oblique angle, shallow focus.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — Wyoming, United States, 2419; all people are English-speaking Americans unless stated. TIME OF DAY (lock): twilight.\n\nSHOT TEXT (authoritative, Korean): 외부 모니터에 휴대폰과 동일한 BATTERY 2% 화면이 뜬 상태\n\nLOCATION (lock): Inside the mountain contact post's underground control room, at the wired isolation box and its external monitor. The active monitor lights the immediate workstation. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: external monitor (displaying the duplicated phone screen with “BATTERY 2%”) — The active display face is seen slightly off-axis, with the duplicated phone interface and battery warning clearly visible; used as Primary focal surface held briefly for text legibility; transparent isolation box (holding the storage box and connected phone) — Its transparent boundary permits a limited view of the enclosed items behind the monitor-edge context; used as Provides peripheral evidence that the phone remains physically isolated; wired connection (connected between the isolation box terminal and the phone); used as Creates a visible physical link supporting the duplicated display.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Neutral, subdued ambient illumination preserves the monitor's austere interface and the urgency of the low-battery reading.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): Inside the transparent isolation case, the wired terminal is plugged into the phone and wireless communication is blocked. The external monitor mirrors the phone’s lit “BATTERY 2%” screen.\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 윌마 디어링 (미국인 여성, 20대 후반, 자연스러운 성인 여성 얼굴, 구체적으로 명시되지 않은 자연색 머리카락). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nWORDS TO RENDER (authoritative — the scene itself calls for these; render each as period-real physical lettering in the native script, exactly as written; add no other readable text anywhere):\n- \"BATTERY 2%\"\n\nThe WORDS TO RENDER above are the only readable writing in this image: render those words exactly as given, in the place and era's own language and script, and nothing else legible. Invent no other wording a viewer could read. No caption, subtitle, watermark, logo or overlay. Surfaces that would carry writing may still be present — stage any wording they would carry out of legibility: a hand across, an oblique angle, shallow focus.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "외부 모니터 화면이 카메라를 향해 비스듬한 각도로 놓여 있음.",
    "built_space": "지하 통제실의 작업대 위. 모니터와 투명 격리 상자가 배치되어 있음.",
    "entities": "외부 모니터, 'BATTERY 2%' 텍스트, 투명 격리 상자, 휴대폰이 존재함. 그러나 격리 상자 내부에 있어야 할 '보관 상자(storage box)'가 완전히 누락됨.",
    "hard_violations": [
     "[gemini-pro] 프롬프트에 명시된 보관 상자(storage box) 누락",
     "[gemini-pro] 격리 단자와 휴대폰을 잇는 유선 연결 누락 (케이블이 상자 외부에만 꽂혀 있고 내부 휴대폰은 단절됨)"
    ],
    "physics": "모니터에서 나온 케이블이 투명 상자 외부 표면에만 연결되어 있고, 정작 내부의 휴대폰에는 아무런 물리적 선이 연결되어 있지 않아 프롬프트의 지시사항과 물리적 개연성을 모두 훼손함. 휴대폰은 바닥에 그냥 놓여 있음."
   },
   {
    "label": "B",
    "direction": "외부 모니터가 카메라를 향해 비스듬히 놓여 있으며 화면의 텍스트가 명확하게 보임.",
    "built_space": "통제실 작업대 위. 모니터와 투명 격리 상자가 공간에 알맞게 배치됨.",
    "entities": "외부 모니터, 'BATTERY 2%' 텍스트, 투명 격리 상자, 내부의 보관 상자, 그 위에 놓인 휴대폰, 그리고 유선 케이블이 모두 정확히 구현됨.",
    "hard_violations": [],
    "physics": "투명 상자 내부의 보관 상자 위에 휴대폰이 안정적으로 거치되어 있으며, 케이블이 단자와 휴대폰에 정상적으로 연결되어 물리적 지지 및 연결 상태가 자연스러움."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "프롬프트가 요구한 모든 사물(보관 상자, 연결된 휴대폰, 투명 격리 상자)과 정확한 텍스트를 물리적 오류 없이 훌륭하게 구현했습니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "격리 상자 내부의 보관 상자가 누락되었고, 휴대폰에 케이블이 연결되지 않아 핵심적인 프롬프트 지시를 위반했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "외부 모니터 화면이 카메라를 향해 비스듬한 각도로 놓여 있음.",
        "built_space": "지하 통제실의 작업대 위. 모니터와 투명 격리 상자가 배치되어 있음.",
        "entities": "외부 모니터, 'BATTERY 2%' 텍스트, 투명 격리 상자, 휴대폰이 존재함. 그러나 격리 상자 내부에 있어야 할 '보관 상자(storage box)'가 완전히 누락됨.",
        "hard_violations": [
         "프롬프트에 명시된 보관 상자(storage box) 누락",
         "격리 단자와 휴대폰을 잇는 유선 연결 누락 (케이블이 상자 외부에만 꽂혀 있고 내부 휴대폰은 단절됨)"
        ],
        "physics": "모니터에서 나온 케이블이 투명 상자 외부 표면에만 연결되어 있고, 정작 내부의 휴대폰에는 아무런 물리적 선이 연결되어 있지 않아 프롬프트의 지시사항과 물리적 개연성을 모두 훼손함. 휴대폰은 바닥에 그냥 놓여 있음."
       },
       {
        "label": "B",
        "direction": "외부 모니터가 카메라를 향해 비스듬히 놓여 있으며 화면의 텍스트가 명확하게 보임.",
        "built_space": "통제실 작업대 위. 모니터와 투명 격리 상자가 공간에 알맞게 배치됨.",
        "entities": "외부 모니터, 'BATTERY 2%' 텍스트, 투명 격리 상자, 내부의 보관 상자, 그 위에 놓인 휴대폰, 그리고 유선 케이블이 모두 정확히 구현됨.",
        "hard_violations": [],
        "physics": "투명 상자 내부의 보관 상자 위에 휴대폰이 안정적으로 거치되어 있으며, 케이블이 단자와 휴대폰에 정상적으로 연결되어 물리적 지지 및 연결 상태가 자연스러움."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "프롬프트가 요구한 모든 사물(보관 상자, 연결된 휴대폰, 투명 격리 상자)과 정확한 텍스트를 물리적 오류 없이 훌륭하게 구현했습니다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "격리 상자 내부의 보관 상자가 누락되었고, 휴대폰에 케이블이 연결되지 않아 핵심적인 프롬프트 지시를 위반했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "외부 모니터 화면이 카메라를 향해 비스듬한 각도로 놓여 있음.",
        "built_space": "지하 통제실의 작업대 위. 모니터와 투명 격리 상자가 배치되어 있음.",
        "entities": "외부 모니터, 'BATTERY 2%' 텍스트, 투명 격리 상자, 휴대폰이 존재함. 그러나 격리 상자 내부에 있어야 할 '보관 상자(storage box)'가 완전히 누락됨.",
        "hard_violations": [
         "프롬프트에 명시된 보관 상자(storage box) 누락",
         "격리 단자와 휴대폰을 잇는 유선 연결 누락 (케이블이 상자 외부에만 꽂혀 있고 내부 휴대폰은 단절됨)"
        ],
        "physics": "모니터에서 나온 케이블이 투명 상자 외부 표면에만 연결되어 있고, 정작 내부의 휴대폰에는 아무런 물리적 선이 연결되어 있지 않아 프롬프트의 지시사항과 물리적 개연성을 모두 훼손함. 휴대폰은 바닥에 그냥 놓여 있음."
       },
       {
        "label": "B",
        "direction": "외부 모니터가 카메라를 향해 비스듬히 놓여 있으며 화면의 텍스트가 명확하게 보임.",
        "built_space": "통제실 작업대 위. 모니터와 투명 격리 상자가 공간에 알맞게 배치됨.",
        "entities": "외부 모니터, 'BATTERY 2%' 텍스트, 투명 격리 상자, 내부의 보관 상자, 그 위에 놓인 휴대폰, 그리고 유선 케이블이 모두 정확히 구현됨.",
        "hard_violations": [],
        "physics": "투명 상자 내부의 보관 상자 위에 휴대폰이 안정적으로 거치되어 있으며, 케이블이 단자와 휴대폰에 정상적으로 연결되어 물리적 지지 및 연결 상태가 자연스러움."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "정확한 클로즈업에서 외부 모니터의 “BATTERY 2%”, 투명 격리함 속 보관함·휴대폰, 단말과 휴대폰을 잇는 배선이 모두 명확해 핵심 장면을 충실히 구현했다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "클로즈업과 지정 문구는 갖췄지만 모니터가 휴대폰 화면 자체가 아니라 휴대폰을 찍은 영상처럼 보이고, 격리함 내부의 보관함과 휴대폰 연결 배선이 빠져 핵심 물리 관계가 틀렸다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "사람의 시선이나 무기, 이동체는 없다. 외부 모니터의 표시 면은 카메라를 향해 약간 비스듬히 놓여 있으며 “BATTERY 2%”가 정면에 가깝게 읽힌다. 격리함 속 휴대폰은 화면이 위쪽을 향한다.",
        "built_space": "지하 제어실 작업대의 근접 구도다. 전경 왼쪽에 외부 모니터 1대, 오른쪽에 투명 격리함 1개가 있고, 그 안에 견고한 보관함 1개와 그 위의 휴대폰 1대가 보인다. 격리함 벽의 단말과 내부 휴대폰 쪽을 잇는 케이블이 보이며, 모니터 가장자리와 격리함이 함께 들어오는 배치도 요구된 근접 프레이밍에 맞는다. 반사와 투시도 카메라 위치상 가능하다.",
        "entities": "외부 모니터는 실제 금속·플라스틱 장비로 보이고 화면에 지정된 유일한 가독 문구 “BATTERY 2%”가 정확히 표시된다. 투명 격리함, 내부 보관함, 휴대폰, 유선 단말과 연결선이 모두 식별된다. 사람은 없으며, 인물을 요구하지 않는 쇼트 텍스트와 맞는다.",
        "hard_violations": [],
        "physics": "모니터는 자체 받침대로 작업대 위에 지지되고, 격리함도 작업대에 놓여 있다. 보관함은 격리함 바닥에 놓이고 휴대폰은 보관함 위에 지지된다. 케이블은 단말과 휴대폰 쪽에 연결되고 일부가 바닥면에 닿아 있어 떠 있는 물체가 없다."
       },
       {
        "label": "B",
        "direction": "사람의 시선이나 무기, 이동체는 없다. 외부 모니터 면은 카메라 쪽으로 약간 비스듬히 향해 문구가 읽힌다. 격리함 속 휴대폰은 화면을 위로 향하지만 외부 모니터에는 그 화면 자체보다 휴대폰 기기 전체를 위에서 촬영한 듯한 영상이 표시된다.",
        "built_space": "전경 왼쪽에 외부 모니터 1대, 오른쪽에 투명 격리함 1개와 벽면 단말 1개가 있다. 격리함 내부에는 휴대폰 1대만 보이고 요구된 보관함은 없다. 외부 케이블은 단말에서 나와 작업대 위에 고리 모양으로 놓였지만, 격리함 안에서 휴대폰까지 이어지는 내부 배선은 보이지 않는다. 반사와 투시는 대체로 가능한 범위다.",
        "entities": "외부 모니터, 투명 격리함, 휴대폰과 외부 케이블은 식별되며 “BATTERY 2%”도 읽힌다. 그러나 모니터 화면은 휴대폰 인터페이스의 직접 복제가 아니라 휴대폰 외형을 담은 영상처럼 보이고, 격리함 속 보관함이 누락되었으며 휴대폰도 유선 연결 상태로 보이지 않는다. 화면 안에는 지정 문구 외의 미세한 인터페이스 글자도 있으나 뚜렷하게 판독되지는 않는다. 사람은 없다.",
        "hard_violations": [],
        "physics": "모니터와 격리함은 작업대가 지지하고 휴대폰은 격리함 바닥에 놓여 있다. 외부 케이블 역시 작업대에 받쳐져 있어 부유 물체는 없다. 다만 케이블이 단말까지만 연결되고 내부 휴대폰에는 닿지 않아, 요구된 유선 연결 기능을 물리적으로 수행하지 못한다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "정확한 클로즈업에서 외부 모니터의 “BATTERY 2%”, 투명 격리함 속 보관함·휴대폰, 단말과 휴대폰을 잇는 배선이 모두 명확해 핵심 장면을 충실히 구현했다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "클로즈업과 지정 문구는 갖췄지만 모니터가 휴대폰 화면 자체가 아니라 휴대폰을 찍은 영상처럼 보이고, 격리함 내부의 보관함과 휴대폰 연결 배선이 빠져 핵심 물리 관계가 틀렸다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "사람의 시선이나 무기, 이동체는 없다. 외부 모니터의 표시 면은 카메라를 향해 약간 비스듬히 놓여 있으며 “BATTERY 2%”가 정면에 가깝게 읽힌다. 격리함 속 휴대폰은 화면이 위쪽을 향한다.",
        "built_space": "지하 제어실 작업대의 근접 구도다. 전경 왼쪽에 외부 모니터 1대, 오른쪽에 투명 격리함 1개가 있고, 그 안에 견고한 보관함 1개와 그 위의 휴대폰 1대가 보인다. 격리함 벽의 단말과 내부 휴대폰 쪽을 잇는 케이블이 보이며, 모니터 가장자리와 격리함이 함께 들어오는 배치도 요구된 근접 프레이밍에 맞는다. 반사와 투시도 카메라 위치상 가능하다.",
        "entities": "외부 모니터는 실제 금속·플라스틱 장비로 보이고 화면에 지정된 유일한 가독 문구 “BATTERY 2%”가 정확히 표시된다. 투명 격리함, 내부 보관함, 휴대폰, 유선 단말과 연결선이 모두 식별된다. 사람은 없으며, 인물을 요구하지 않는 쇼트 텍스트와 맞는다.",
        "hard_violations": [],
        "physics": "모니터는 자체 받침대로 작업대 위에 지지되고, 격리함도 작업대에 놓여 있다. 보관함은 격리함 바닥에 놓이고 휴대폰은 보관함 위에 지지된다. 케이블은 단말과 휴대폰 쪽에 연결되고 일부가 바닥면에 닿아 있어 떠 있는 물체가 없다."
       },
       {
        "label": "A",
        "direction": "사람의 시선이나 무기, 이동체는 없다. 외부 모니터 면은 카메라 쪽으로 약간 비스듬히 향해 문구가 읽힌다. 격리함 속 휴대폰은 화면을 위로 향하지만 외부 모니터에는 그 화면 자체보다 휴대폰 기기 전체를 위에서 촬영한 듯한 영상이 표시된다.",
        "built_space": "전경 왼쪽에 외부 모니터 1대, 오른쪽에 투명 격리함 1개와 벽면 단말 1개가 있다. 격리함 내부에는 휴대폰 1대만 보이고 요구된 보관함은 없다. 외부 케이블은 단말에서 나와 작업대 위에 고리 모양으로 놓였지만, 격리함 안에서 휴대폰까지 이어지는 내부 배선은 보이지 않는다. 반사와 투시는 대체로 가능한 범위다.",
        "entities": "외부 모니터, 투명 격리함, 휴대폰과 외부 케이블은 식별되며 “BATTERY 2%”도 읽힌다. 그러나 모니터 화면은 휴대폰 인터페이스의 직접 복제가 아니라 휴대폰 외형을 담은 영상처럼 보이고, 격리함 속 보관함이 누락되었으며 휴대폰도 유선 연결 상태로 보이지 않는다. 화면 안에는 지정 문구 외의 미세한 인터페이스 글자도 있으나 뚜렷하게 판독되지는 않는다. 사람은 없다.",
        "hard_violations": [],
        "physics": "모니터와 격리함은 작업대가 지지하고 휴대폰은 격리함 바닥에 놓여 있다. 외부 케이블 역시 작업대에 받쳐져 있어 부유 물체는 없다. 다만 케이블이 단말까지만 연결되고 내부 휴대폰에는 닿지 않아, 요구된 유선 연결 기능을 물리적으로 수행하지 못한다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt"
   ],
   "normalized": {
    "A": 0.889,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.639,
    "B": 2.0
   },
   "violations": {
    "A": [
     "[gemini-pro] 프롬프트에 명시된 보관 상자(storage box) 누락",
     "[gemini-pro] 격리 단자와 휴대폰을 잇는 유선 연결 누락 (케이블이 상자 외부에만 꽂혀 있고 내부 휴대폰은 단절됨)"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 2000,
   "A": 639
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "프롬프트가 요구한 모든 사물(보관 상자, 연결된 휴대폰, 투명 격리 상자)과 정확한 텍스트를 물리적 오류 없이 훌륭하게 구현했습니다."
   },
   {
    "label": "A",
    "score": 639,
    "verdict_ko": "격리 상자 내부의 보관 상자가 누락되었고, 휴대폰에 케이블이 연결되지 않아 핵심적인 프롬프트 지시를 위반했습니다.  ★위반: [gemini-pro] 프롬프트에 명시된 보관 상자(storage box) 누락 / [gemini-pro] 격리 단자와 휴대폰을 잇는 유선 연결 누락 (케이블이 상자 외부에만 꽂혀 있고 내부 휴대폰은 단절됨)"
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/37154490-a4ed-4b37-9092-489c054d8a72/episodes/cc1d3d0b-626e-4f20-bc33-f5f9e2c6ee7e/images/background_chain/L11B02.png",
    "asset_id": "eb7f72c2-429d-4eb2-880a-e3c8b93c34f5",
    "role": "location_plate"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": true,
  "shot_run_uid": "06a9bd92-333d-71ed-b1c8-2e9ea88126ac",
  "ref_mode": "플레이트만 (배경 전용)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S2sh8::cine": {
  "applied": true,
  "attempted_at": "2026-09-05T08:56:51.676506+00:00",
  "fingerprint": "32a1eb3c9c2c124fc84eefc94a5a319064dd6f6f216913ea7239eeb3b85e0ff1",
  "fingerprint_version": 2,
  "provider": "grok",
  "endpoint": "openrouter/chat-completions",
  "model": "x-ai/grok-imagine-image-2.0",
  "pack": "24.202608252115",
  "source_file": "S2sh8_sel.png",
  "source_sha256": "0e23c9eae3e0d2cc951366c45aca449a9fcc7ad89587143cab74c9f598230b2b",
  "file": "S2sh8_cine.png",
  "staged_sha256": "df2347bbc6a06eb48a66df309cf1c6b39c291bba7a12da3b645aa6f647f179b9",
  "latency_ms": 13288
 },
 "S2sh11::signage": {
  "fp": "9da8445e5134d83e",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S2sh11": {
  "input_fingerprint": "386cd22b081a2d47",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — Wyoming, United States, 2419; all people are English-speaking Americans unless stated. TIME OF DAY (lock): twilight.\n\nSHOT TEXT (authoritative, Korean): 미간을 찌푸린 채 모니터를 바라보는 윌마의 얼굴\n\nLOCATION (lock): Inside the mountain contact post's underground control room, directly in front of the external monitoring station. Screen light falls across the observer's face. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: external monitor (active with the duplicated phone display) — Only the monitor's near side and a sliver of its display-facing edge are visible; the content face is angled away from the camera toward Wilma; used as Soft foreground edge that anchors Wilma's off-screen eyeline.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Neutral low-key ambient light shapes Wilma's tense features without introducing an unmotivated color source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established — the place, its fixed features and wear, each person's clothing and state — persists. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The phone remains wired inside the transparent isolation case, with its active display mirrored to the external monitor at 2% battery. The rock door remains closed under red interior lighting.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 윌마 디어링 (미국인 여성, 20대 후반, 자연스러운 성인 여성 얼굴, 구체적으로 명시되지 않은 자연색 머리카락) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — Wyoming, United States, 2419; all people are English-speaking Americans unless stated. TIME OF DAY (lock): twilight.\n\nSHOT TEXT (authoritative, Korean): 미간을 찌푸린 채 모니터를 바라보는 윌마의 얼굴\n\nLOCATION (lock): Inside the mountain contact post's underground control room, directly in front of the external monitoring station. Screen light falls across the observer's face. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: external monitor (active with the duplicated phone display) — Only the monitor's near side and a sliver of its display-facing edge are visible; the content face is angled away from the camera toward Wilma; used as Soft foreground edge that anchors Wilma's off-screen eyeline.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Neutral low-key ambient light shapes Wilma's tense features without introducing an unmotivated color source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established — the place, its fixed features and wear, each person's clothing and state — persists. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The phone remains wired inside the transparent isolation case, with its active display mirrored to the external monitor at 2% battery. The rock door remains closed under red interior lighting.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 윌마 디어링 (미국인 여성, 20대 후반, 자연스러운 성인 여성 얼굴, 구체적으로 명시되지 않은 자연색 머리카락) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — Wyoming, United States, 2419; all people are English-speaking Americans unless stated. TIME OF DAY (lock): twilight.\n\nSHOT TEXT (authoritative, Korean): 미간을 찌푸린 채 모니터를 바라보는 윌마의 얼굴\n\nLOCATION (lock): Inside the mountain contact post's underground control room, directly in front of the external monitoring station. Screen light falls across the observer's face. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: external monitor (active with the duplicated phone display) — Only the monitor's near side and a sliver of its display-facing edge are visible; the content face is angled away from the camera toward Wilma; used as Soft foreground edge that anchors Wilma's off-screen eyeline.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Neutral low-key ambient light shapes Wilma's tense features without introducing an unmotivated color source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established — the place, its fixed features and wear, each person's clothing and state — persists. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The phone remains wired inside the transparent isolation case, with its active display mirrored to the external monitor at 2% battery. The rock door remains closed under red interior lighting.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 윌마 디어링 (미국인 여성, 20대 후반, 자연스러운 성인 여성 얼굴, 구체적으로 명시되지 않은 자연색 머리카락) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "윌마의 시선은 화면 우측 전경에 부드러운 경계를 형성하고 있는 모니터를 향해 뚜렷하게 고정되어 있다.",
    "built_space": "지하 통제실 내부. 윌마가 중앙에 있고 우측 전경에 모니터가 배치되어 있다. 이전 샷에서 모니터 옆에 있던 아크릴 격리 상자가 윌마의 좌측 뒤편 배경으로 이동해 있어 공간 연속성에 오류가 있으나, 우측 배경에는 프롬프트가 지시한 붉은 조명이 묘사되어 있다.",
    "entities": "윌마의 얼굴과 묶은 머리, 녹색 비행복은 캐릭터 레퍼런스와 완벽히 일치한다. 우측 전경에 두꺼운 베젤을 가진 모니터 측면이 보이며, 좌측 배경에는 투명 아크릴 격리 상자와 그 안의 황갈색 케이스가 존재한다.",
    "hard_violations": [],
    "physics": "윌마의 자세는 안정적이며 무게중심이 자연스럽다. 전경의 모니터와 배경의 상자들 역시 표면 위에 정상적으로 지지되어 있다."
   },
   {
    "label": "B",
    "direction": "윌마의 시선은 화면 좌측 전경에 앵글이 틀어진 채 배치된 모니터 쪽을 향하고 있다.",
    "built_space": "좌측 전경에 모니터가 있고 그 바로 뒤편 책상 위에 황갈색 케이스가 놓여 있어 이전 샷의 기기 배치를 잘 따르고 있다. 하지만 프롬프트에 명시된 붉은 조명이 있는 문은 배경에서 보이지 않는다.",
    "entities": "윌마의 이목구비는 레퍼런스와 일치하지만, 의상 지침을 어기고 레퍼런스에 없는 두꺼운 검은색 가슴 하네스와 성조기 패치를 착용하고 있다. 또한 폰을 덮고 있어야 할 투명 아크릴 격리 상자가 묘사되지 않아 폰이 케이스 위에 그대로 노출되어 있다.",
    "hard_violations": [
     "[gemini-pro] invented objects (캐릭터 레퍼런스에 없는 검은색 가슴 하네스 및 패치 추가)"
    ],
    "physics": "인물의 어깨와 자세, 그리고 책상 위에 놓인 장비들은 모두 물리 법칙에 맞게 자연스럽게 지지되어 있다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "레퍼런스와 일치하는 복장을 입고 명시된 투명 격리 상자를 묘사하여 프롬프트의 요구사항을 충실히 따랐으나, 격리 상자의 공간적 위치가 이전 샷과 어긋나게 배경으로 이동한 점이 아쉽습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "모니터와 황갈색 케이스의 공간 배치는 이전 샷과 잘 맞으나, 필수적인 투명 아크릴 격리 상자가 완전히 누락되었고 레퍼런스에 없는 가슴 하네스를 착용하여 주요 지침을 위반했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "윌마의 시선은 화면 우측 전경에 부드러운 경계를 형성하고 있는 모니터를 향해 뚜렷하게 고정되어 있다.",
        "built_space": "지하 통제실 내부. 윌마가 중앙에 있고 우측 전경에 모니터가 배치되어 있다. 이전 샷에서 모니터 옆에 있던 아크릴 격리 상자가 윌마의 좌측 뒤편 배경으로 이동해 있어 공간 연속성에 오류가 있으나, 우측 배경에는 프롬프트가 지시한 붉은 조명이 묘사되어 있다.",
        "entities": "윌마의 얼굴과 묶은 머리, 녹색 비행복은 캐릭터 레퍼런스와 완벽히 일치한다. 우측 전경에 두꺼운 베젤을 가진 모니터 측면이 보이며, 좌측 배경에는 투명 아크릴 격리 상자와 그 안의 황갈색 케이스가 존재한다.",
        "hard_violations": [],
        "physics": "윌마의 자세는 안정적이며 무게중심이 자연스럽다. 전경의 모니터와 배경의 상자들 역시 표면 위에 정상적으로 지지되어 있다."
       },
       {
        "label": "B",
        "direction": "윌마의 시선은 화면 좌측 전경에 앵글이 틀어진 채 배치된 모니터 쪽을 향하고 있다.",
        "built_space": "좌측 전경에 모니터가 있고 그 바로 뒤편 책상 위에 황갈색 케이스가 놓여 있어 이전 샷의 기기 배치를 잘 따르고 있다. 하지만 프롬프트에 명시된 붉은 조명이 있는 문은 배경에서 보이지 않는다.",
        "entities": "윌마의 이목구비는 레퍼런스와 일치하지만, 의상 지침을 어기고 레퍼런스에 없는 두꺼운 검은색 가슴 하네스와 성조기 패치를 착용하고 있다. 또한 폰을 덮고 있어야 할 투명 아크릴 격리 상자가 묘사되지 않아 폰이 케이스 위에 그대로 노출되어 있다.",
        "hard_violations": [
         "invented objects (캐릭터 레퍼런스에 없는 검은색 가슴 하네스 및 패치 추가)"
        ],
        "physics": "인물의 어깨와 자세, 그리고 책상 위에 놓인 장비들은 모두 물리 법칙에 맞게 자연스럽게 지지되어 있다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "레퍼런스와 일치하는 복장을 입고 명시된 투명 격리 상자를 묘사하여 프롬프트의 요구사항을 충실히 따랐으나, 격리 상자의 공간적 위치가 이전 샷과 어긋나게 배경으로 이동한 점이 아쉽습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "모니터와 황갈색 케이스의 공간 배치는 이전 샷과 잘 맞으나, 필수적인 투명 아크릴 격리 상자가 완전히 누락되었고 레퍼런스에 없는 가슴 하네스를 착용하여 주요 지침을 위반했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "윌마의 시선은 화면 우측 전경에 부드러운 경계를 형성하고 있는 모니터를 향해 뚜렷하게 고정되어 있다.",
        "built_space": "지하 통제실 내부. 윌마가 중앙에 있고 우측 전경에 모니터가 배치되어 있다. 이전 샷에서 모니터 옆에 있던 아크릴 격리 상자가 윌마의 좌측 뒤편 배경으로 이동해 있어 공간 연속성에 오류가 있으나, 우측 배경에는 프롬프트가 지시한 붉은 조명이 묘사되어 있다.",
        "entities": "윌마의 얼굴과 묶은 머리, 녹색 비행복은 캐릭터 레퍼런스와 완벽히 일치한다. 우측 전경에 두꺼운 베젤을 가진 모니터 측면이 보이며, 좌측 배경에는 투명 아크릴 격리 상자와 그 안의 황갈색 케이스가 존재한다.",
        "hard_violations": [],
        "physics": "윌마의 자세는 안정적이며 무게중심이 자연스럽다. 전경의 모니터와 배경의 상자들 역시 표면 위에 정상적으로 지지되어 있다."
       },
       {
        "label": "B",
        "direction": "윌마의 시선은 화면 좌측 전경에 앵글이 틀어진 채 배치된 모니터 쪽을 향하고 있다.",
        "built_space": "좌측 전경에 모니터가 있고 그 바로 뒤편 책상 위에 황갈색 케이스가 놓여 있어 이전 샷의 기기 배치를 잘 따르고 있다. 하지만 프롬프트에 명시된 붉은 조명이 있는 문은 배경에서 보이지 않는다.",
        "entities": "윌마의 이목구비는 레퍼런스와 일치하지만, 의상 지침을 어기고 레퍼런스에 없는 두꺼운 검은색 가슴 하네스와 성조기 패치를 착용하고 있다. 또한 폰을 덮고 있어야 할 투명 아크릴 격리 상자가 묘사되지 않아 폰이 케이스 위에 그대로 노출되어 있다.",
        "hard_violations": [
         "invented objects (캐릭터 레퍼런스에 없는 검은색 가슴 하네스 및 패치 추가)"
        ],
        "physics": "인물의 어깨와 자세, 그리고 책상 위에 놓인 장비들은 모두 물리 법칙에 맞게 자연스럽게 지지되어 있다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "모니터를 향한 시선과 찌푸린 표정은 정확하지만, 얼굴 중심 클로즈업이 B보다 느슨하고 레퍼런스에 없는 어깨 하네스와 성조기 패치가 의상 연속성을 해친다."
       },
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "윌마의 얼굴을 밀도 높은 클로즈업으로 담고, 화면이 보이지 않는 모니터 가장자리를 전경에 둔 채 그 모니터를 정확히 응시하는 순간을 가장 충실히 구현했다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "윌마의 두 눈과 얼굴은 프레임 왼쪽 전경의 외부 모니터를 향한다. 모니터의 화면 면은 윌마 쪽으로 돌아가 있어 카메라에는 내용이 보이지 않으며, 시선은 실제 모니터 위치에 닿는다.",
        "built_space": "지하 통제실 내부로 보이며 외부 모니터 1대가 왼쪽 전경에 있다. 뒤에는 케이블이 연결된 휴대전화와 낡은 장비 상자, 희미한 투명 격리 구조가 보인다. 윌마는 모니터 바로 앞에 앉거나 몸을 기울인 위치이며 광학적으로 불가능한 반사는 없다. 다만 모니터가 부드러운 가장자리라기보다 화면 왼쪽의 큰 덩어리로 보이고, 격리 케이스의 형태가 레퍼런스보다 불분명하다.",
        "entities": "보이는 사람은 윌마 한 명뿐이며, 20대 후반의 백인 미국 여성으로 읽히고 밝은 갈색 머리와 얼굴 인상은 캐릭터 레퍼런스에 대체로 부합한다. 녹색 작업복은 맞지만 레퍼런스에 없는 검은 어깨 하네스와 성조기 패치가 추가되었다. 외부 모니터와 유선 휴대전화 관련 장비는 존재하나 2% 화면 내용은 카메라에서 보이지 않는다.",
        "hard_violations": [],
        "physics": "윌마의 상체는 모니터 쪽으로 자연스럽게 기울어 있고 아래 프레임 밖의 좌석 또는 작업대에 기대어 있는 자세로 읽힌다. 팔은 하단에서 잘렸지만 몸이나 물체가 공중에 떠 있지 않으며, 휴대전화는 장비 상자 위에 놓이고 케이블이 연결되어 있다."
       },
       {
        "label": "B",
        "direction": "윌마의 두 눈은 프레임 오른쪽 전경의 외부 모니터를 정확히 응시한다. 모니터 화면은 윌마 쪽을 향하고 카메라는 뒷면과 얇은 측면만 보므로 사용 방향이 올바르며, 시선의 목표도 명확하다.",
        "built_space": "외부 모니터 1대가 오른쪽 전경에 있고 윌마는 그 바로 앞 왼쪽에 자리한다. 뒤에는 투명 격리 케이스 1개와 내부 장비가 흐릿하게 보이며, 통제실의 금속·산업 설비와 붉은 내부 조명이 이어진다. 반사나 중복 피팅은 없고 모니터는 얕은 초점의 전경 가장자리로 기능한다.",
        "entities": "사람은 윌마 한 명만 보인다. 20대 후반의 백인 미국 여성, 밝은 갈색으로 묶은 머리, 얼굴 구조와 체격이 캐릭터 레퍼런스에 가깝고 녹색 작업복도 일치한다. 외부 모니터와 투명 격리 케이스가 확인되며, 휴대전화와 2% 표시는 얕은 초점과 가림 때문에 판독되지 않아 금지된 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "윌마는 모니터 앞에 앉거나 기대어 있는 안정된 상체 자세이며 목과 어깨의 하중이 자연스럽다. 프레임 밖 하체를 억지로 드러내지 않았고, 공중에 뜬 신체나 지지 없이 떠 있는 물체는 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "모니터를 향한 시선과 찌푸린 표정은 정확하지만, 얼굴 중심 클로즈업이 B보다 느슨하고 레퍼런스에 없는 어깨 하네스와 성조기 패치가 의상 연속성을 해친다."
       },
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "윌마의 얼굴을 밀도 높은 클로즈업으로 담고, 화면이 보이지 않는 모니터 가장자리를 전경에 둔 채 그 모니터를 정확히 응시하는 순간을 가장 충실히 구현했다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "윌마의 두 눈과 얼굴은 프레임 왼쪽 전경의 외부 모니터를 향한다. 모니터의 화면 면은 윌마 쪽으로 돌아가 있어 카메라에는 내용이 보이지 않으며, 시선은 실제 모니터 위치에 닿는다.",
        "built_space": "지하 통제실 내부로 보이며 외부 모니터 1대가 왼쪽 전경에 있다. 뒤에는 케이블이 연결된 휴대전화와 낡은 장비 상자, 희미한 투명 격리 구조가 보인다. 윌마는 모니터 바로 앞에 앉거나 몸을 기울인 위치이며 광학적으로 불가능한 반사는 없다. 다만 모니터가 부드러운 가장자리라기보다 화면 왼쪽의 큰 덩어리로 보이고, 격리 케이스의 형태가 레퍼런스보다 불분명하다.",
        "entities": "보이는 사람은 윌마 한 명뿐이며, 20대 후반의 백인 미국 여성으로 읽히고 밝은 갈색 머리와 얼굴 인상은 캐릭터 레퍼런스에 대체로 부합한다. 녹색 작업복은 맞지만 레퍼런스에 없는 검은 어깨 하네스와 성조기 패치가 추가되었다. 외부 모니터와 유선 휴대전화 관련 장비는 존재하나 2% 화면 내용은 카메라에서 보이지 않는다.",
        "hard_violations": [],
        "physics": "윌마의 상체는 모니터 쪽으로 자연스럽게 기울어 있고 아래 프레임 밖의 좌석 또는 작업대에 기대어 있는 자세로 읽힌다. 팔은 하단에서 잘렸지만 몸이나 물체가 공중에 떠 있지 않으며, 휴대전화는 장비 상자 위에 놓이고 케이블이 연결되어 있다."
       },
       {
        "label": "A",
        "direction": "윌마의 두 눈은 프레임 오른쪽 전경의 외부 모니터를 정확히 응시한다. 모니터 화면은 윌마 쪽을 향하고 카메라는 뒷면과 얇은 측면만 보므로 사용 방향이 올바르며, 시선의 목표도 명확하다.",
        "built_space": "외부 모니터 1대가 오른쪽 전경에 있고 윌마는 그 바로 앞 왼쪽에 자리한다. 뒤에는 투명 격리 케이스 1개와 내부 장비가 흐릿하게 보이며, 통제실의 금속·산업 설비와 붉은 내부 조명이 이어진다. 반사나 중복 피팅은 없고 모니터는 얕은 초점의 전경 가장자리로 기능한다.",
        "entities": "사람은 윌마 한 명만 보인다. 20대 후반의 백인 미국 여성, 밝은 갈색으로 묶은 머리, 얼굴 구조와 체격이 캐릭터 레퍼런스에 가깝고 녹색 작업복도 일치한다. 외부 모니터와 투명 격리 케이스가 확인되며, 휴대전화와 2% 표시는 얕은 초점과 가림 때문에 판독되지 않아 금지된 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "윌마는 모니터 앞에 앉거나 기대어 있는 안정된 상체 자세이며 목과 어깨의 하중이 자연스럽다. 프레임 밖 하체를 억지로 드러내지 않았고, 공중에 뜬 신체나 지지 없이 떠 있는 물체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.278
   },
   "adjusted": {
    "A": 2.0,
    "B": 1.028
   },
   "violations": {
    "B": [
     "[gemini-pro] invented objects (캐릭터 레퍼런스에 없는 검은색 가슴 하네스 및 패치 추가)"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 1028
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "레퍼런스와 일치하는 복장을 입고 명시된 투명 격리 상자를 묘사하여 프롬프트의 요구사항을 충실히 따랐으나, 격리 상자의 공간적 위치가 이전 샷과 어긋나게 배경으로 이동한 점이 아쉽습니다."
   },
   {
    "label": "B",
    "score": 1028,
    "verdict_ko": "모니터와 황갈색 케이스의 공간 배치는 이전 샷과 잘 맞으나, 필수적인 투명 아크릴 격리 상자가 완전히 누락되었고 레퍼런스에 없는 가슴 하네스를 착용하여 주요 지침을 위반했습니다.  ★위반: [gemini-pro] invented objects (캐릭터 레퍼런스에 없는 검은색 가슴 하네스 및 패치 추가)"
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features, lighting mood and each person's clothing are LOCKED to this photo; never copy its camera framing. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/37154490-a4ed-4b37-9092-489c054d8a72/images/cc1d3d0b-626e-4f20-bc33-f5f9e2c6ee7e/scene/recipe/S2sh8_sel.png",
    "asset_id": "a8c49db8-e119-49e9-b25c-cf8fe45dace4",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 윌마 디어링: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:766962>",
    "asset_id": "4982a15f-bf33-4ea0-b9c8-11e549bdf723",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": true,
  "shot_run_uid": "06a9bd96-1183-7361-ad8a-1f82a014b5d3",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S2sh8"
  }
 },
 "S2sh11::cine": {
  "applied": true,
  "attempted_at": "2026-09-05T08:58:32.469703+00:00",
  "fingerprint": "e2d2b682e46ddfe50252847242bba257cde0fcbdf177c932c92307bfe1a1cbfb",
  "fingerprint_version": 2,
  "provider": "grok",
  "endpoint": "openrouter/chat-completions",
  "model": "x-ai/grok-imagine-image-2.0",
  "pack": "24.202608252115",
  "source_file": "S2sh11_sel.png",
  "source_sha256": "2b0648df839830ce9ee3b300e73ef076f13210cbc3c98bf718aa09fe9ed5896a",
  "file": "S2sh11_cine.png",
  "staged_sha256": "62fa08b0c0c9ea12ea97fd5410ba57e717ef3e3508100ee05b1c97fe63c119df",
  "latency_ms": 13114
 },
 "S3sh8::signage": {
  "fp": "c2803be887a837c3",
  "inscriptions": [],
  "cues": [
   {
    "text_native": "",
    "source": "scene_text_implied",
    "source_quote": "사원증"
   }
  ],
  "dropped": []
 },
 "S3sh8::bgfirst_bg": {
  "input_fingerprint": "452d12fb812f2851",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 사원증을 든 채 홀로그램 빌 헌을 응시하는 토니의 측면\n\nLOCATION (lock): Inside the mountain contact post's underground control room, beside the active holographic communication disc. Its monochrome projected light reaches the man holding the broken identification card.\n\nTIME OF DAY (lock): twilight.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: 토니(앤서니 로저스) in the middle-left of the frame, foreground, looks toward Bill Hern hologram; 빌 헌 (홀로그램 투사) in the upper-right of the frame, midground, looks toward Tony; ultraphone disc in the lower-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: ultraphone disc (on and projecting Bill Hern's hologram) — The upper face is visible beneath the projected figure, angled diagonally away from the camera; used as Lower-right spatial anchor establishing the hologram's projection origin; cracked identification card (held up by Tony for inspection) — The damaged identification face is angled partly toward Bill and partly toward the camera, keeping its cracked condition readable; used as Foreground evidence linking Tony's hand to the interrogation.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Muted low-key ambient rendering preserves the monochrome holographic outline as the shot's restrained technological accent.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 사원증을 든 채 홀로그램 빌 헌을 응시하는 토니의 측면\n\nLOCATION (lock): Inside the mountain contact post's underground control room, beside the active holographic communication disc. Its monochrome projected light reaches the man holding the broken identification card.\n\nTIME OF DAY (lock): twilight.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: 토니(앤서니 로저스) in the middle-left of the frame, foreground, looks toward Bill Hern hologram; 빌 헌 (홀로그램 투사) in the upper-right of the frame, midground, looks toward Tony; ultraphone disc in the lower-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: ultraphone disc (on and projecting Bill Hern's hologram) — The upper face is visible beneath the projected figure, angled diagonally away from the camera; used as Lower-right spatial anchor establishing the hologram's projection origin; cracked identification card (held up by Tony for inspection) — The damaged identification face is angled partly toward Bill and partly toward the camera, keeping its cracked condition readable; used as Foreground evidence linking Tony's hand to the interrogation.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Muted low-key ambient rendering preserves the monochrome holographic outline as the shot's restrained technological accent.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/37154490-a4ed-4b37-9092-489c054d8a72/images/cc1d3d0b-626e-4f20-bc33-f5f9e2c6ee7e/scene/recipe/S3sh8__bgfirst_bg.png",
  "asset_id": "69d82add-2a90-4e95-84cf-1ef03031fdd4",
  "input_asset_ids": [
   "0d6982f0-e81b-4f25-81ca-85e895df43a3",
   "eb7f72c2-429d-4eb2-880a-e3c8b93c34f5"
  ]
 },
 "S3sh8": {
  "input_fingerprint": "894a805822b29f55",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — Wyoming, United States, 2419; all people are English-speaking Americans unless stated. TIME OF DAY (lock): twilight.\n\nSHOT TEXT (authoritative, Korean): 사원증을 든 채 홀로그램 빌 헌을 응시하는 토니의 측면\n\nLOCATION (lock): Inside the mountain contact post's underground control room, beside the active holographic communication disc. Its monochrome projected light reaches the man holding the broken identification card. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: 토니(앤서니 로저스) in the middle-left of the frame, foreground, looks toward Bill Hern hologram; 빌 헌 (홀로그램 투사) in the upper-right of the frame, midground, looks toward Tony; ultraphone disc in the lower-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: ultraphone disc (on and projecting Bill Hern's hologram) — The upper face is visible beneath the projected figure, angled diagonally away from the camera; used as Lower-right spatial anchor establishing the hologram's projection origin; cracked identification card (held up by Tony for inspection) — The damaged identification face is angled partly toward Bill and partly toward the camera, keeping its cracked condition readable; used as Foreground evidence linking Tony's hand to the interrogation.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Muted low-key ambient rendering preserves the monochrome holographic outline as the shot's restrained technological accent.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The wired phone remains isolated and mirrored to the external monitor, where the unfinished message to Mara is available above an unpressed playback button. The ultraphone disc projects monochrome holographic outlines. 토니(앤서니 로저스): He holds his broken employee ID badge while the long-gun muzzle remains pressed to his chest.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 토니 right now, so 토니's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 토니: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 토니(앤서니 로저스) (미국인 남성, 30대 중반, 자연스러운 성인 남성 얼굴, 구체적으로 명시되지 않은 자연색 머리카락) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — Wyoming, United States, 2419; all people are English-speaking Americans unless stated. TIME OF DAY (lock): twilight.\n\nSHOT TEXT (authoritative, Korean): 사원증을 든 채 홀로그램 빌 헌을 응시하는 토니의 측면\n\nLOCATION (lock): Inside the mountain contact post's underground control room, beside the active holographic communication disc. Its monochrome projected light reaches the man holding the broken identification card. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: 토니(앤서니 로저스) in the middle-left of the frame, foreground, looks toward Bill Hern hologram; 빌 헌 (홀로그램 투사) in the upper-right of the frame, midground, looks toward Tony; ultraphone disc in the lower-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: ultraphone disc (on and projecting Bill Hern's hologram) — The upper face is visible beneath the projected figure, angled diagonally away from the camera; used as Lower-right spatial anchor establishing the hologram's projection origin; cracked identification card (held up by Tony for inspection) — The damaged identification face is angled partly toward Bill and partly toward the camera, keeping its cracked condition readable; used as Foreground evidence linking Tony's hand to the interrogation.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Muted low-key ambient rendering preserves the monochrome holographic outline as the shot's restrained technological accent.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The wired phone remains isolated and mirrored to the external monitor, where the unfinished message to Mara is available above an unpressed playback button. The ultraphone disc projects monochrome holographic outlines. 토니(앤서니 로저스): He holds his broken employee ID badge while the long-gun muzzle remains pressed to his chest.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 토니 right now, so 토니's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 토니: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 토니(앤서니 로저스) (미국인 남성, 30대 중반, 자연스러운 성인 남성 얼굴, 구체적으로 명시되지 않은 자연색 머리카락) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — Wyoming, United States, 2419; all people are English-speaking Americans unless stated. TIME OF DAY (lock): twilight.\n\nSHOT TEXT (authoritative, Korean): 사원증을 든 채 홀로그램 빌 헌을 응시하는 토니의 측면\n\nLOCATION (lock): Inside the mountain contact post's underground control room, beside the active holographic communication disc. Its monochrome projected light reaches the man holding the broken identification card. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: 토니(앤서니 로저스) in the middle-left of the frame, foreground, looks toward Bill Hern hologram; 빌 헌 (홀로그램 투사) in the upper-right of the frame, midground, looks toward Tony; ultraphone disc in the lower-right of the frame, midground.\n- KEY BACKGROUND ELEMENTS: ultraphone disc (on and projecting Bill Hern's hologram) — The upper face is visible beneath the projected figure, angled diagonally away from the camera; used as Lower-right spatial anchor establishing the hologram's projection origin; cracked identification card (held up by Tony for inspection) — The damaged identification face is angled partly toward Bill and partly toward the camera, keeping its cracked condition readable; used as Foreground evidence linking Tony's hand to the interrogation.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Muted low-key ambient rendering preserves the monochrome holographic outline as the shot's restrained technological accent.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The wired phone remains isolated and mirrored to the external monitor, where the unfinished message to Mara is available above an unpressed playback button. The ultraphone disc projects monochrome holographic outlines. 토니(앤서니 로저스): He holds his broken employee ID badge while the long-gun muzzle remains pressed to his chest.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 토니 right now, so 토니's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 토니: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 토니(앤서니 로저스) (미국인 남성, 30대 중반, 자연스러운 성인 남성 얼굴, 구체적으로 명시되지 않은 자연색 머리카락) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/37154490-a4ed-4b37-9092-489c054d8a72/images/cc1d3d0b-626e-4f20-bc33-f5f9e2c6ee7e/scene/recipe/S3sh8__bgfirst_bg.png",
     "asset_id": "69d82add-2a90-4e95-84cf-1ef03031fdd4",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/37154490-a4ed-4b37-9092-489c054d8a72/images/cc1d3d0b-626e-4f20-bc33-f5f9e2c6ee7e/conti/conti_S3sh8.png",
     "asset_id": "0d6982f0-e81b-4f25-81ca-85e895df43a3",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 토니(앤서니 로저스): the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:929851>",
     "asset_id": "1d3c6fe6-553b-4c00-b68f-025be6d786c1",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/37154490-a4ed-4b37-9092-489c054d8a72/episodes/cc1d3d0b-626e-4f20-bc33-f5f9e2c6ee7e/images/background_chain/L11B02.png",
     "asset_id": "eb7f72c2-429d-4eb2-880a-e3c8b93c34f5",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 토니(앤서니 로저스): the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:929851>",
     "asset_id": "1d3c6fe6-553b-4c00-b68f-025be6d786c1",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "토니는 우측 상단의 빌 헌 홀로그램을 응시하고 있으며, 홀로그램 역시 토니를 마주보고 있습니다. 토니가 든 사원증은 카메라와 홀로그램 양쪽을 향해 비스듬히 들려 있습니다.",
    "built_space": "지하 통제실의 배경(좌측의 두꺼운 철문, 뒤쪽의 모니터 시스템, 우측의 책상과 투명한 상자)이 레퍼런스 이미지와 일치하게 배치되어 있습니다. 홀로그램 디스크는 우측 하단 구조물 위에 정확히 위치합니다.",
    "entities": "토니는 레퍼런스와 일치하는 얼굴과 복장을 하고 있으며 손에 금이 간 사원증을 들고 있습니다. 그의 다른 손은 장총을 잡고 있으며 총구는 가슴 쪽에 닿아 있습니다. 빌 헌은 흑백 홀로그램 형태로 온전하게 투사되어 있습니다.",
    "hard_violations": [],
    "physics": "토니는 안정적으로 서서 사원증과 총을 지지하고 있으며, 홀로그램은 디스크 기기에서 위쪽으로 자연스럽게 투사되고 있습니다. 지지되지 않고 떠 있는 물체는 없습니다."
   },
   {
    "label": "B",
    "direction": "토니는 우측의 홀로그램을 향해 시선을 두고 있으며, 홀로그램의 시선도 토니를 향하고 있습니다. 총구 하나가 좌측에서 토니의 등 쪽을 향하고 있습니다.",
    "built_space": "붉은 조명이 도는 지하 통제실의 배경 구조물(벽면의 지도, 모니터, 투명 상자 및 홀로그램 디스크)이 레퍼런스와 매우 흡사하게 잘 구현되어 있습니다.",
    "entities": "토니의 외모와 복장은 레퍼런스와 일치하며 장갑을 낀 손으로 사원증을 들고 있습니다. 홀로그램은 상반신만 투사되어 있습니다. 화면 좌측에 명시되지 않은 제3자의 손과 장총이 등장했습니다.",
    "hard_violations": [
     "[gemini-pro] 프롬프트에 명시되지 않은 인물의 신체 일부(좌측에서 총을 든 두 손과 팔)가 프레임에 등장함 (extra bodies/invented people)",
     "[gpt] 화면 왼쪽에서 정체불명의 추가 인물에 속한 두 손과 팔이 들어와 장총을 들고 있다.",
     "[gpt] 왼쪽 지도 화면의 영문 표제와 표식 등 읽을 수 있는 문자가 노출되어 있다."
    ],
    "physics": "토니가 사원증을 들고 있고 홀로그램이 기기에서 투사되는 물리적 상태는 자연스러우나, 좌측의 총을 든 손은 프레임 밖의 보이지 않는 몸에 의존하고 있어 연출상 오류를 낳습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "지정된 인물과 배경, 사원증 및 가슴에 댄 장총의 위치 등 프롬프트의 요구사항을 훌륭하게 구현한 이미지입니다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "화면 좌측에 프롬프트에 없는 정체불명의 팔과 총이 등장하여 중대한 위반(Hard Violation)이 발생했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "토니는 우측 상단의 빌 헌 홀로그램을 응시하고 있으며, 홀로그램 역시 토니를 마주보고 있습니다. 토니가 든 사원증은 카메라와 홀로그램 양쪽을 향해 비스듬히 들려 있습니다.",
        "built_space": "지하 통제실의 배경(좌측의 두꺼운 철문, 뒤쪽의 모니터 시스템, 우측의 책상과 투명한 상자)이 레퍼런스 이미지와 일치하게 배치되어 있습니다. 홀로그램 디스크는 우측 하단 구조물 위에 정확히 위치합니다.",
        "entities": "토니는 레퍼런스와 일치하는 얼굴과 복장을 하고 있으며 손에 금이 간 사원증을 들고 있습니다. 그의 다른 손은 장총을 잡고 있으며 총구는 가슴 쪽에 닿아 있습니다. 빌 헌은 흑백 홀로그램 형태로 온전하게 투사되어 있습니다.",
        "hard_violations": [],
        "physics": "토니는 안정적으로 서서 사원증과 총을 지지하고 있으며, 홀로그램은 디스크 기기에서 위쪽으로 자연스럽게 투사되고 있습니다. 지지되지 않고 떠 있는 물체는 없습니다."
       },
       {
        "label": "B",
        "direction": "토니는 우측의 홀로그램을 향해 시선을 두고 있으며, 홀로그램의 시선도 토니를 향하고 있습니다. 총구 하나가 좌측에서 토니의 등 쪽을 향하고 있습니다.",
        "built_space": "붉은 조명이 도는 지하 통제실의 배경 구조물(벽면의 지도, 모니터, 투명 상자 및 홀로그램 디스크)이 레퍼런스와 매우 흡사하게 잘 구현되어 있습니다.",
        "entities": "토니의 외모와 복장은 레퍼런스와 일치하며 장갑을 낀 손으로 사원증을 들고 있습니다. 홀로그램은 상반신만 투사되어 있습니다. 화면 좌측에 명시되지 않은 제3자의 손과 장총이 등장했습니다.",
        "hard_violations": [
         "프롬프트에 명시되지 않은 인물의 신체 일부(좌측에서 총을 든 두 손과 팔)가 프레임에 등장함 (extra bodies/invented people)"
        ],
        "physics": "토니가 사원증을 들고 있고 홀로그램이 기기에서 투사되는 물리적 상태는 자연스러우나, 좌측의 총을 든 손은 프레임 밖의 보이지 않는 몸에 의존하고 있어 연출상 오류를 낳습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "지정된 인물과 배경, 사원증 및 가슴에 댄 장총의 위치 등 프롬프트의 요구사항을 훌륭하게 구현한 이미지입니다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "화면 좌측에 프롬프트에 없는 정체불명의 팔과 총이 등장하여 중대한 위반(Hard Violation)이 발생했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "토니는 우측 상단의 빌 헌 홀로그램을 응시하고 있으며, 홀로그램 역시 토니를 마주보고 있습니다. 토니가 든 사원증은 카메라와 홀로그램 양쪽을 향해 비스듬히 들려 있습니다.",
        "built_space": "지하 통제실의 배경(좌측의 두꺼운 철문, 뒤쪽의 모니터 시스템, 우측의 책상과 투명한 상자)이 레퍼런스 이미지와 일치하게 배치되어 있습니다. 홀로그램 디스크는 우측 하단 구조물 위에 정확히 위치합니다.",
        "entities": "토니는 레퍼런스와 일치하는 얼굴과 복장을 하고 있으며 손에 금이 간 사원증을 들고 있습니다. 그의 다른 손은 장총을 잡고 있으며 총구는 가슴 쪽에 닿아 있습니다. 빌 헌은 흑백 홀로그램 형태로 온전하게 투사되어 있습니다.",
        "hard_violations": [],
        "physics": "토니는 안정적으로 서서 사원증과 총을 지지하고 있으며, 홀로그램은 디스크 기기에서 위쪽으로 자연스럽게 투사되고 있습니다. 지지되지 않고 떠 있는 물체는 없습니다."
       },
       {
        "label": "B",
        "direction": "토니는 우측의 홀로그램을 향해 시선을 두고 있으며, 홀로그램의 시선도 토니를 향하고 있습니다. 총구 하나가 좌측에서 토니의 등 쪽을 향하고 있습니다.",
        "built_space": "붉은 조명이 도는 지하 통제실의 배경 구조물(벽면의 지도, 모니터, 투명 상자 및 홀로그램 디스크)이 레퍼런스와 매우 흡사하게 잘 구현되어 있습니다.",
        "entities": "토니의 외모와 복장은 레퍼런스와 일치하며 장갑을 낀 손으로 사원증을 들고 있습니다. 홀로그램은 상반신만 투사되어 있습니다. 화면 좌측에 명시되지 않은 제3자의 손과 장총이 등장했습니다.",
        "hard_violations": [
         "프롬프트에 명시되지 않은 인물의 신체 일부(좌측에서 총을 든 두 손과 팔)가 프레임에 등장함 (extra bodies/invented people)"
        ],
        "physics": "토니가 사원증을 들고 있고 홀로그램이 기기에서 투사되는 물리적 상태는 자연스러우나, 좌측의 총을 든 손은 프레임 밖의 보이지 않는 몸에 의존하고 있어 연출상 오류를 낳습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "배치와 상호 시선은 대체로 맞지만, 화면 왼쪽에서 침입한 정체불명의 추가 인물의 두 손이 총을 들고 있으며 배경 문구도 읽혀 사용 불가다."
       },
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "토니의 측면 중경, 상호 시선, 깨진 사원증, 가슴에 닿은 장총 총구, 우상단 빌 홀로그램과 우하단 투사 원반을 정확히 구현했다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "토니는 오른쪽의 빌 헌 홀로그램 얼굴을 응시하고, 빌도 왼쪽의 토니 쪽을 바라본다. 깨진 사원증의 앞면은 빌과 카메라 사이로 향한다. 장총 총구는 왼쪽에서 오른쪽으로 향해 토니의 가슴에 닿아 있다.",
        "built_space": "활성 투사 원반 1개가 우하단에 있고, 그 뒤로 오른쪽 작업대와 투명 장비 케이스 1개, 벽 제어반, 왼쪽 지도 화면이 보인다. 원반 상면은 카메라에 보이며 투사 장치가 홀로그램 바로 아래에 있다. 다만 기준 장소의 후면 중앙 작업석과 출입문 중심의 공간 관계가 크게 바뀌었고, 지도와 작업대 배치도 기준 사진보다 덜 정확하다. 불가능한 반사는 보이지 않는다.",
        "entities": "토니는 30대 중반 백인 미국인 남성으로 보이며 기준 인물의 짧은 갈색 머리, 얼굴, 체격과 어두운 작업복을 잘 따른다. 빌은 완전한 성인 남성의 청백색 단색 홀로그램으로 표현됐다. 깨진 신분증과 투사 원반은 모두 존재한다. 그러나 화면 왼쪽에 토니가 아닌 인물의 손과 팔이 추가됐고, 지도 및 모니터에 일부 읽을 수 있는 영문이 남아 있다.",
        "hard_violations": [
         "화면 왼쪽에서 정체불명의 추가 인물에 속한 두 손과 팔이 들어와 장총을 들고 있다.",
         "왼쪽 지도 화면의 영문 표제와 표식 등 읽을 수 있는 문자가 노출되어 있다."
        ],
        "physics": "토니는 바닥에 서 있는 자세이며 프레임 밖 하체로 체중이 이어진다. 사원증은 토니의 장갑 낀 오른손이 확실히 잡고 있다. 홀로그램은 원반 위 발광 투사기에서 나온 광선으로 지지된다. 장총은 공중에 뜨지 않고 추가 인물의 두 손이 받치지만, 그 지지 주체 자체가 허용되지 않은 인물이다."
       },
       {
        "label": "B",
        "direction": "토니는 왼쪽 측면에서 오른쪽 위의 빌 홀로그램을 똑바로 바라보고, 빌의 머리와 눈도 왼쪽 아래의 토니를 향한다. 토니가 든 깨진 사원증 앞면은 자신의 시야와 빌 쪽을 향하면서 카메라에도 비스듬히 보여 균열이 읽힌다. 장총은 위쪽을 향하며 총구가 토니의 윗가슴에 닿아 있다.",
        "built_space": "우하단에 원형 투사 원반 1개가 있고 상면과 중앙 투사기가 보이며, 홀로그램은 그 위에 정확히 선다. 왼쪽 암반 벽과 금속 출입문, 후면 중앙 지도 화면과 작업대, 작업용 의자 2개, 오른쪽 벽 제어반과 투명 장비 케이스 1개가 기준 장소와 일관되게 배치됐다. 토니는 출입문 옆 전경에 서 있고 빌은 원반 위 중경에 있어 구조물과 충돌하지 않는다. 불가능한 반사는 없다.",
        "entities": "토니는 기준 사진과 부합하는 30대 중반 백인 남성의 얼굴, 짧은 갈색 머리, 체격과 낡은 어두운 복장을 갖췄다. 빌은 정장 차림의 완전한 성인 남성 형상을 지닌 청백색 단색 홀로그램이다. 깨진 사원증, 장총, 활성 투사 원반이 모두 있으며 사원증의 파손도 선명하다. 눈에 확실히 읽히는 문구나 로고는 없다.",
        "hard_violations": [],
        "physics": "토니는 바닥에 선 자연스러운 체중 자세이며, 깨진 사원증은 오른손 손가락으로 잡혀 있다. 장총은 왼손이 확실히 움켜쥐고 몸통에도 기대어 있어 지지되며 총구는 가슴에 접촉한다. 빌의 가상 신체는 원반 중앙 투사기에서 올라오는 빛으로 형성되어 투사 기원이 명확하고, 다른 물체는 떠 있지 않는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "배치와 상호 시선은 대체로 맞지만, 화면 왼쪽에서 침입한 정체불명의 추가 인물의 두 손이 총을 들고 있으며 배경 문구도 읽혀 사용 불가다."
       },
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "토니의 측면 중경, 상호 시선, 깨진 사원증, 가슴에 닿은 장총 총구, 우상단 빌 홀로그램과 우하단 투사 원반을 정확히 구현했다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "토니는 오른쪽의 빌 헌 홀로그램 얼굴을 응시하고, 빌도 왼쪽의 토니 쪽을 바라본다. 깨진 사원증의 앞면은 빌과 카메라 사이로 향한다. 장총 총구는 왼쪽에서 오른쪽으로 향해 토니의 가슴에 닿아 있다.",
        "built_space": "활성 투사 원반 1개가 우하단에 있고, 그 뒤로 오른쪽 작업대와 투명 장비 케이스 1개, 벽 제어반, 왼쪽 지도 화면이 보인다. 원반 상면은 카메라에 보이며 투사 장치가 홀로그램 바로 아래에 있다. 다만 기준 장소의 후면 중앙 작업석과 출입문 중심의 공간 관계가 크게 바뀌었고, 지도와 작업대 배치도 기준 사진보다 덜 정확하다. 불가능한 반사는 보이지 않는다.",
        "entities": "토니는 30대 중반 백인 미국인 남성으로 보이며 기준 인물의 짧은 갈색 머리, 얼굴, 체격과 어두운 작업복을 잘 따른다. 빌은 완전한 성인 남성의 청백색 단색 홀로그램으로 표현됐다. 깨진 신분증과 투사 원반은 모두 존재한다. 그러나 화면 왼쪽에 토니가 아닌 인물의 손과 팔이 추가됐고, 지도 및 모니터에 일부 읽을 수 있는 영문이 남아 있다.",
        "hard_violations": [
         "화면 왼쪽에서 정체불명의 추가 인물에 속한 두 손과 팔이 들어와 장총을 들고 있다.",
         "왼쪽 지도 화면의 영문 표제와 표식 등 읽을 수 있는 문자가 노출되어 있다."
        ],
        "physics": "토니는 바닥에 서 있는 자세이며 프레임 밖 하체로 체중이 이어진다. 사원증은 토니의 장갑 낀 오른손이 확실히 잡고 있다. 홀로그램은 원반 위 발광 투사기에서 나온 광선으로 지지된다. 장총은 공중에 뜨지 않고 추가 인물의 두 손이 받치지만, 그 지지 주체 자체가 허용되지 않은 인물이다."
       },
       {
        "label": "A",
        "direction": "토니는 왼쪽 측면에서 오른쪽 위의 빌 홀로그램을 똑바로 바라보고, 빌의 머리와 눈도 왼쪽 아래의 토니를 향한다. 토니가 든 깨진 사원증 앞면은 자신의 시야와 빌 쪽을 향하면서 카메라에도 비스듬히 보여 균열이 읽힌다. 장총은 위쪽을 향하며 총구가 토니의 윗가슴에 닿아 있다.",
        "built_space": "우하단에 원형 투사 원반 1개가 있고 상면과 중앙 투사기가 보이며, 홀로그램은 그 위에 정확히 선다. 왼쪽 암반 벽과 금속 출입문, 후면 중앙 지도 화면과 작업대, 작업용 의자 2개, 오른쪽 벽 제어반과 투명 장비 케이스 1개가 기준 장소와 일관되게 배치됐다. 토니는 출입문 옆 전경에 서 있고 빌은 원반 위 중경에 있어 구조물과 충돌하지 않는다. 불가능한 반사는 없다.",
        "entities": "토니는 기준 사진과 부합하는 30대 중반 백인 남성의 얼굴, 짧은 갈색 머리, 체격과 낡은 어두운 복장을 갖췄다. 빌은 정장 차림의 완전한 성인 남성 형상을 지닌 청백색 단색 홀로그램이다. 깨진 사원증, 장총, 활성 투사 원반이 모두 있으며 사원증의 파손도 선명하다. 눈에 확실히 읽히는 문구나 로고는 없다.",
        "hard_violations": [],
        "physics": "토니는 바닥에 선 자연스러운 체중 자세이며, 깨진 사원증은 오른손 손가락으로 잡혀 있다. 장총은 왼손이 확실히 움켜쥐고 몸통에도 기대어 있어 지지되며 총구는 가슴에 접촉한다. 빌의 가상 신체는 원반 중앙 투사기에서 올라오는 빛으로 형성되어 투사 기원이 명확하고, 다른 물체는 떠 있지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.444
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.194
   },
   "violations": {
    "B": [
     "[gemini-pro] 프롬프트에 명시되지 않은 인물의 신체 일부(좌측에서 총을 든 두 손과 팔)가 프레임에 등장함 (extra bodies/invented people)",
     "[gpt] 화면 왼쪽에서 정체불명의 추가 인물에 속한 두 손과 팔이 들어와 장총을 들고 있다.",
     "[gpt] 왼쪽 지도 화면의 영문 표제와 표식 등 읽을 수 있는 문자가 노출되어 있다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 194
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "지정된 인물과 배경, 사원증 및 가슴에 댄 장총의 위치 등 프롬프트의 요구사항을 훌륭하게 구현한 이미지입니다."
   },
   {
    "label": "B",
    "score": 194,
    "verdict_ko": "화면 좌측에 프롬프트에 없는 정체불명의 팔과 총이 등장하여 중대한 위반(Hard Violation)이 발생했습니다.  ★위반: [gemini-pro] 프롬프트에 명시되지 않은 인물의 신체 일부(좌측에서 총을 든 두 손과 팔)가 프레임에 등장함 (extra bodies/invented people) / [gpt] 화면 왼쪽에서 정체불명의 추가 인물에 속한 두 손과 팔이 들어와 장총을 들고 있다. / [gpt] 왼쪽 지도 화면의 영문 표제와 표식 등 읽을 수 있는 문자가 노출되어 있다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/37154490-a4ed-4b37-9092-489c054d8a72/episodes/cc1d3d0b-626e-4f20-bc33-f5f9e2c6ee7e/images/background_chain/L11B02.png",
    "asset_id": "eb7f72c2-429d-4eb2-880a-e3c8b93c34f5",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 토니(앤서니 로저스): the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:929851>",
    "asset_id": "1d3c6fe6-553b-4c00-b68f-025be6d786c1",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": true,
  "shot_run_uid": "06a9bd9c-5b5f-7d69-b510-ded75d4b0cf7",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/37154490-a4ed-4b37-9092-489c054d8a72/images/cc1d3d0b-626e-4f20-bc33-f5f9e2c6ee7e/scene/recipe/S3sh8__bgfirst_bg.png",
   "bg_asset_id": "69d82add-2a90-4e95-84cf-1ef03031fdd4",
   "bg_record_key": "S3sh8::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S3sh8::cine": {
  "applied": true,
  "attempted_at": "2026-09-05T09:01:34.219471+00:00",
  "fingerprint": "d68564555f9b47ed5dbce19d583d983e39aaae568fb24b8055d2a21da8798f80",
  "fingerprint_version": 2,
  "provider": "grok",
  "endpoint": "openrouter/chat-completions",
  "model": "x-ai/grok-imagine-image-2.0",
  "pack": "24.202608252115",
  "source_file": "S3sh8_sel.png",
  "source_sha256": "cf556a4e4d6820f68deb849f6ec75773c8832f96fccb47ba85fcc34b62e417f5",
  "file": "S3sh8_cine.png",
  "staged_sha256": "ddc50467709ad54f298dda07c1733f17855ddacf4a233d810617fa6396b4c52d",
  "latency_ms": 14775
 },
 "S3sh12::signage": {
  "fp": "158ebcc39fa9bd45",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S3sh12": {
  "input_fingerprint": "c706af09837b1a3a",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — Wyoming, United States, 2419; all people are English-speaking Americans unless stated. TIME OF DAY (lock): twilight.\n\nSHOT TEXT (authoritative, Korean): 토니의 가슴에 닿아 있는 장총 총구의 클로즈업\n\nLOCATION (lock): Inside the mountain contact post's underground control room, in the confrontation area facing the holographic link. The rifle remains pressed against the captive's chest. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: insert close-up on a detail\n- KEY BACKGROUND ELEMENTS: long rifle (muzzle held against Tony's chest) — The barrel crosses the crop side-on and terminates against Tony's chest, never pointing along the camera axis; used as Isolated threat detail and primary focus.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Neutral low-key ambient light emphasizes the pressure point and restrained tactile contrast without adding an unsupported light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 토니(앤서니 로저스) — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Take the same rock-walled contact-room surfaces, red emergency illumination, and fixed communication fixtures from the reference. Exclude the identity-card emphasis and the prior holographic caller from this shot.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The wired phone remains in the isolation case with its display mirrored externally. The ultraphone disc continues projecting the two monochrome outlines, and the archival system shows a noisy 2026 facility plan labeled “MANUAL DUAL-ACTION SEAL.” 토니(앤서니 로저스): He continues holding the broken employee ID badge with the long-gun muzzle touching his chest.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 토니(앤서니 로저스) (미국인 남성, 30대 중반, 자연스러운 성인 남성 얼굴, 구체적으로 명시되지 않은 자연색 머리카락) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — Wyoming, United States, 2419; all people are English-speaking Americans unless stated. TIME OF DAY (lock): twilight.\n\nSHOT TEXT (authoritative, Korean): 토니의 가슴에 닿아 있는 장총 총구의 클로즈업\n\nLOCATION (lock): Inside the mountain contact post's underground control room, in the confrontation area facing the holographic link. The rifle remains pressed against the captive's chest. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: insert close-up on a detail\n- KEY BACKGROUND ELEMENTS: long rifle (muzzle held against Tony's chest) — The barrel crosses the crop side-on and terminates against Tony's chest, never pointing along the camera axis; used as Isolated threat detail and primary focus.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Neutral low-key ambient light emphasizes the pressure point and restrained tactile contrast without adding an unsupported light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 토니(앤서니 로저스) — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Take the same rock-walled contact-room surfaces, red emergency illumination, and fixed communication fixtures from the reference. Exclude the identity-card emphasis and the prior holographic caller from this shot.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The wired phone remains in the isolation case with its display mirrored externally. The ultraphone disc continues projecting the two monochrome outlines, and the archival system shows a noisy 2026 facility plan labeled “MANUAL DUAL-ACTION SEAL.” 토니(앤서니 로저스): He continues holding the broken employee ID badge with the long-gun muzzle touching his chest.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 토니(앤서니 로저스) (미국인 남성, 30대 중반, 자연스러운 성인 남성 얼굴, 구체적으로 명시되지 않은 자연색 머리카락) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — Wyoming, United States, 2419; all people are English-speaking Americans unless stated. TIME OF DAY (lock): twilight.\n\nSHOT TEXT (authoritative, Korean): 토니의 가슴에 닿아 있는 장총 총구의 클로즈업\n\nLOCATION (lock): Inside the mountain contact post's underground control room, in the confrontation area facing the holographic link. The rifle remains pressed against the captive's chest. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: insert close-up on a detail\n- KEY BACKGROUND ELEMENTS: long rifle (muzzle held against Tony's chest) — The barrel crosses the crop side-on and terminates against Tony's chest, never pointing along the camera axis; used as Isolated threat detail and primary focus.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Neutral low-key ambient light emphasizes the pressure point and restrained tactile contrast without adding an unsupported light source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The clothing, hair and overall look of 토니(앤서니 로저스) — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Take the same rock-walled contact-room surfaces, red emergency illumination, and fixed communication fixtures from the reference. Exclude the identity-card emphasis and the prior holographic caller from this shot.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The wired phone remains in the isolation case with its display mirrored externally. The ultraphone disc continues projecting the two monochrome outlines, and the archival system shows a noisy 2026 facility plan labeled “MANUAL DUAL-ACTION SEAL.” 토니(앤서니 로저스): He continues holding the broken employee ID badge with the long-gun muzzle touching his chest.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 토니(앤서니 로저스) (미국인 남성, 30대 중반, 자연스러운 성인 남성 얼굴, 구체적으로 명시되지 않은 자연색 머리카락) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "장총의 총구가 화면 오른쪽에서 왼쪽으로 향하며 토니의 가슴 정중앙을 겨냥하여 맞닿아 있음.",
    "built_space": "붉은 조명의 지하 통제실로, 뒤쪽에 암석 벽과 흐릿한 제어 콘솔이 이전 샷의 환경과 일치하게 배치되어 있음.",
    "entities": "토니의 가슴 부분(전술 재킷, 하네스, 무전기)이 캐릭터 참조 이미지와 완벽히 일치함. 장총의 총열, 상단 레일, 총구 제퇴기(muzzle brake) 역시 프롭 참조와 일치함.",
    "hard_violations": [],
    "physics": "토니는 화면 중앙에 서 있으며, 장총은 화면 오른쪽 밖에서 누군가에 의해 지지되어 물리적으로 타당하게 가슴을 압박하고 있음."
   },
   {
    "label": "B",
    "direction": "장총의 총구가 화면 왼쪽에서 오른쪽으로 향하며 토니의 가슴(무전기 부근)에 닿아 있음. 배경의 손은 신분증을 들고 카메라 쪽을 향함.",
    "built_space": "암석 벽과 붉은 비상 조명, 배관 및 전기 박스가 배경에 묘사되어 있음.",
    "entities": "오른쪽에 토니의 몸통이 있으나, 화면 왼쪽 중앙에 부서진 신분증을 든 정체불명의 두 손이 나타남. 장총의 외형은 참조와 유사함.",
    "hard_violations": [
     "[gemini-pro] 물리적으로 불가능한 해부학 및 스테이징: 화면 오른쪽에 있는 몸통과 구조적으로 연결될 수 없는 위치(화면 왼쪽 하단/중앙)에서 신분증을 든 손과 팔이 뻗어 나옴."
    ],
    "physics": "장총은 화면 왼쪽 밖에서 지지되고 있음. 그러나 신분증을 든 손은 오른쪽에 위치한 몸통과 분리되어 허공에서 나타나는 물리적 모순을 보임."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 10,
        "verdict_ko": "프롬프트가 요구한 인서트 클로즈업 앵글을 완벽히 구현했으며, 앵글에 따라 가슴에 맞닿은 장총의 디테일과 참조된 의상을 정확히 반영했습니다."
       },
       {
        "label": "B",
        "score": 0,
        "verdict_ko": "오른쪽에 있는 피사체의 몸통과 물리적으로 연결될 수 없는 위치(화면 왼쪽)에 신분증을 든 손이 등장하여 심각한 해부학적/스테이징 오류가 발생했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "장총의 총구가 화면 오른쪽에서 왼쪽으로 향하며 토니의 가슴 정중앙을 겨냥하여 맞닿아 있음.",
        "built_space": "붉은 조명의 지하 통제실로, 뒤쪽에 암석 벽과 흐릿한 제어 콘솔이 이전 샷의 환경과 일치하게 배치되어 있음.",
        "entities": "토니의 가슴 부분(전술 재킷, 하네스, 무전기)이 캐릭터 참조 이미지와 완벽히 일치함. 장총의 총열, 상단 레일, 총구 제퇴기(muzzle brake) 역시 프롭 참조와 일치함.",
        "hard_violations": [],
        "physics": "토니는 화면 중앙에 서 있으며, 장총은 화면 오른쪽 밖에서 누군가에 의해 지지되어 물리적으로 타당하게 가슴을 압박하고 있음."
       },
       {
        "label": "B",
        "direction": "장총의 총구가 화면 왼쪽에서 오른쪽으로 향하며 토니의 가슴(무전기 부근)에 닿아 있음. 배경의 손은 신분증을 들고 카메라 쪽을 향함.",
        "built_space": "암석 벽과 붉은 비상 조명, 배관 및 전기 박스가 배경에 묘사되어 있음.",
        "entities": "오른쪽에 토니의 몸통이 있으나, 화면 왼쪽 중앙에 부서진 신분증을 든 정체불명의 두 손이 나타남. 장총의 외형은 참조와 유사함.",
        "hard_violations": [
         "물리적으로 불가능한 해부학 및 스테이징: 화면 오른쪽에 있는 몸통과 구조적으로 연결될 수 없는 위치(화면 왼쪽 하단/중앙)에서 신분증을 든 손과 팔이 뻗어 나옴."
        ],
        "physics": "장총은 화면 왼쪽 밖에서 지지되고 있음. 그러나 신분증을 든 손은 오른쪽에 위치한 몸통과 분리되어 허공에서 나타나는 물리적 모순을 보임."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 10,
        "verdict_ko": "프롬프트가 요구한 인서트 클로즈업 앵글을 완벽히 구현했으며, 앵글에 따라 가슴에 맞닿은 장총의 디테일과 참조된 의상을 정확히 반영했습니다."
       },
       {
        "label": "B",
        "score": 0,
        "verdict_ko": "오른쪽에 있는 피사체의 몸통과 물리적으로 연결될 수 없는 위치(화면 왼쪽)에 신분증을 든 손이 등장하여 심각한 해부학적/스테이징 오류가 발생했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "장총의 총구가 화면 오른쪽에서 왼쪽으로 향하며 토니의 가슴 정중앙을 겨냥하여 맞닿아 있음.",
        "built_space": "붉은 조명의 지하 통제실로, 뒤쪽에 암석 벽과 흐릿한 제어 콘솔이 이전 샷의 환경과 일치하게 배치되어 있음.",
        "entities": "토니의 가슴 부분(전술 재킷, 하네스, 무전기)이 캐릭터 참조 이미지와 완벽히 일치함. 장총의 총열, 상단 레일, 총구 제퇴기(muzzle brake) 역시 프롭 참조와 일치함.",
        "hard_violations": [],
        "physics": "토니는 화면 중앙에 서 있으며, 장총은 화면 오른쪽 밖에서 누군가에 의해 지지되어 물리적으로 타당하게 가슴을 압박하고 있음."
       },
       {
        "label": "B",
        "direction": "장총의 총구가 화면 왼쪽에서 오른쪽으로 향하며 토니의 가슴(무전기 부근)에 닿아 있음. 배경의 손은 신분증을 들고 카메라 쪽을 향함.",
        "built_space": "암석 벽과 붉은 비상 조명, 배관 및 전기 박스가 배경에 묘사되어 있음.",
        "entities": "오른쪽에 토니의 몸통이 있으나, 화면 왼쪽 중앙에 부서진 신분증을 든 정체불명의 두 손이 나타남. 장총의 외형은 참조와 유사함.",
        "hard_violations": [
         "물리적으로 불가능한 해부학 및 스테이징: 화면 오른쪽에 있는 몸통과 구조적으로 연결될 수 없는 위치(화면 왼쪽 하단/중앙)에서 신분증을 든 손과 팔이 뻗어 나옴."
        ],
        "physics": "장총은 화면 왼쪽 밖에서 지지되고 있음. 그러나 신분증을 든 손은 오른쪽에 위치한 몸통과 분리되어 허공에서 나타나는 물리적 모순을 보임."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "장총 총구가 화면축이 아닌 측면 방향으로 토니의 가슴에 정확히 닿고 압박 지점을 삽입 클로즈업으로 고립해, 요구된 프레이밍과 순간을 가장 충실히 구현했다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "총구 방향과 가슴 접촉은 정확하지만 손과 두 장처럼 보이는 배지 조각 및 천장등까지 넓게 담아, 배지 강조를 배제한 총구 압박점 중심의 삽입 클로즈업보다 주의가 분산된다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "장총은 화면 왼쪽에서 오른쪽으로 측면 횡단하며, 카메라축을 향하지 않는다. 총열 끝의 소염기형 총구가 화면 오른쪽 토니의 가슴 장비와 멜빵 부위에 직접 닿아 토니를 겨눈다. 토니의 얼굴은 턱 일부만 보여 시선은 판별할 수 없다.",
        "built_space": "암벽 벽면, 노출 배관과 케이블, 전기함, 붉은 천장등 한 개가 보이며 이전 통제실의 지하 암반 구조와 비상 조명을 이어간다. 토니는 화면 오른쪽 전경에 서 있고 장총은 왼쪽 바깥에서 들어온다. 고정 설비의 중복이나 불가능한 반사는 없다.",
        "entities": "토니는 30대 중반 미국인 남성의 턱과 참조와 일치하는 낡은 갈색 전술복·하네스로 나타난다. 참조 장총의 검은 금속 총열, 천공형 소염기, 방열 덮개 특징이 확인된다. 토니가 양손에 든 두 직사각형 조각은 파손 배지로 읽힐 수 있으나 서로 별개의 카드 두 장처럼 보여 정확성이 다소 떨어지며, 글자는 판독되지 않는다. 다른 인물은 없다.",
        "hard_violations": [],
        "physics": "토니의 몸은 하단에서 잘린 정상적인 입상 자세이며 부유 징후가 없다. 배지 조각들은 양손이 각각 직접 쥐고 있어 지지된다. 장총은 화면 밖 왼쪽의 가해자가 들고 있는 것으로 자연스럽게 프레임 밖으로 이어지며, 총구가 가슴 장비에 접촉해 압력을 전달하는 배치다."
       },
       {
        "label": "B",
        "direction": "장총은 화면 오른쪽에서 왼쪽으로 측면 횡단하고 카메라축을 향하지 않는다. 총열 끝의 소염기형 총구가 화면 중앙의 토니 가슴 하네스와 장비 바로 아래에 닿아 토니를 정확히 겨눈다. 얼굴과 눈은 프레임 밖이어서 시선은 보이지 않는다.",
        "built_space": "배경에 암벽 벽면, 통제 콘솔과 작업대가 작고 흐리게 배치되며 붉은 저조도 조명이 유지된다. 토니는 전경 중앙에 서 있고 장총은 오른쪽 프레임 밖에서 들어온다. 통제실 설비가 총구보다 과장되지 않았고, 중복 설비나 불가능한 반사는 보이지 않는다.",
        "entities": "보이는 인물은 토니 한 명뿐이며, 참조와 맞는 남성 체격, 낡은 갈색 전술 재킷, 하네스, 가슴 통신장비와 허리 로프를 갖췄다. 장총은 참조의 검은 금속 재질, 긴 총열, 천공형 소염기와 방열 덮개 형태를 충실히 따른다. 배지는 올바른 삽입 프레이밍 때문에 보이지 않으며 이를 감점할 이유가 없다. 판독 가능한 문자는 없다.",
        "hard_violations": [],
        "physics": "토니의 상체는 정상적인 직립 상태로 프레임 아래의 다리와 지면에 의해 지지되는 구도이며 부유하지 않는다. 장총은 오른쪽 프레임 밖의 가해자가 들고 있는 것으로 총열이 자연스럽게 이어지고, 총구는 토니의 가슴 하네스에 실제 접촉한다. 프레임 안에서 지지 없이 떠 있는 물체는 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "장총 총구가 화면축이 아닌 측면 방향으로 토니의 가슴에 정확히 닿고 압박 지점을 삽입 클로즈업으로 고립해, 요구된 프레이밍과 순간을 가장 충실히 구현했다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "총구 방향과 가슴 접촉은 정확하지만 손과 두 장처럼 보이는 배지 조각 및 천장등까지 넓게 담아, 배지 강조를 배제한 총구 압박점 중심의 삽입 클로즈업보다 주의가 분산된다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "장총은 화면 왼쪽에서 오른쪽으로 측면 횡단하며, 카메라축을 향하지 않는다. 총열 끝의 소염기형 총구가 화면 오른쪽 토니의 가슴 장비와 멜빵 부위에 직접 닿아 토니를 겨눈다. 토니의 얼굴은 턱 일부만 보여 시선은 판별할 수 없다.",
        "built_space": "암벽 벽면, 노출 배관과 케이블, 전기함, 붉은 천장등 한 개가 보이며 이전 통제실의 지하 암반 구조와 비상 조명을 이어간다. 토니는 화면 오른쪽 전경에 서 있고 장총은 왼쪽 바깥에서 들어온다. 고정 설비의 중복이나 불가능한 반사는 없다.",
        "entities": "토니는 30대 중반 미국인 남성의 턱과 참조와 일치하는 낡은 갈색 전술복·하네스로 나타난다. 참조 장총의 검은 금속 총열, 천공형 소염기, 방열 덮개 특징이 확인된다. 토니가 양손에 든 두 직사각형 조각은 파손 배지로 읽힐 수 있으나 서로 별개의 카드 두 장처럼 보여 정확성이 다소 떨어지며, 글자는 판독되지 않는다. 다른 인물은 없다.",
        "hard_violations": [],
        "physics": "토니의 몸은 하단에서 잘린 정상적인 입상 자세이며 부유 징후가 없다. 배지 조각들은 양손이 각각 직접 쥐고 있어 지지된다. 장총은 화면 밖 왼쪽의 가해자가 들고 있는 것으로 자연스럽게 프레임 밖으로 이어지며, 총구가 가슴 장비에 접촉해 압력을 전달하는 배치다."
       },
       {
        "label": "A",
        "direction": "장총은 화면 오른쪽에서 왼쪽으로 측면 횡단하고 카메라축을 향하지 않는다. 총열 끝의 소염기형 총구가 화면 중앙의 토니 가슴 하네스와 장비 바로 아래에 닿아 토니를 정확히 겨눈다. 얼굴과 눈은 프레임 밖이어서 시선은 보이지 않는다.",
        "built_space": "배경에 암벽 벽면, 통제 콘솔과 작업대가 작고 흐리게 배치되며 붉은 저조도 조명이 유지된다. 토니는 전경 중앙에 서 있고 장총은 오른쪽 프레임 밖에서 들어온다. 통제실 설비가 총구보다 과장되지 않았고, 중복 설비나 불가능한 반사는 보이지 않는다.",
        "entities": "보이는 인물은 토니 한 명뿐이며, 참조와 맞는 남성 체격, 낡은 갈색 전술 재킷, 하네스, 가슴 통신장비와 허리 로프를 갖췄다. 장총은 참조의 검은 금속 재질, 긴 총열, 천공형 소염기와 방열 덮개 형태를 충실히 따른다. 배지는 올바른 삽입 프레이밍 때문에 보이지 않으며 이를 감점할 이유가 없다. 판독 가능한 문자는 없다.",
        "hard_violations": [],
        "physics": "토니의 상체는 정상적인 직립 상태로 프레임 아래의 다리와 지면에 의해 지지되는 구도이며 부유하지 않는다. 장총은 오른쪽 프레임 밖의 가해자가 들고 있는 것으로 총열이 자연스럽게 이어지고, 총구는 토니의 가슴 하네스에 실제 접촉한다. 프레임 안에서 지지 없이 떠 있는 물체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.778
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.528
   },
   "violations": {
    "B": [
     "[gemini-pro] 물리적으로 불가능한 해부학 및 스테이징: 화면 오른쪽에 있는 몸통과 구조적으로 연결될 수 없는 위치(화면 왼쪽 하단/중앙)에서 신분증을 든 손과 팔이 뻗어 나옴."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 528
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "프롬프트가 요구한 인서트 클로즈업 앵글을 완벽히 구현했으며, 앵글에 따라 가슴에 맞닿은 장총의 디테일과 참조된 의상을 정확히 반영했습니다."
   },
   {
    "label": "B",
    "score": 528,
    "verdict_ko": "오른쪽에 있는 피사체의 몸통과 물리적으로 연결될 수 없는 위치(화면 왼쪽)에 신분증을 든 손이 등장하여 심각한 해부학적/스테이징 오류가 발생했습니다.  ★위반: [gemini-pro] 물리적으로 불가능한 해부학 및 스테이징: 화면 오른쪽에 있는 몸통과 구조적으로 연결될 수 없는 위치(화면 왼쪽 하단/중앙)에서 신분증을 든 손과 팔이 뻗어 나옴."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The clothing, hair and overall look of 토니(앤서니 로저스) — who appear both in that photo and in this shot — are LOCKED to that photo. Anyone else visible in that photo is NOT in this shot: never carry their face, body or clothing onto anyone here. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/37154490-a4ed-4b37-9092-489c054d8a72/images/cc1d3d0b-626e-4f20-bc33-f5f9e2c6ee7e/scene/recipe/S1sh9_sel.png",
    "asset_id": "2344892f-3713-4f6f-8eac-d95f72b5b919",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 토니(앤서니 로저스): the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:929851>",
    "asset_id": "1d3c6fe6-553b-4c00-b68f-025be6d786c1",
    "role": "character_ref"
   },
   {
    "label": "PROP REFERENCE — 장총: the exact object appearing in this shot; match its look, material and wear exactly.",
    "path": "<bytes:785972>",
    "asset_id": "3c9dd37a-da89-4fa6-a8b0-83c72e84ff1f",
    "role": "prop_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": true,
  "shot_run_uid": "06a9bda7-d1f8-78a8-ac5d-d3af55d7c96d",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S1sh9"
  }
 },
 "S3sh12::cine": {
  "applied": true,
  "attempted_at": "2026-09-05T09:02:57.600664+00:00",
  "fingerprint": "bbde9ae33d83501dca00f81e57863d48fde093e03d90ac62b24166d70b6c6a2c",
  "fingerprint_version": 2,
  "provider": "grok",
  "endpoint": "openrouter/chat-completions",
  "model": "x-ai/grok-imagine-image-2.0",
  "pack": "24.202608252115",
  "source_file": "S3sh12_sel.png",
  "source_sha256": "663f9921089696e659621781151a54fb2a64837fe712fa629b8e1fa8c307d790",
  "file": "S3sh12_cine.png",
  "staged_sha256": "364e1ef445d6421ba1b9cce6d2c9a03001ccb2fbe3d7d51dd6f8034b8bec1a6e",
  "latency_ms": 17909
 },
 "S3sh13::signage": {
  "fp": "0aba99d186ea91b5",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S3sh13": {
  "input_fingerprint": "1e37dbb2e88ab290",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — Wyoming, United States, 2419; all people are English-speaking Americans unless stated. TIME OF DAY (lock): twilight.\n\nSHOT TEXT (authoritative, Korean): 앨런의 두 손에 쥐어진 장총 총구가 아래로 살짝 기울어진 찰나\n\nLOCATION (lock): Inside the mountain contact post's underground control room, between the archive display and the standing captive. Electronic displays and the monochrome link provide the visible light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: records display (showing the 2026 facility drawing and “MANUAL DUAL-ACTION SEAL” instruction through noise) — The active content face lies beneath Alan's downward eyeline, with the facility drawing and seal instruction visible before the tilt leaves it; used as Verification evidence at the top of the tilt path; long rifle (held in both hands as the muzzle dips slightly) — Seen from above and behind Alan's working side, the barrel angles away from Tony's chest and toward the lower frame; used as Primary reaction detail revealing Alan's first reduction of force.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Neutral low-key ambient illumination keeps the verification display readable while holding Alan's hands and weapon in restrained contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Take the same rough stone walls, red ambient light, and installed communication equipment from the reference. Exclude the displayed identity card and any prior holographic figure not specified for this shot.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The 2026 facility plan labeled “MANUAL DUAL-ACTION SEAL” remains on the archival display. The wired phone stays isolated and mirrored, while the ultraphone disc maintains the monochrome projections. 앨런: He still holds the long gun, but its muzzle is now lowered a few centimeters from its former aim.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 앨런 right now, so 앨런's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 앨런: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앨런 (미국계 성인 남성, 자연스러운 인간 얼굴, 자연 색상의 머리카락, 구체적 얼굴형은 명시되지 않음) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — Wyoming, United States, 2419; all people are English-speaking Americans unless stated. TIME OF DAY (lock): twilight.\n\nSHOT TEXT (authoritative, Korean): 앨런의 두 손에 쥐어진 장총 총구가 아래로 살짝 기울어진 찰나\n\nLOCATION (lock): Inside the mountain contact post's underground control room, between the archive display and the standing captive. Electronic displays and the monochrome link provide the visible light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: records display (showing the 2026 facility drawing and “MANUAL DUAL-ACTION SEAL” instruction through noise) — The active content face lies beneath Alan's downward eyeline, with the facility drawing and seal instruction visible before the tilt leaves it; used as Verification evidence at the top of the tilt path; long rifle (held in both hands as the muzzle dips slightly) — Seen from above and behind Alan's working side, the barrel angles away from Tony's chest and toward the lower frame; used as Primary reaction detail revealing Alan's first reduction of force.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Neutral low-key ambient illumination keeps the verification display readable while holding Alan's hands and weapon in restrained contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Take the same rough stone walls, red ambient light, and installed communication equipment from the reference. Exclude the displayed identity card and any prior holographic figure not specified for this shot.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The 2026 facility plan labeled “MANUAL DUAL-ACTION SEAL” remains on the archival display. The wired phone stays isolated and mirrored, while the ultraphone disc maintains the monochrome projections. 앨런: He still holds the long gun, but its muzzle is now lowered a few centimeters from its former aim.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 앨런 right now, so 앨런's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 앨런: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앨런 (미국계 성인 남성, 자연스러운 인간 얼굴, 자연 색상의 머리카락, 구체적 얼굴형은 명시되지 않음) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — Wyoming, United States, 2419; all people are English-speaking Americans unless stated. TIME OF DAY (lock): twilight.\n\nSHOT TEXT (authoritative, Korean): 앨런의 두 손에 쥐어진 장총 총구가 아래로 살짝 기울어진 찰나\n\nLOCATION (lock): Inside the mountain contact post's underground control room, between the archive display and the standing captive. Electronic displays and the monochrome link provide the visible light. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: records display (showing the 2026 facility drawing and “MANUAL DUAL-ACTION SEAL” instruction through noise) — The active content face lies beneath Alan's downward eyeline, with the facility drawing and seal instruction visible before the tilt leaves it; used as Verification evidence at the top of the tilt path; long rifle (held in both hands as the muzzle dips slightly) — Seen from above and behind Alan's working side, the barrel angles away from Tony's chest and toward the lower frame; used as Primary reaction detail revealing Alan's first reduction of force.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Neutral low-key ambient illumination keeps the verification display readable while holding Alan's hands and weapon in restrained contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Take the same rough stone walls, red ambient light, and installed communication equipment from the reference. Exclude the displayed identity card and any prior holographic figure not specified for this shot.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The 2026 facility plan labeled “MANUAL DUAL-ACTION SEAL” remains on the archival display. The wired phone stays isolated and mirrored, while the ultraphone disc maintains the monochrome projections. 앨런: He still holds the long gun, but its muzzle is now lowered a few centimeters from its former aim.\n\nTHE HAND THAT IS DOING THIS: the object at the centre of this shot is being held, operated, read, aimed or handed over by 앨런 right now, so 앨런's hand — and as much of the wrist and forearm as the framing reaches — is in the frame, gripping or resting on that object exactly the way its real use requires. Match that hand to 앨런: its size, build, skin, age, grooming, sleeve and anything worn on it belong to that person and to no one else.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앨런 (미국계 성인 남성, 자연스러운 인간 얼굴, 자연 색상의 머리카락, 구체적 얼굴형은 명시되지 않음) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "initial_roll_all_fail": true,
  "readings": [
   {
    "label": "A",
    "direction": "앨런의 시선은 아래쪽에 놓인 디스플레이와 장총을 향함. 장총의 총구는 화면 좌측 하단을 향해 살짝 기울어짐.",
    "built_space": "거친 돌벽과 붉은 조명이 배경을 이루며, 디스플레이가 책상 위에 놓여 있음. 그러나 배경의 벽면과 앞쪽 책상 위에 명시되지 않은 사무용 유선 전화기가 두 개나 배치되어 공간의 일관성을 해침.",
    "entities": "앨런은 레퍼런스와 일치하는 복장(비니, 재킷)을 착용함. 장총은 두 손에 쥐어져 있으나 프랍 레퍼런스에 있는 대형 스코프가 완전히 누락됨. 디스플레이에는 시설 도면과 'MANUAL DUAL-ACTION SEAL' 텍스트가 정확히 표기됨.",
    "hard_violations": [
     "[gemini-pro] physically impossible anatomy (방아쇠를 쥔 손이 손등을 보이고 엄지가 앞을 향하는 등 정상적인 오른손이나 왼손으로 불가능한 구조임)",
     "[gemini-pro] invented objects (책상 위에 원본에 없는 90년대풍 현대식 유선 전화기가 생성됨)",
     "[gemini-pro] duplicated fitting (전화기가 벽면과 책상 두 곳에 중복으로 존재함)",
     "[gpt] 기록 디스플레이에 금지된 판독 가능한 영어 문구가 선명하게 노출되어 있다."
    ],
    "physics": "장총은 앨런의 양손에 의해 허공에 지지되고 있으나, 개머리판 쪽을 쥔 손의 관절과 엄지손가락 위치가 물리적/해부학적으로 불가능하게 뒤틀려 있음."
   },
   {
    "label": "B",
    "direction": "앨런의 시선은 손에 쥔 장총과 디스플레이 쪽을 향함. 장총의 총구는 화면 우측 하단으로 향함.",
    "built_space": "거친 돌벽과 붉은 조명 아래, 책상 위에 디스플레이가 놓인 구조. 공간적 피팅은 비교적 단순하게 처리됨.",
    "entities": "앨런의 얼굴은 유지되었으나 캐릭터 레퍼런스의 필수 복장인 비니가 누락됨. 장총은 두 손에 쥐어졌으나 대형 스코프가 누락됨. 디스플레이에 'MANUAL DUAL-ACTION SEAL' 텍스트가 확인됨.",
    "hard_violations": [
     "[gemini-pro] physically impossible anatomy (오른팔에 달린 방아쇠 쪽 손의 엄지가 카메라 쪽인 우측에 위치해 있어 사실상 두 개의 왼손을 가진 셈이 됨)",
     "[gpt] 기록 디스플레이에 금지된 판독 가능한 영어 문구가 선명하게 노출되어 있다."
    ],
    "physics": "양손이 장총을 받치고 있으나, 카메라와 가까운 쪽인 오른손이 총의 그립을 쥘 때 엄지가 총의 우측(카메라 방향)에 위치해 있어 물리적으로 불가능한 쥐기 자세를 연출함."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "캐릭터의 복장(비니)과 디스플레이 내용 구현은 우수하나, 방아쇠를 쥔 손의 치명적인 해부학적 오류, 장총 스코프 누락, 그리고 지시된 카메라 앵글 위반 및 불필요한 사물 추가로 인해 사용할 수 없습니다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "장총의 스코프 누락과 캐릭터의 지정된 복장(비니) 누락이 있으며, 방아쇠를 쥔 오른팔에 왼손이 달려있는 심각한 해부학적 오류로 인해 완전히 실패한 결과물입니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "앨런의 시선은 아래쪽에 놓인 디스플레이와 장총을 향함. 장총의 총구는 화면 좌측 하단을 향해 살짝 기울어짐.",
        "built_space": "거친 돌벽과 붉은 조명이 배경을 이루며, 디스플레이가 책상 위에 놓여 있음. 그러나 배경의 벽면과 앞쪽 책상 위에 명시되지 않은 사무용 유선 전화기가 두 개나 배치되어 공간의 일관성을 해침.",
        "entities": "앨런은 레퍼런스와 일치하는 복장(비니, 재킷)을 착용함. 장총은 두 손에 쥐어져 있으나 프랍 레퍼런스에 있는 대형 스코프가 완전히 누락됨. 디스플레이에는 시설 도면과 'MANUAL DUAL-ACTION SEAL' 텍스트가 정확히 표기됨.",
        "hard_violations": [
         "physically impossible anatomy (방아쇠를 쥔 손이 손등을 보이고 엄지가 앞을 향하는 등 정상적인 오른손이나 왼손으로 불가능한 구조임)",
         "invented objects (책상 위에 원본에 없는 90년대풍 현대식 유선 전화기가 생성됨)",
         "duplicated fitting (전화기가 벽면과 책상 두 곳에 중복으로 존재함)"
        ],
        "physics": "장총은 앨런의 양손에 의해 허공에 지지되고 있으나, 개머리판 쪽을 쥔 손의 관절과 엄지손가락 위치가 물리적/해부학적으로 불가능하게 뒤틀려 있음."
       },
       {
        "label": "B",
        "direction": "앨런의 시선은 손에 쥔 장총과 디스플레이 쪽을 향함. 장총의 총구는 화면 우측 하단으로 향함.",
        "built_space": "거친 돌벽과 붉은 조명 아래, 책상 위에 디스플레이가 놓인 구조. 공간적 피팅은 비교적 단순하게 처리됨.",
        "entities": "앨런의 얼굴은 유지되었으나 캐릭터 레퍼런스의 필수 복장인 비니가 누락됨. 장총은 두 손에 쥐어졌으나 대형 스코프가 누락됨. 디스플레이에 'MANUAL DUAL-ACTION SEAL' 텍스트가 확인됨.",
        "hard_violations": [
         "physically impossible anatomy (오른팔에 달린 방아쇠 쪽 손의 엄지가 카메라 쪽인 우측에 위치해 있어 사실상 두 개의 왼손을 가진 셈이 됨)"
        ],
        "physics": "양손이 장총을 받치고 있으나, 카메라와 가까운 쪽인 오른손이 총의 그립을 쥘 때 엄지가 총의 우측(카메라 방향)에 위치해 있어 물리적으로 불가능한 쥐기 자세를 연출함."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "캐릭터의 복장(비니)과 디스플레이 내용 구현은 우수하나, 방아쇠를 쥔 손의 치명적인 해부학적 오류, 장총 스코프 누락, 그리고 지시된 카메라 앵글 위반 및 불필요한 사물 추가로 인해 사용할 수 없습니다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "장총의 스코프 누락과 캐릭터의 지정된 복장(비니) 누락이 있으며, 방아쇠를 쥔 오른팔에 왼손이 달려있는 심각한 해부학적 오류로 인해 완전히 실패한 결과물입니다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "A",
        "direction": "앨런의 시선은 아래쪽에 놓인 디스플레이와 장총을 향함. 장총의 총구는 화면 좌측 하단을 향해 살짝 기울어짐.",
        "built_space": "거친 돌벽과 붉은 조명이 배경을 이루며, 디스플레이가 책상 위에 놓여 있음. 그러나 배경의 벽면과 앞쪽 책상 위에 명시되지 않은 사무용 유선 전화기가 두 개나 배치되어 공간의 일관성을 해침.",
        "entities": "앨런은 레퍼런스와 일치하는 복장(비니, 재킷)을 착용함. 장총은 두 손에 쥐어져 있으나 프랍 레퍼런스에 있는 대형 스코프가 완전히 누락됨. 디스플레이에는 시설 도면과 'MANUAL DUAL-ACTION SEAL' 텍스트가 정확히 표기됨.",
        "hard_violations": [
         "physically impossible anatomy (방아쇠를 쥔 손이 손등을 보이고 엄지가 앞을 향하는 등 정상적인 오른손이나 왼손으로 불가능한 구조임)",
         "invented objects (책상 위에 원본에 없는 90년대풍 현대식 유선 전화기가 생성됨)",
         "duplicated fitting (전화기가 벽면과 책상 두 곳에 중복으로 존재함)"
        ],
        "physics": "장총은 앨런의 양손에 의해 허공에 지지되고 있으나, 개머리판 쪽을 쥔 손의 관절과 엄지손가락 위치가 물리적/해부학적으로 불가능하게 뒤틀려 있음."
       },
       {
        "label": "B",
        "direction": "앨런의 시선은 손에 쥔 장총과 디스플레이 쪽을 향함. 장총의 총구는 화면 우측 하단으로 향함.",
        "built_space": "거친 돌벽과 붉은 조명 아래, 책상 위에 디스플레이가 놓인 구조. 공간적 피팅은 비교적 단순하게 처리됨.",
        "entities": "앨런의 얼굴은 유지되었으나 캐릭터 레퍼런스의 필수 복장인 비니가 누락됨. 장총은 두 손에 쥐어졌으나 대형 스코프가 누락됨. 디스플레이에 'MANUAL DUAL-ACTION SEAL' 텍스트가 확인됨.",
        "hard_violations": [
         "physically impossible anatomy (오른팔에 달린 방아쇠 쪽 손의 엄지가 카메라 쪽인 우측에 위치해 있어 사실상 두 개의 왼손을 가진 셈이 됨)"
        ],
        "physics": "양손이 장총을 받치고 있으나, 카메라와 가까운 쪽인 오른손이 총의 그립을 쥘 때 엄지가 총의 우측(카메라 방향)에 위치해 있어 물리적으로 불가능한 쥐기 자세를 연출함."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "총구를 토니의 가슴선에서 벗어나 하단으로 내리는 순간과 앨런의 신원·복장이 A보다 정확하지만, 화면 문구가 명확히 읽히고 장총의 조준경이 없어 완전한 충실도에는 미달한다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "두 손으로 장총을 지지하며 총구를 소폭 내린 동작은 보이나, 카메라가 앨런의 작업 측 뒤쪽보다 앞쪽에 가깝고 이전 인물의 전술 복장을 답습했으며 금지된 화면 문구도 읽힌다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "앨런의 시선은 자기 손과 장총의 작동부 쪽으로 내려가 있다. 총구는 화면 오른쪽 아래를 향하며, 화면 밖 토니의 가슴을 계속 겨누기보다는 그 선에서 아래로 조금 벗어난 것으로 읽힌다. 다만 카메라는 앨런의 작업 측 뒤쪽보다는 얼굴과 가슴 앞쪽에서 보는 각도에 가깝다.",
        "built_space": "거친 암벽 벽면, 긴 작업대, 기울어진 기록 디스플레이 1대가 보이며 지하 통제실의 붉은 환경광은 이전 스틸과 대체로 이어진다. 디스플레이는 앨런의 내려간 눈높이 아래쪽이 아니라 얼굴 뒤·옆에 놓였고, 유선 전화기와 울트라폰 디스크 및 단색 투사는 이 클로즈업에 보이지 않는다. 인물은 서서 작업대 앞에서 총을 들고 있으며 반사상은 없다.",
        "entities": "앨런으로 제시된 성인 미국계 남성 1명이 보이지만 얼굴의 수염과 인상, 하네스·로프가 달린 전술 복장은 캐릭터 레퍼런스보다 이전 스틸의 인물 복장에 더 가깝다. 장총은 검은 볼트액션 장총이고 양손에 들렸으나, 레퍼런스의 대형 조준경이 빠져 정확한 동일 소품은 아니다. 기록 화면에는 시설 도면과 영어 지시문이 보인다. 토니나 다른 사람은 등장하지 않는다.",
        "hard_violations": [
         "기록 디스플레이에 금지된 판독 가능한 영어 문구가 선명하게 노출되어 있다."
        ],
        "physics": "장총은 방아쇠 부근을 잡은 장갑 낀 손과 총열덮개 아래를 받치는 다른 손으로 지지된다. 총의 무게가 두 팔에 실려 있고 공중에 무지지 상태로 뜬 물체는 없다. 손가락과 손목은 다소 빽빽하지만 총구를 몇 센티미터 낮추는 동작으로 물리적으로 가능하다."
       },
       {
        "label": "B",
        "direction": "앨런의 시선은 장총의 리시버와 손 쪽으로 내려간다. 총구는 화면 왼쪽 아래를 향해 토니의 화면 밖 가슴선에서 확실히 벗어나 하단으로 향한다. 하강 각도가 ‘몇 센티미터’보다 다소 커 보이지만, 오른쪽 어깨 뒤·위쪽에 가까운 관찰축이라 A보다 요구된 작업 측 후방 시점에 가깝다.",
        "built_space": "거친 암벽, 붉은 저조도 조명, 작업대, 기울어진 기록 디스플레이 1대가 보인다. 작업대에는 유선 탁상전화 1대가 있고 벽에는 수화기가 달린 통신 장치 1대가 별도로 보이며, 이를 같은 전화기의 중복이라고 단정할 정도로 형태가 같지는 않다. 기록 화면은 앨런의 내려간 시선 아래쪽과 총구 하강 경로의 상부 배경에 놓인다. 울트라폰 디스크와 단색 투사는 보이지 않으며 불가능한 반사는 없다.",
        "entities": "앨런은 성인 미국계 남성 1명이며 비니, 짧은 수염, 얼굴형과 낡은 어두운 재킷이 캐릭터 레퍼런스에 A보다 가깝다. 장총은 검은 볼트액션 구조와 긴 총열·양각대를 갖췄지만 레퍼런스의 대형 조준경이 없어 정확한 동일 소품은 아니다. 두 손 모두 앨런의 손으로 보이고 다른 인물은 없다. 기록 디스플레이에는 시설 도면과 영어 지시문이 표시된다.",
        "hard_violations": [
         "기록 디스플레이에 금지된 판독 가능한 영어 문구가 선명하게 노출되어 있다."
        ],
        "physics": "앨런의 뒤쪽 손이 권총손잡이와 방아쇠울 부근을 쥐고 앞손이 총열덮개 아래를 받쳐 장총 전체를 지지한다. 양각대는 접힌 채 아래로 매달려 있으며 총을 떠받치는 주 지지점은 두 손이다. 총구가 내려가는 자세는 팔꿈치와 손목의 회전으로 만들 수 있고, 지지 없이 떠 있는 신체나 물체는 없다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "총구를 토니의 가슴선에서 벗어나 하단으로 내리는 순간과 앨런의 신원·복장이 A보다 정확하지만, 화면 문구가 명확히 읽히고 장총의 조준경이 없어 완전한 충실도에는 미달한다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "두 손으로 장총을 지지하며 총구를 소폭 내린 동작은 보이나, 카메라가 앨런의 작업 측 뒤쪽보다 앞쪽에 가깝고 이전 인물의 전술 복장을 답습했으며 금지된 화면 문구도 읽힌다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "앨런의 시선은 자기 손과 장총의 작동부 쪽으로 내려가 있다. 총구는 화면 오른쪽 아래를 향하며, 화면 밖 토니의 가슴을 계속 겨누기보다는 그 선에서 아래로 조금 벗어난 것으로 읽힌다. 다만 카메라는 앨런의 작업 측 뒤쪽보다는 얼굴과 가슴 앞쪽에서 보는 각도에 가깝다.",
        "built_space": "거친 암벽 벽면, 긴 작업대, 기울어진 기록 디스플레이 1대가 보이며 지하 통제실의 붉은 환경광은 이전 스틸과 대체로 이어진다. 디스플레이는 앨런의 내려간 눈높이 아래쪽이 아니라 얼굴 뒤·옆에 놓였고, 유선 전화기와 울트라폰 디스크 및 단색 투사는 이 클로즈업에 보이지 않는다. 인물은 서서 작업대 앞에서 총을 들고 있으며 반사상은 없다.",
        "entities": "앨런으로 제시된 성인 미국계 남성 1명이 보이지만 얼굴의 수염과 인상, 하네스·로프가 달린 전술 복장은 캐릭터 레퍼런스보다 이전 스틸의 인물 복장에 더 가깝다. 장총은 검은 볼트액션 장총이고 양손에 들렸으나, 레퍼런스의 대형 조준경이 빠져 정확한 동일 소품은 아니다. 기록 화면에는 시설 도면과 영어 지시문이 보인다. 토니나 다른 사람은 등장하지 않는다.",
        "hard_violations": [
         "기록 디스플레이에 금지된 판독 가능한 영어 문구가 선명하게 노출되어 있다."
        ],
        "physics": "장총은 방아쇠 부근을 잡은 장갑 낀 손과 총열덮개 아래를 받치는 다른 손으로 지지된다. 총의 무게가 두 팔에 실려 있고 공중에 무지지 상태로 뜬 물체는 없다. 손가락과 손목은 다소 빽빽하지만 총구를 몇 센티미터 낮추는 동작으로 물리적으로 가능하다."
       },
       {
        "label": "A",
        "direction": "앨런의 시선은 장총의 리시버와 손 쪽으로 내려간다. 총구는 화면 왼쪽 아래를 향해 토니의 화면 밖 가슴선에서 확실히 벗어나 하단으로 향한다. 하강 각도가 ‘몇 센티미터’보다 다소 커 보이지만, 오른쪽 어깨 뒤·위쪽에 가까운 관찰축이라 A보다 요구된 작업 측 후방 시점에 가깝다.",
        "built_space": "거친 암벽, 붉은 저조도 조명, 작업대, 기울어진 기록 디스플레이 1대가 보인다. 작업대에는 유선 탁상전화 1대가 있고 벽에는 수화기가 달린 통신 장치 1대가 별도로 보이며, 이를 같은 전화기의 중복이라고 단정할 정도로 형태가 같지는 않다. 기록 화면은 앨런의 내려간 시선 아래쪽과 총구 하강 경로의 상부 배경에 놓인다. 울트라폰 디스크와 단색 투사는 보이지 않으며 불가능한 반사는 없다.",
        "entities": "앨런은 성인 미국계 남성 1명이며 비니, 짧은 수염, 얼굴형과 낡은 어두운 재킷이 캐릭터 레퍼런스에 A보다 가깝다. 장총은 검은 볼트액션 구조와 긴 총열·양각대를 갖췄지만 레퍼런스의 대형 조준경이 없어 정확한 동일 소품은 아니다. 두 손 모두 앨런의 손으로 보이고 다른 인물은 없다. 기록 디스플레이에는 시설 도면과 영어 지시문이 표시된다.",
        "hard_violations": [
         "기록 디스플레이에 금지된 판독 가능한 영어 문구가 선명하게 노출되어 있다."
        ],
        "physics": "앨런의 뒤쪽 손이 권총손잡이와 방아쇠울 부근을 쥐고 앞손이 총열덮개 아래를 받쳐 장총 전체를 지지한다. 양각대는 접힌 채 아래로 매달려 있으며 총을 떠받치는 주 지지점은 두 손이다. 총구가 내려가는 자세는 팔꿈치와 손목의 회전으로 만들 수 있고, 지지 없이 떠 있는 신체나 물체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.333
   },
   "adjusted": {
    "A": 1.75,
    "B": 1.083
   },
   "violations": {
    "A": [
     "[gemini-pro] physically impossible anatomy (방아쇠를 쥔 손이 손등을 보이고 엄지가 앞을 향하는 등 정상적인 오른손이나 왼손으로 불가능한 구조임)",
     "[gemini-pro] invented objects (책상 위에 원본에 없는 90년대풍 현대식 유선 전화기가 생성됨)",
     "[gemini-pro] duplicated fitting (전화기가 벽면과 책상 두 곳에 중복으로 존재함)",
     "[gpt] 기록 디스플레이에 금지된 판독 가능한 영어 문구가 선명하게 노출되어 있다."
    ],
    "B": [
     "[gemini-pro] physically impossible anatomy (오른팔에 달린 방아쇠 쪽 손의 엄지가 카메라 쪽인 우측에 위치해 있어 사실상 두 개의 왼손을 가진 셈이 됨)",
     "[gpt] 기록 디스플레이에 금지된 판독 가능한 영어 문구가 선명하게 노출되어 있다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 1750,
   "B": 1083
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1750,
    "verdict_ko": "캐릭터의 복장(비니)과 디스플레이 내용 구현은 우수하나, 방아쇠를 쥔 손의 치명적인 해부학적 오류, 장총 스코프 누락, 그리고 지시된 카메라 앵글 위반 및 불필요한 사물 추가로 인해 사용할 수 없습니다.  ★위반: [gemini-pro] physically impossible anatomy (방아쇠를 쥔 손이 손등을 보이고 엄지가 앞을 향하는 등 정상적인 오른손이나 왼손으로 불가능한 구조임) / [gemini-pro] invented objects (책상 위에 원본에 없는 90년대풍 현대식 유선 전화기가 생성됨) / [gemini-pro] duplicated fitting (전화기가 벽면과 책상 두 곳에 중복으로 존재함) / [gpt] 기록 디스플레이에 금지된 판독 가능한 영어 문구가 선명하게 노출되어 있다."
   },
   {
    "label": "B",
    "score": 1083,
    "verdict_ko": "장총의 스코프 누락과 캐릭터의 지정된 복장(비니) 누락이 있으며, 방아쇠를 쥔 오른팔에 왼손이 달려있는 심각한 해부학적 오류로 인해 완전히 실패한 결과물입니다.  ★위반: [gemini-pro] physically impossible anatomy (오른팔에 달린 방아쇠 쪽 손의 엄지가 카메라 쪽인 우측에 위치해 있어 사실상 두 개의 왼손을 가진 셈이 됨) / [gpt] 기록 디스플레이에 금지된 판독 가능한 영어 문구가 선명하게 노출되어 있다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/37154490-a4ed-4b37-9092-489c054d8a72/images/cc1d3d0b-626e-4f20-bc33-f5f9e2c6ee7e/scene/recipe/S3sh12_sel.png",
    "asset_id": "ed9f6f98-fd63-4e13-a7dd-1dec0a98444b",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 앨런: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1001252>",
    "asset_id": "b426ff26-eb7d-4b66-a4a7-de7c648b9df3",
    "role": "character_ref"
   },
   {
    "label": "PROP REFERENCE — 장총: the exact object appearing in this shot; match its look, material and wear exactly.",
    "path": "<bytes:785972>",
    "asset_id": "3c9dd37a-da89-4fa6-a8b0-83c72e84ff1f",
    "role": "prop_ref"
   }
  ],
  "critique_skipped": true,
  "needs_reshoot": true,
  "shot_run_produced": true,
  "shot_run_uid": "06a9bdad-3a59-78bf-8792-1d648b10484a",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S3sh12"
  }
 },
 "S3sh13::cine": {
  "applied": true,
  "attempted_at": "2026-09-05T09:05:13.955066+00:00",
  "fingerprint": "3b583c777aa18dc7654f4b75c394772c55d532a510395c8da7ecdb69aac1d66b",
  "fingerprint_version": 2,
  "provider": "grok",
  "endpoint": "openrouter/chat-completions",
  "model": "x-ai/grok-imagine-image-2.0",
  "pack": "24.202608252115",
  "source_file": "S3sh13_sel.png",
  "source_sha256": "8e277b8900b9e47d6464e61b50876a31c65934b6f9622b0a2b20853dbe0f1905",
  "file": "S3sh13_cine.png",
  "staged_sha256": "6b6511dc8a8e5c2ec13974b4ffc497f58cdfb6d031bce3bad1b2b6b5272d004a",
  "latency_ms": 16942
 },
 "S4sh7::signage": {
  "fp": "fa567a20290ca7ad",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S4sh7": {
  "input_fingerprint": "2eb36348b2674b29",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — Wyoming, United States, 2419; all people are English-speaking Americans unless stated. TIME OF DAY (lock): twilight.\n\nSHOT TEXT (authoritative, Korean): 원판 위로 푸른빛을 발하는 단색 홀로그램 하트의 형상\n\nLOCATION (lock): Inside the mountain contact post's underground control room, over the active communication disc. A blue-tinted monochrome holographic figure glows above the disc. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: ultraphone disc in the lower-center of the frame, foreground; 하트 (홀로그램 투사) in the upper-center of the frame, midground.\n- KEY BACKGROUND ELEMENTS: ultraphone disc (on and projecting Hart's hologram) — The upper face is visible along the lower frame, receding diagonally beneath Hart's projected figure; used as Lower-frame projection anchor and scale reference for the hologram.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Hart's explicitly blue holographic emission supplies the restrained cool accent within an otherwise muted, low-key composition.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The isolated phone remains wired to the contact station and shows “EXTERNAL POWER — UNSTABLE,” with its certificate lineage displayed externally. The ultraphone disc continues to cast blue-toned monochrome holographic light.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — Wyoming, United States, 2419; all people are English-speaking Americans unless stated. TIME OF DAY (lock): twilight.\n\nSHOT TEXT (authoritative, Korean): 원판 위로 푸른빛을 발하는 단색 홀로그램 하트의 형상\n\nLOCATION (lock): Inside the mountain contact post's underground control room, over the active communication disc. A blue-tinted monochrome holographic figure glows above the disc. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: ultraphone disc in the lower-center of the frame, foreground; 하트 (홀로그램 투사) in the upper-center of the frame, midground.\n- KEY BACKGROUND ELEMENTS: ultraphone disc (on and projecting Hart's hologram) — The upper face is visible along the lower frame, receding diagonally beneath Hart's projected figure; used as Lower-frame projection anchor and scale reference for the hologram.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Hart's explicitly blue holographic emission supplies the restrained cool accent within an otherwise muted, low-key composition.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The isolated phone remains wired to the contact station and shows “EXTERNAL POWER — UNSTABLE,” with its certificate lineage displayed externally. The ultraphone disc continues to cast blue-toned monochrome holographic light.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — Wyoming, United States, 2419; all people are English-speaking Americans unless stated. TIME OF DAY (lock): twilight.\n\nSHOT TEXT (authoritative, Korean): 원판 위로 푸른빛을 발하는 단색 홀로그램 하트의 형상\n\nLOCATION (lock): Inside the mountain contact post's underground control room, over the active communication disc. A blue-tinted monochrome holographic figure glows above the disc. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: ultraphone disc in the lower-center of the frame, foreground; 하트 (홀로그램 투사) in the upper-center of the frame, midground.\n- KEY BACKGROUND ELEMENTS: ultraphone disc (on and projecting Hart's hologram) — The upper face is visible along the lower frame, receding diagonally beneath Hart's projected figure; used as Lower-frame projection anchor and scale reference for the hologram.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Hart's explicitly blue holographic emission supplies the restrained cool accent within an otherwise muted, low-key composition.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established about the place — its fixed features, wear and lighting — persists. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The isolated phone remains wired to the contact station and shows “EXTERNAL POWER — UNSTABLE,” with its certificate lineage displayed externally. The ultraphone disc continues to cast blue-toned monochrome holographic light.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "원판의 중심부에서 위쪽으로 푸른빛이 투사되어 홀로그램 하트를 형성하고 있음.",
    "built_space": "레퍼런스에 나타난 지하 통제실의 고정된 요소(좌측의 빛나는 링이 있는 문, 후경의 지도와 모니터가 있는 책상, 우측의 투명한 전시 상자 등)가 정확히 동일한 위치와 형태로 유지됨. 공간의 구조와 카메라 시점이 일치함.",
    "entities": "푸른색의 단색 홀로그램 하트가 원판 위 공중에 떠 있음. 프롬프트의 지시대로 사람은 아무도 존재하지 않음.",
    "hard_violations": [],
    "physics": "홀로그램 하트는 기기에서 쏘아 올린 빛의 투사에 의해 논리적으로 공중에 떠 있음."
   },
   {
    "label": "B",
    "direction": "원판 위에 놓인 사각형 형태의 장치에서 위쪽으로 빛이 뻗어 나가 홀로그램 하트를 비추고 있음.",
    "built_space": "레퍼런스의 공간 구조가 완전히 변경됨. 좌측에 있던 돌벽이 우측으로 이동했고, 지하 통제실임에도 불구하고 후경 상단에 바깥빛이 들어오는 창문이 추가됨. 책상과 모니터의 배치도 완전히 달라짐.",
    "entities": "푸른색 와이어프레임 형태의 홀로그램 하트가 원판 위에 생성됨. 사람은 존재하지 않음.",
    "hard_violations": [
     "[gemini-pro] LOCATION 지정 위반: 레퍼런스 이미지의 공간(지하 통제실) 구조 및 고정 시설물의 배치를 무시하고 완전히 다른 실내로 렌더링함.",
     "[gemini-pro] 물리적/논리적 불가능: '지하(underground)' 공간임에도 바깥이 보이는 창문을 추가함.",
     "[gpt] 잠긴 동일 장소에 존재하지 않던 밝은 창문형 개구부 2개를 추가해 지하 통제실의 고정 건축 구조를 변경했다."
    ],
    "physics": "홀로그램 하트는 투사 장치에서 뿜어져 나오는 빛에 의해 시각적으로 지탱되며 떠 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "레퍼런스 이미지의 지하 통제실 배경(좌측 문, 후경의 모니터와 지도, 우측 투명 상자 등)을 완벽하게 재현했으며, 지시된 대로 인물 없이 홀로그램 하트를 훌륭하게 연출했습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "홀로그램 하트는 잘 표현되었으나, 지시된 레퍼런스의 공간 구조를 무시하고 지하 통제실에 창문을 추가하거나 구조물의 위치를 임의로 변경하여 Location Lock을 심각하게 위반했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "원판의 중심부에서 위쪽으로 푸른빛이 투사되어 홀로그램 하트를 형성하고 있음.",
        "built_space": "레퍼런스에 나타난 지하 통제실의 고정된 요소(좌측의 빛나는 링이 있는 문, 후경의 지도와 모니터가 있는 책상, 우측의 투명한 전시 상자 등)가 정확히 동일한 위치와 형태로 유지됨. 공간의 구조와 카메라 시점이 일치함.",
        "entities": "푸른색의 단색 홀로그램 하트가 원판 위 공중에 떠 있음. 프롬프트의 지시대로 사람은 아무도 존재하지 않음.",
        "hard_violations": [],
        "physics": "홀로그램 하트는 기기에서 쏘아 올린 빛의 투사에 의해 논리적으로 공중에 떠 있음."
       },
       {
        "label": "B",
        "direction": "원판 위에 놓인 사각형 형태의 장치에서 위쪽으로 빛이 뻗어 나가 홀로그램 하트를 비추고 있음.",
        "built_space": "레퍼런스의 공간 구조가 완전히 변경됨. 좌측에 있던 돌벽이 우측으로 이동했고, 지하 통제실임에도 불구하고 후경 상단에 바깥빛이 들어오는 창문이 추가됨. 책상과 모니터의 배치도 완전히 달라짐.",
        "entities": "푸른색 와이어프레임 형태의 홀로그램 하트가 원판 위에 생성됨. 사람은 존재하지 않음.",
        "hard_violations": [
         "LOCATION 지정 위반: 레퍼런스 이미지의 공간(지하 통제실) 구조 및 고정 시설물의 배치를 무시하고 완전히 다른 실내로 렌더링함.",
         "물리적/논리적 불가능: '지하(underground)' 공간임에도 바깥이 보이는 창문을 추가함."
        ],
        "physics": "홀로그램 하트는 투사 장치에서 뿜어져 나오는 빛에 의해 시각적으로 지탱되며 떠 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "레퍼런스 이미지의 지하 통제실 배경(좌측 문, 후경의 모니터와 지도, 우측 투명 상자 등)을 완벽하게 재현했으며, 지시된 대로 인물 없이 홀로그램 하트를 훌륭하게 연출했습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "홀로그램 하트는 잘 표현되었으나, 지시된 레퍼런스의 공간 구조를 무시하고 지하 통제실에 창문을 추가하거나 구조물의 위치를 임의로 변경하여 Location Lock을 심각하게 위반했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "원판의 중심부에서 위쪽으로 푸른빛이 투사되어 홀로그램 하트를 형성하고 있음.",
        "built_space": "레퍼런스에 나타난 지하 통제실의 고정된 요소(좌측의 빛나는 링이 있는 문, 후경의 지도와 모니터가 있는 책상, 우측의 투명한 전시 상자 등)가 정확히 동일한 위치와 형태로 유지됨. 공간의 구조와 카메라 시점이 일치함.",
        "entities": "푸른색의 단색 홀로그램 하트가 원판 위 공중에 떠 있음. 프롬프트의 지시대로 사람은 아무도 존재하지 않음.",
        "hard_violations": [],
        "physics": "홀로그램 하트는 기기에서 쏘아 올린 빛의 투사에 의해 논리적으로 공중에 떠 있음."
       },
       {
        "label": "B",
        "direction": "원판 위에 놓인 사각형 형태의 장치에서 위쪽으로 빛이 뻗어 나가 홀로그램 하트를 비추고 있음.",
        "built_space": "레퍼런스의 공간 구조가 완전히 변경됨. 좌측에 있던 돌벽이 우측으로 이동했고, 지하 통제실임에도 불구하고 후경 상단에 바깥빛이 들어오는 창문이 추가됨. 책상과 모니터의 배치도 완전히 달라짐.",
        "entities": "푸른색 와이어프레임 형태의 홀로그램 하트가 원판 위에 생성됨. 사람은 존재하지 않음.",
        "hard_violations": [
         "LOCATION 지정 위반: 레퍼런스 이미지의 공간(지하 통제실) 구조 및 고정 시설물의 배치를 무시하고 완전히 다른 실내로 렌더링함.",
         "물리적/논리적 불가능: '지하(underground)' 공간임에도 바깥이 보이는 창문을 추가함."
        ],
        "physics": "홀로그램 하트는 투사 장치에서 뿜어져 나오는 빛에 의해 시각적으로 지탱되며 떠 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "원판을 하단 중앙에 두고 그 중심에서 투사된 푸른 단색 하트 홀로그램을 상단 중앙에 정확히 배치했으며, 이전 숏의 지하 통제실 구조와 무인 상태도 가장 충실히 유지했다."
       },
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "하트와 원판의 배치는 맞지만 이전 숏에 없던 밝은 창문 두 개로 잠긴 지하 통제실의 고정 구조를 바꾸고 격리 전화도 명확히 재현하지 못했다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "시선이나 무기, 이동하는 몸은 없다. 하트의 아래쪽 꼭짓점은 원판 중앙의 사각 투사기를 향하며, 푸른 투사광도 그 투사기에서 하트 쪽으로 위로 뻗어 서로 정렬된다.",
        "built_space": "하단 중앙에 원형 울트라폰 원판 1개와 그 위의 사각 투사기 1개가 있다. 좌측에는 작업대와 기기·케이블, 양쪽에는 장비 랙이 보이지만, 후면 상부에 밝은 창문 형태의 개구부 2개가 새로 생겼다. 이전 숏의 좌측 금속문과 우측 투명 보호함 중심의 고정 배치는 사라지거나 크게 바뀌었다. 광학적으로 불가능한 반사는 보이지 않는다.",
        "entities": "푸른 와이어프레임 하트 홀로그램 1개와 작동 중인 원판 1개가 있으며 사람·얼굴·신체는 없다. 하트는 파란색 계열의 단색 투사 형상으로 읽힌다. 좌측 화면형 장치는 보이지만 격리 전화인지, 외부 인증 계보를 가진 유선 전화인지 명확하지 않으며 문자는 판독되지 않는다.",
        "hard_violations": [
         "잠긴 동일 장소에 존재하지 않던 밝은 창문형 개구부 2개를 추가해 지하 통제실의 고정 건축 구조를 변경했다."
        ],
        "physics": "하트는 물체처럼 무지지 부유하는 것이 아니라 원판 중앙 투사기에서 올라오는 가시적인 광선으로 형성된 홀로그램이다. 원판은 바닥에 놓인 원통형 설비로 안정적으로 지지되며, 떠 있는 사람이나 손에 잡혀야 할 물체는 없다."
       },
       {
        "label": "B",
        "direction": "시선이나 무기, 이동하는 몸은 없다. 하트의 아래쪽 꼭짓점이 원판의 발광 중심을 정확히 가리키고, 투사광은 원판 중심에서 하트로 수직 상승해 명시된 투사 관계가 성립한다.",
        "built_space": "하단 중앙 전경에 원형 울트라폰 원판 1개가 있고 윗면이 보인 채 뒤쪽으로 물러난다. 좌측에는 거친 암벽과 금속문 1개, 후면에는 통제 콘솔과 의자·장비 랙, 우측에는 긴 작업대와 투명 보호함 1개가 있어 이전 숏의 공간 구성을 유지한다. 사람은 없다. 투명함과 금속 표면의 반사는 주변 조명과 카메라 위치에서 가능한 범위다.",
        "entities": "푸른 단색 계열의 하트 홀로그램 1개와 활성 울트라폰 원판 1개가 명확하다. 살아 있는 사람·얼굴·의복은 전혀 없다. 우측 보호함 안에 격리 전화로 보이는 소형 단말이 남아 있고 배선도 주변 설비와 연결되어 있으나, 요구된 경고문과 인증 계보는 판독 불가능하게 처리되어 무문자 조건을 지킨다. 원판 중앙 투사부는 이전 숏의 사각 받침 대신 원형 개구로 단순화되었다.",
        "hard_violations": [],
        "physics": "하트는 원판 중앙의 밝은 방출구와 연결된 투사광에 의해 형성되어 홀로그램으로서 자연스럽게 떠 있다. 원판은 바닥에 놓인 설비로 지지되고, 전화형 단말은 보호함 내부 받침면 위에 놓여 있으며, 지지 없이 떠 있는 신체나 물체는 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "원판을 하단 중앙에 두고 그 중심에서 투사된 푸른 단색 하트 홀로그램을 상단 중앙에 정확히 배치했으며, 이전 숏의 지하 통제실 구조와 무인 상태도 가장 충실히 유지했다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "하트와 원판의 배치는 맞지만 이전 숏에 없던 밝은 창문 두 개로 잠긴 지하 통제실의 고정 구조를 바꾸고 격리 전화도 명확히 재현하지 못했다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "시선이나 무기, 이동하는 몸은 없다. 하트의 아래쪽 꼭짓점은 원판 중앙의 사각 투사기를 향하며, 푸른 투사광도 그 투사기에서 하트 쪽으로 위로 뻗어 서로 정렬된다.",
        "built_space": "하단 중앙에 원형 울트라폰 원판 1개와 그 위의 사각 투사기 1개가 있다. 좌측에는 작업대와 기기·케이블, 양쪽에는 장비 랙이 보이지만, 후면 상부에 밝은 창문 형태의 개구부 2개가 새로 생겼다. 이전 숏의 좌측 금속문과 우측 투명 보호함 중심의 고정 배치는 사라지거나 크게 바뀌었다. 광학적으로 불가능한 반사는 보이지 않는다.",
        "entities": "푸른 와이어프레임 하트 홀로그램 1개와 작동 중인 원판 1개가 있으며 사람·얼굴·신체는 없다. 하트는 파란색 계열의 단색 투사 형상으로 읽힌다. 좌측 화면형 장치는 보이지만 격리 전화인지, 외부 인증 계보를 가진 유선 전화인지 명확하지 않으며 문자는 판독되지 않는다.",
        "hard_violations": [
         "잠긴 동일 장소에 존재하지 않던 밝은 창문형 개구부 2개를 추가해 지하 통제실의 고정 건축 구조를 변경했다."
        ],
        "physics": "하트는 물체처럼 무지지 부유하는 것이 아니라 원판 중앙 투사기에서 올라오는 가시적인 광선으로 형성된 홀로그램이다. 원판은 바닥에 놓인 원통형 설비로 안정적으로 지지되며, 떠 있는 사람이나 손에 잡혀야 할 물체는 없다."
       },
       {
        "label": "A",
        "direction": "시선이나 무기, 이동하는 몸은 없다. 하트의 아래쪽 꼭짓점이 원판의 발광 중심을 정확히 가리키고, 투사광은 원판 중심에서 하트로 수직 상승해 명시된 투사 관계가 성립한다.",
        "built_space": "하단 중앙 전경에 원형 울트라폰 원판 1개가 있고 윗면이 보인 채 뒤쪽으로 물러난다. 좌측에는 거친 암벽과 금속문 1개, 후면에는 통제 콘솔과 의자·장비 랙, 우측에는 긴 작업대와 투명 보호함 1개가 있어 이전 숏의 공간 구성을 유지한다. 사람은 없다. 투명함과 금속 표면의 반사는 주변 조명과 카메라 위치에서 가능한 범위다.",
        "entities": "푸른 단색 계열의 하트 홀로그램 1개와 활성 울트라폰 원판 1개가 명확하다. 살아 있는 사람·얼굴·의복은 전혀 없다. 우측 보호함 안에 격리 전화로 보이는 소형 단말이 남아 있고 배선도 주변 설비와 연결되어 있으나, 요구된 경고문과 인증 계보는 판독 불가능하게 처리되어 무문자 조건을 지킨다. 원판 중앙 투사부는 이전 숏의 사각 받침 대신 원형 개구로 단순화되었다.",
        "hard_violations": [],
        "physics": "하트는 원판 중앙의 밝은 방출구와 연결된 투사광에 의해 형성되어 홀로그램으로서 자연스럽게 떠 있다. 원판은 바닥에 놓인 설비로 지지되고, 전화형 단말은 보호함 내부 받침면 위에 놓여 있으며, 지지 없이 떠 있는 신체나 물체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.0
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.75
   },
   "violations": {
    "B": [
     "[gemini-pro] LOCATION 지정 위반: 레퍼런스 이미지의 공간(지하 통제실) 구조 및 고정 시설물의 배치를 무시하고 완전히 다른 실내로 렌더링함.",
     "[gemini-pro] 물리적/논리적 불가능: '지하(underground)' 공간임에도 바깥이 보이는 창문을 추가함.",
     "[gpt] 잠긴 동일 장소에 존재하지 않던 밝은 창문형 개구부 2개를 추가해 지하 통제실의 고정 건축 구조를 변경했다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 750
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "레퍼런스 이미지의 지하 통제실 배경(좌측 문, 후경의 모니터와 지도, 우측 투명 상자 등)을 완벽하게 재현했으며, 지시된 대로 인물 없이 홀로그램 하트를 훌륭하게 연출했습니다."
   },
   {
    "label": "B",
    "score": 750,
    "verdict_ko": "홀로그램 하트는 잘 표현되었으나, 지시된 레퍼런스의 공간 구조를 무시하고 지하 통제실에 창문을 추가하거나 구조물의 위치를 임의로 변경하여 Location Lock을 심각하게 위반했습니다.  ★위반: [gemini-pro] LOCATION 지정 위반: 레퍼런스 이미지의 공간(지하 통제실) 구조 및 고정 시설물의 배치를 무시하고 완전히 다른 실내로 렌더링함. / [gemini-pro] 물리적/논리적 불가능: '지하(underground)' 공간임에도 바깥이 보이는 창문을 추가함. / [gpt] 잠긴 동일 장소에 존재하지 않던 밝은 창문형 개구부 2개를 추가해 지하 통제실의 고정 건축 구조를 변경했다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features and lighting mood are LOCKED to this photo; never copy its camera framing. The people visible in that photo are NOT in this shot: never carry their faces, bodies or clothing onto anyone here. Each person in THIS shot is defined solely by the PEOPLE list, their own CHARACTER REFERENCE images and their wardrobe notes. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/37154490-a4ed-4b37-9092-489c054d8a72/images/cc1d3d0b-626e-4f20-bc33-f5f9e2c6ee7e/scene/recipe/S3sh8_sel.png",
    "asset_id": "c78987c9-ea46-4c17-9171-403ed623c2aa",
    "role": "prev_still"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": true,
  "shot_run_uid": "06a9bdb5-b04c-774b-b5aa-cd240d970a7c",
  "ref_mode": "prev만 (배경 전용·공유 계획)",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S3sh8"
  },
  "lane_policy": "share_plan_prev_bgonly"
 },
 "S4sh7::cine": {
  "applied": true,
  "attempted_at": "2026-09-05T09:06:25.560257+00:00",
  "fingerprint": "78050124a86c2a6a2dcadc74091481910481e745366f3c2ea730b8ac75d0f4ac",
  "fingerprint_version": 2,
  "provider": "grok",
  "endpoint": "openrouter/chat-completions",
  "model": "x-ai/grok-imagine-image-2.0",
  "pack": "24.202608252115",
  "source_file": "S4sh7_sel.png",
  "source_sha256": "f0119a69ee97e38149af8114db5bfcb9ae22516f9fd98771c88fa82ad1bc5cea",
  "file": "S4sh7_cine.png",
  "staged_sha256": "fdfffa114a21ad283f5717db526f41238a08eb074704217ab1feacf7da1daa0f",
  "latency_ms": 10110
 },
 "S4sh10::signage": {
  "fp": "e7572c759ad18b0e",
  "inscriptions": [],
  "cues": [
   {
    "text_native": "",
    "source": "scene_text_implied",
    "source_quote": "확대된 텍스트"
   }
  ],
  "dropped": []
 },
 "S4sh10::bgfirst_bg": {
  "input_fingerprint": "2fefb363e1a108dc",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 확대된 텍스트를 가리킨 채 앨런을 쳐다보는 토니의 상체\n\nLOCATION (lock): Inside the mountain contact post's underground control room, at the certificate display beside the wired phone workstation. Enlarged text on the screen lights the speaker's upper body.\n\nTIME OF DAY (lock): twilight.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Monitor with enlarged authority text (Displaying ROOT RECOVERY CLASS: AR-12 in enlarged form) — The content-bearing face is oblique but readable to camera, with the enlarged authority line visible behind Tony's pointing hand; used as Connects Tony's extended finger to the evidence while remaining secondary to his expression.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Low-key illumination appropriate to the twilight interior preserves a cool, muted palette and restrained contrast around the readable monitor text.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 확대된 텍스트를 가리킨 채 앨런을 쳐다보는 토니의 상체\n\nLOCATION (lock): Inside the mountain contact post's underground control room, at the certificate display beside the wired phone workstation. Enlarged text on the screen lights the speaker's upper body.\n\nTIME OF DAY (lock): twilight.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Monitor with enlarged authority text (Displaying ROOT RECOVERY CLASS: AR-12 in enlarged form) — The content-bearing face is oblique but readable to camera, with the enlarged authority line visible behind Tony's pointing hand; used as Connects Tony's extended finger to the evidence while remaining secondary to his expression.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Low-key illumination appropriate to the twilight interior preserves a cool, muted palette and restrained contrast around the readable monitor text.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/37154490-a4ed-4b37-9092-489c054d8a72/images/cc1d3d0b-626e-4f20-bc33-f5f9e2c6ee7e/scene/recipe/S4sh10__bgfirst_bg.png",
  "asset_id": "e8939552-c2bd-4281-8182-5b67265ca03c",
  "input_asset_ids": [
   "c6bc4fcb-ca37-49cc-9352-ba77eac59b11",
   "eb7f72c2-429d-4eb2-880a-e3c8b93c34f5"
  ]
 },
 "S4sh10": {
  "input_fingerprint": "0bd66b4f937dd2aa",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — Wyoming, United States, 2419; all people are English-speaking Americans unless stated. TIME OF DAY (lock): twilight.\n\nSHOT TEXT (authoritative, Korean): 확대된 텍스트를 가리킨 채 앨런을 쳐다보는 토니의 상체\n\nLOCATION (lock): Inside the mountain contact post's underground control room, at the certificate display beside the wired phone workstation. Enlarged text on the screen lights the speaker's upper body. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Monitor with enlarged authority text (Displaying ROOT RECOVERY CLASS: AR-12 in enlarged form) — The content-bearing face is oblique but readable to camera, with the enlarged authority line visible behind Tony's pointing hand; used as Connects Tony's extended finger to the evidence while remaining secondary to his expression.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Low-key illumination appropriate to the twilight interior preserves a cool, muted palette and restrained contrast around the readable monitor text.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The phone remains on unstable external power inside the wired isolation setup. The certificate display is enlarged to show “ROOT RECOVERY CLASS: AR-12,” and the ultraphone projections remain active. 토니(앤서니 로저스): He continues holding the broken employee ID badge and indicates the enlarged authority text.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 토니(앤서니 로저스) (미국인 남성, 30대 중반, 자연스러운 성인 남성 얼굴, 구체적으로 명시되지 않은 자연색 머리카락) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — Wyoming, United States, 2419; all people are English-speaking Americans unless stated. TIME OF DAY (lock): twilight.\n\nSHOT TEXT (authoritative, Korean): 확대된 텍스트를 가리킨 채 앨런을 쳐다보는 토니의 상체\n\nLOCATION (lock): Inside the mountain contact post's underground control room, at the certificate display beside the wired phone workstation. Enlarged text on the screen lights the speaker's upper body. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Monitor with enlarged authority text (Displaying ROOT RECOVERY CLASS: AR-12 in enlarged form) — The content-bearing face is oblique but readable to camera, with the enlarged authority line visible behind Tony's pointing hand; used as Connects Tony's extended finger to the evidence while remaining secondary to his expression.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Low-key illumination appropriate to the twilight interior preserves a cool, muted palette and restrained contrast around the readable monitor text.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The phone remains on unstable external power inside the wired isolation setup. The certificate display is enlarged to show “ROOT RECOVERY CLASS: AR-12,” and the ultraphone projections remain active. 토니(앤서니 로저스): He continues holding the broken employee ID badge and indicates the enlarged authority text.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 토니(앤서니 로저스) (미국인 남성, 30대 중반, 자연스러운 성인 남성 얼굴, 구체적으로 명시되지 않은 자연색 머리카락) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — Wyoming, United States, 2419; all people are English-speaking Americans unless stated. TIME OF DAY (lock): twilight.\n\nSHOT TEXT (authoritative, Korean): 확대된 텍스트를 가리킨 채 앨런을 쳐다보는 토니의 상체\n\nLOCATION (lock): Inside the mountain contact post's underground control room, at the certificate display beside the wired phone workstation. Enlarged text on the screen lights the speaker's upper body. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Monitor with enlarged authority text (Displaying ROOT RECOVERY CLASS: AR-12 in enlarged form) — The content-bearing face is oblique but readable to camera, with the enlarged authority line visible behind Tony's pointing hand; used as Connects Tony's extended finger to the evidence while remaining secondary to his expression.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Low-key illumination appropriate to the twilight interior preserves a cool, muted palette and restrained contrast around the readable monitor text.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The phone remains on unstable external power inside the wired isolation setup. The certificate display is enlarged to show “ROOT RECOVERY CLASS: AR-12,” and the ultraphone projections remain active. 토니(앤서니 로저스): He continues holding the broken employee ID badge and indicates the enlarged authority text.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 토니(앤서니 로저스) (미국인 남성, 30대 중반, 자연스러운 성인 남성 얼굴, 구체적으로 명시되지 않은 자연색 머리카락) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/37154490-a4ed-4b37-9092-489c054d8a72/images/cc1d3d0b-626e-4f20-bc33-f5f9e2c6ee7e/scene/recipe/S4sh10__bgfirst_bg.png",
     "asset_id": "e8939552-c2bd-4281-8182-5b67265ca03c",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/37154490-a4ed-4b37-9092-489c054d8a72/images/cc1d3d0b-626e-4f20-bc33-f5f9e2c6ee7e/conti/conti_S4sh10.png",
     "asset_id": "c6bc4fcb-ca37-49cc-9352-ba77eac59b11",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 토니(앤서니 로저스): the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:929851>",
     "asset_id": "1d3c6fe6-553b-4c00-b68f-025be6d786c1",
     "role": "character_ref"
    },
    {
     "label": "PROP REFERENCE — 균열된 2020년대 스마트폰: the exact object appearing in this shot; match its look, material and wear exactly.",
     "path": "<bytes:992213>",
     "asset_id": "0f8d1056-90bf-4776-ab97-2224cc50fa6e",
     "role": "prop_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/37154490-a4ed-4b37-9092-489c054d8a72/episodes/cc1d3d0b-626e-4f20-bc33-f5f9e2c6ee7e/images/background_chain/L11B02.png",
     "asset_id": "eb7f72c2-429d-4eb2-880a-e3c8b93c34f5",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 토니(앤서니 로저스): the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:929851>",
     "asset_id": "1d3c6fe6-553b-4c00-b68f-025be6d786c1",
     "role": "character_ref"
    },
    {
     "label": "PROP REFERENCE — 균열된 2020년대 스마트폰: the exact object appearing in this shot; match its look, material and wear exactly.",
     "path": "<bytes:992213>",
     "asset_id": "0f8d1056-90bf-4776-ab97-2224cc50fa6e",
     "role": "prop_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "토니는 화면 밖을 응시하며, 오른손으로 모니터의 'ROOT RECOVERY CLASS: AR-12' 텍스트를 가리키고 있습니다.",
    "built_space": "암석 벽과 제어 기기들이 있는 지하 통제실입니다. 책상 위에는 와이어가 연결된 일반 전화기와 텍스트가 띄워진 모니터가 배치되어 있습니다.",
    "entities": "토니는 참조 이미지의 전술복이 아닌 검은색 반팔 티셔츠를 입고 있습니다. 지시된 신분증 대신 부서진 스마트폰을 들고 있습니다.",
    "hard_violations": [
     "[gemini-pro] 물리적 법칙 위반: 토니의 왼손에 있는 스마트폰이 손가락에 닿지 않은 채 공중에 떠 있습니다 (아무것도 지탱하지 않음)."
    ],
    "physics": "오른팔은 화면을 가리키기 위해 뻗어 있으나, 왼손 근처의 스마트폰은 손가락의 그립 없이 공중에 떠 있어 아무것도 사물을 지탱하지 않습니다."
   },
   {
    "label": "B",
    "direction": "토니는 전경 우측에 있는 인물(앨런)을 바라보며, 왼손으로 모니터의 'ROOT RECOVERY CLASS: AR-12' 텍스트를 가리키고 있습니다.",
    "built_space": "지하 통제실 환경이 잘 구현되어 있으며, 책상 위에 모니터와 스마트폰이 자연스럽게 배치되어 있습니다.",
    "entities": "토니는 참조 이미지와 동일한 전술복을 입고 있으며, 오른손에 부서진 투명 신분증을 들고 있습니다. 책상 위의 스마트폰은 와이어에 연결되어 홀로그램을 투사하고 있습니다. 샷 텍스트에 명시된 인물(앨런)의 어깨와 뒷모습이 오버 더 숄더 구도로 보입니다.",
    "hard_violations": [
     "[gpt] 토니 외의 두 번째 남성 신체를 화면 오른쪽 전경에 추가했다."
    ],
    "physics": "토니는 오른손으로 신분증을 단단히 쥐고 있으며, 스마트폰은 책상 위에 안정적으로 놓여 홀로그램을 투사합니다. 모든 사물과 인물의 자세가 올바르게 지탱되고 있습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "스마트폰이 손가락에 닿지 않고 공중에 떠 있는 심각한 물리적 오류가 있으며, 복장과 소품 상태(신분증 대신 폰을 들고 있음)가 지시문과 전혀 일치하지 않습니다."
       },
       {
        "label": "B",
        "score": 10,
        "verdict_ko": "부서진 신분증, 홀로그램을 투사하는 유선 연결 스마트폰, 모니터 텍스트, 참조 이미지와 일치하는 복장까지 복잡한 지시 사항들을 클로즈업 구도로 완벽하게 구현했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "토니는 화면 밖을 응시하며, 오른손으로 모니터의 'ROOT RECOVERY CLASS: AR-12' 텍스트를 가리키고 있습니다.",
        "built_space": "암석 벽과 제어 기기들이 있는 지하 통제실입니다. 책상 위에는 와이어가 연결된 일반 전화기와 텍스트가 띄워진 모니터가 배치되어 있습니다.",
        "entities": "토니는 참조 이미지의 전술복이 아닌 검은색 반팔 티셔츠를 입고 있습니다. 지시된 신분증 대신 부서진 스마트폰을 들고 있습니다.",
        "hard_violations": [
         "물리적 법칙 위반: 토니의 왼손에 있는 스마트폰이 손가락에 닿지 않은 채 공중에 떠 있습니다 (아무것도 지탱하지 않음)."
        ],
        "physics": "오른팔은 화면을 가리키기 위해 뻗어 있으나, 왼손 근처의 스마트폰은 손가락의 그립 없이 공중에 떠 있어 아무것도 사물을 지탱하지 않습니다."
       },
       {
        "label": "B",
        "direction": "토니는 전경 우측에 있는 인물(앨런)을 바라보며, 왼손으로 모니터의 'ROOT RECOVERY CLASS: AR-12' 텍스트를 가리키고 있습니다.",
        "built_space": "지하 통제실 환경이 잘 구현되어 있으며, 책상 위에 모니터와 스마트폰이 자연스럽게 배치되어 있습니다.",
        "entities": "토니는 참조 이미지와 동일한 전술복을 입고 있으며, 오른손에 부서진 투명 신분증을 들고 있습니다. 책상 위의 스마트폰은 와이어에 연결되어 홀로그램을 투사하고 있습니다. 샷 텍스트에 명시된 인물(앨런)의 어깨와 뒷모습이 오버 더 숄더 구도로 보입니다.",
        "hard_violations": [],
        "physics": "토니는 오른손으로 신분증을 단단히 쥐고 있으며, 스마트폰은 책상 위에 안정적으로 놓여 홀로그램을 투사합니다. 모든 사물과 인물의 자세가 올바르게 지탱되고 있습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 2,
        "verdict_ko": "스마트폰이 손가락에 닿지 않고 공중에 떠 있는 심각한 물리적 오류가 있으며, 복장과 소품 상태(신분증 대신 폰을 들고 있음)가 지시문과 전혀 일치하지 않습니다."
       },
       {
        "label": "B",
        "score": 10,
        "verdict_ko": "부서진 신분증, 홀로그램을 투사하는 유선 연결 스마트폰, 모니터 텍스트, 참조 이미지와 일치하는 복장까지 복잡한 지시 사항들을 클로즈업 구도로 완벽하게 구현했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "토니는 화면 밖을 응시하며, 오른손으로 모니터의 'ROOT RECOVERY CLASS: AR-12' 텍스트를 가리키고 있습니다.",
        "built_space": "암석 벽과 제어 기기들이 있는 지하 통제실입니다. 책상 위에는 와이어가 연결된 일반 전화기와 텍스트가 띄워진 모니터가 배치되어 있습니다.",
        "entities": "토니는 참조 이미지의 전술복이 아닌 검은색 반팔 티셔츠를 입고 있습니다. 지시된 신분증 대신 부서진 스마트폰을 들고 있습니다.",
        "hard_violations": [
         "물리적 법칙 위반: 토니의 왼손에 있는 스마트폰이 손가락에 닿지 않은 채 공중에 떠 있습니다 (아무것도 지탱하지 않음)."
        ],
        "physics": "오른팔은 화면을 가리키기 위해 뻗어 있으나, 왼손 근처의 스마트폰은 손가락의 그립 없이 공중에 떠 있어 아무것도 사물을 지탱하지 않습니다."
       },
       {
        "label": "B",
        "direction": "토니는 전경 우측에 있는 인물(앨런)을 바라보며, 왼손으로 모니터의 'ROOT RECOVERY CLASS: AR-12' 텍스트를 가리키고 있습니다.",
        "built_space": "지하 통제실 환경이 잘 구현되어 있으며, 책상 위에 모니터와 스마트폰이 자연스럽게 배치되어 있습니다.",
        "entities": "토니는 참조 이미지와 동일한 전술복을 입고 있으며, 오른손에 부서진 투명 신분증을 들고 있습니다. 책상 위의 스마트폰은 와이어에 연결되어 홀로그램을 투사하고 있습니다. 샷 텍스트에 명시된 인물(앨런)의 어깨와 뒷모습이 오버 더 숄더 구도로 보입니다.",
        "hard_violations": [],
        "physics": "토니는 오른손으로 신분증을 단단히 쥐고 있으며, 스마트폰은 책상 위에 안정적으로 놓여 홀로그램을 투사합니다. 모든 사물과 인물의 자세가 올바르게 지탱되고 있습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "상체 클로즈업과 앨런을 향한 시선, 토니의 장비·파손 배지 표현은 우수하지만, 금지된 두 번째 인물을 화면에 추가했고 손가락도 권한 문구가 아니라 그 인물을 향해 치명적으로 어긋난다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "추가 인물 없이 손가락이 ‘ROOT RECOVERY CLASS: AR-12’를 정확히 가리키지만, 클로즈업보다 넓고 토니가 앨런이 아닌 화면을 보며 복장과 손의 소지품도 지시와 다르다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "토니의 눈은 화면 오른쪽 전경에 보이는 남성, 즉 앨런으로 연출된 인물을 향한다. 그러나 뻗은 검지는 모니터의 확대 문구를 지나 화면 오른쪽의 그 남성을 향하며, 손끝이 ‘ROOT RECOVERY CLASS: AR-12’에 착지하지 않는다.",
        "built_space": "토니는 제어실 워크스테이션 의자에 앉아 있으며, 뒤에 확대 문구 모니터 1대와 상부 상황 모니터 일부, 제어 랙, 책상 위 투사 장치가 보인다. 정확한 클로즈업이라 유선 전화는 프레임 밖이다. 화면 오른쪽 전경에는 지정되지 않은 두 번째 사람이 워크스테이션 앞에 배치되어 있다. 광학적으로 불가능한 반사는 보이지 않는다.",
        "entities": "토니는 미국인 30대 남성으로 보이고 얼굴·짧은 갈색 머리·체격과 전술 재킷 및 하네스가 참조에 비교적 가깝지만 참조의 헬멧은 없다. 왼손에는 균열되고 일부 파손된 직원 배지처럼 보이는 얇은 카드가 잡혀 있다. 확대 모니터에는 요구된 권한 문구가 보인다. 다만 화면 오른쪽에 토니 외 남성 한 명이 추가되었다.",
        "hard_violations": [
         "토니 외의 두 번째 남성 신체를 화면 오른쪽 전경에 추가했다."
        ],
        "physics": "토니의 몸은 의자에 앉아 지지되고, 오른팔은 어깨와 팔꿈치에서 자연스럽게 뻗어 있다. 파손 배지는 왼손 장갑으로 확실히 쥐고 있어 떠 있지 않는다. 추가 인물도 전경에 서거나 앉은 정상적인 몸으로 보이며 공중에 뜬 물체는 없다."
       },
       {
        "label": "B",
        "direction": "토니의 검지 손끝은 오른쪽 모니터의 확대된 ‘ROOT RECOVERY CLASS: AR-12’ 문구에 정확히 닿는다. 하지만 눈도 같은 모니터를 바라보며, 샷 텍스트가 요구하는 앨런을 향한 시선은 성립하지 않는다.",
        "built_space": "지하 제어실을 넓게 보여 주며 암벽과 금속문, 서버 캐비닛, 긴 워크스테이션, 확대 문구 모니터와 하단 모니터, 유선 탁상전화 1대, 램프 1개, 투사 장치와 원형 중앙 설비가 보인다. 토니는 전화 워크스테이션 옆에 서 있다. 모니터에 비친 손의 반사는 손이 화면 바로 앞에 있으므로 가능한 반사다. 다만 구도는 요구된 클로즈업이 아니라 허리 부근까지 포함한 미디엄 와이드에 가깝다.",
        "entities": "토니 한 명만 보이며 미국인 30대 남성의 얼굴·머리·체격은 대체로 참조 인물과 유사하다. 그러나 검은 반소매 티셔츠는 참조의 헬멧, 전술 재킷, 하네스 복장과 크게 다르다. 왼손 물체는 참조와 비슷한 균열 스마트폰이며 파손 직원 배지로 보이지 않는다. 확대 모니터의 요구 문구와 유선 전화는 존재한다.",
        "hard_violations": [],
        "physics": "토니는 바닥에 서 있는 자세로 읽히고, 뻗은 오른팔은 어깨에서 자연스럽게 연결되어 검지가 화면 표면에 접근한다. 균열 스마트폰은 왼손이 확실히 쥐고 있다. 전화와 램프는 책상 위에 놓여 있고 케이블도 연결되어 있어 지지되지 않은 물체는 없다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "상체 클로즈업과 앨런을 향한 시선, 토니의 장비·파손 배지 표현은 우수하지만, 금지된 두 번째 인물을 화면에 추가했고 손가락도 권한 문구가 아니라 그 인물을 향해 치명적으로 어긋난다."
       },
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "추가 인물 없이 손가락이 ‘ROOT RECOVERY CLASS: AR-12’를 정확히 가리키지만, 클로즈업보다 넓고 토니가 앨런이 아닌 화면을 보며 복장과 손의 소지품도 지시와 다르다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "토니의 눈은 화면 오른쪽 전경에 보이는 남성, 즉 앨런으로 연출된 인물을 향한다. 그러나 뻗은 검지는 모니터의 확대 문구를 지나 화면 오른쪽의 그 남성을 향하며, 손끝이 ‘ROOT RECOVERY CLASS: AR-12’에 착지하지 않는다.",
        "built_space": "토니는 제어실 워크스테이션 의자에 앉아 있으며, 뒤에 확대 문구 모니터 1대와 상부 상황 모니터 일부, 제어 랙, 책상 위 투사 장치가 보인다. 정확한 클로즈업이라 유선 전화는 프레임 밖이다. 화면 오른쪽 전경에는 지정되지 않은 두 번째 사람이 워크스테이션 앞에 배치되어 있다. 광학적으로 불가능한 반사는 보이지 않는다.",
        "entities": "토니는 미국인 30대 남성으로 보이고 얼굴·짧은 갈색 머리·체격과 전술 재킷 및 하네스가 참조에 비교적 가깝지만 참조의 헬멧은 없다. 왼손에는 균열되고 일부 파손된 직원 배지처럼 보이는 얇은 카드가 잡혀 있다. 확대 모니터에는 요구된 권한 문구가 보인다. 다만 화면 오른쪽에 토니 외 남성 한 명이 추가되었다.",
        "hard_violations": [
         "토니 외의 두 번째 남성 신체를 화면 오른쪽 전경에 추가했다."
        ],
        "physics": "토니의 몸은 의자에 앉아 지지되고, 오른팔은 어깨와 팔꿈치에서 자연스럽게 뻗어 있다. 파손 배지는 왼손 장갑으로 확실히 쥐고 있어 떠 있지 않는다. 추가 인물도 전경에 서거나 앉은 정상적인 몸으로 보이며 공중에 뜬 물체는 없다."
       },
       {
        "label": "A",
        "direction": "토니의 검지 손끝은 오른쪽 모니터의 확대된 ‘ROOT RECOVERY CLASS: AR-12’ 문구에 정확히 닿는다. 하지만 눈도 같은 모니터를 바라보며, 샷 텍스트가 요구하는 앨런을 향한 시선은 성립하지 않는다.",
        "built_space": "지하 제어실을 넓게 보여 주며 암벽과 금속문, 서버 캐비닛, 긴 워크스테이션, 확대 문구 모니터와 하단 모니터, 유선 탁상전화 1대, 램프 1개, 투사 장치와 원형 중앙 설비가 보인다. 토니는 전화 워크스테이션 옆에 서 있다. 모니터에 비친 손의 반사는 손이 화면 바로 앞에 있으므로 가능한 반사다. 다만 구도는 요구된 클로즈업이 아니라 허리 부근까지 포함한 미디엄 와이드에 가깝다.",
        "entities": "토니 한 명만 보이며 미국인 30대 남성의 얼굴·머리·체격은 대체로 참조 인물과 유사하다. 그러나 검은 반소매 티셔츠는 참조의 헬멧, 전술 재킷, 하네스 복장과 크게 다르다. 왼손 물체는 참조와 비슷한 균열 스마트폰이며 파손 직원 배지로 보이지 않는다. 확대 모니터의 요구 문구와 유선 전화는 존재한다.",
        "hard_violations": [],
        "physics": "토니는 바닥에 서 있는 자세로 읽히고, 뻗은 오른팔은 어깨에서 자연스럽게 연결되어 검지가 화면 표면에 접근한다. 균열 스마트폰은 왼손이 확실히 쥐고 있다. 전화와 램프는 책상 위에 놓여 있고 케이블도 연결되어 있어 지지되지 않은 물체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt": "A"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt"
   ],
   "normalized": {
    "A": 1.2,
    "B": 1.6
   },
   "adjusted": {
    "A": 0.95,
    "B": 1.35
   },
   "violations": {
    "A": [
     "[gemini-pro] 물리적 법칙 위반: 토니의 왼손에 있는 스마트폰이 손가락에 닿지 않은 채 공중에 떠 있습니다 (아무것도 지탱하지 않음)."
    ],
    "B": [
     "[gpt] 토니 외의 두 번째 남성 신체를 화면 오른쪽 전경에 추가했다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt": "A"
   },
   "agreed": false
  },
  "totals": {
   "A": 950,
   "B": 1350
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 950,
    "verdict_ko": "스마트폰이 손가락에 닿지 않고 공중에 떠 있는 심각한 물리적 오류가 있으며, 복장과 소품 상태(신분증 대신 폰을 들고 있음)가 지시문과 전혀 일치하지 않습니다.  ★위반: [gemini-pro] 물리적 법칙 위반: 토니의 왼손에 있는 스마트폰이 손가락에 닿지 않은 채 공중에 떠 있습니다 (아무것도 지탱하지 않음)."
   },
   {
    "label": "B",
    "score": 1350,
    "verdict_ko": "부서진 신분증, 홀로그램을 투사하는 유선 연결 스마트폰, 모니터 텍스트, 참조 이미지와 일치하는 복장까지 복잡한 지시 사항들을 클로즈업 구도로 완벽하게 구현했습니다.  ★위반: [gpt] 토니 외의 두 번째 남성 신체를 화면 오른쪽 전경에 추가했다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/37154490-a4ed-4b37-9092-489c054d8a72/episodes/cc1d3d0b-626e-4f20-bc33-f5f9e2c6ee7e/images/background_chain/L11B02.png",
    "asset_id": "eb7f72c2-429d-4eb2-880a-e3c8b93c34f5",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 토니(앤서니 로저스): the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:929851>",
    "asset_id": "1d3c6fe6-553b-4c00-b68f-025be6d786c1",
    "role": "character_ref"
   },
   {
    "label": "PROP REFERENCE — 균열된 2020년대 스마트폰: the exact object appearing in this shot; match its look, material and wear exactly.",
    "path": "<bytes:992213>",
    "asset_id": "0f8d1056-90bf-4776-ab97-2224cc50fa6e",
    "role": "prop_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": true,
  "shot_run_uid": "06a9bdb9-bcdd-7462-a847-2be727921895",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/37154490-a4ed-4b37-9092-489c054d8a72/images/cc1d3d0b-626e-4f20-bc33-f5f9e2c6ee7e/scene/recipe/S4sh10__bgfirst_bg.png",
   "bg_asset_id": "e8939552-c2bd-4281-8182-5b67265ca03c",
   "bg_record_key": "S4sh10::bgfirst_bg",
   "chain_winner": false,
   "authority": "plate"
  },
  "ref_mode": "플레이트+엔티티 (2택1: 무콘티 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S4sh10::cine": {
  "applied": true,
  "attempted_at": "2026-09-05T09:09:22.124251+00:00",
  "fingerprint": "7561d209b12b295388fbadee685ca437195140db2cd975466cd751747071b04b",
  "fingerprint_version": 2,
  "provider": "grok",
  "endpoint": "openrouter/chat-completions",
  "model": "x-ai/grok-imagine-image-2.0",
  "pack": "24.202608252115",
  "source_file": "S4sh10_sel.png",
  "source_sha256": "a79f2c19b637fbbaabd3815675af1c309ea85196fba29d62376a4f87bc02c794",
  "file": "S4sh10_cine.png",
  "staged_sha256": "87554d72b57b80579906e3a51531f1ff47d73a2301720cea8d2cbc8530be84d7",
  "latency_ms": 13772
 },
 "S5sh7::signage": {
  "fp": "919590f3d048adeb",
  "inscriptions": [
   {
    "text_native": "UNSENT — 393 YEARS AGO",
    "source": "scene_text_quoted",
    "reason_ko": "모니터 화면 팝업에 표시되는 미전송 알림 문구이다.",
    "source_quote": "UNSENT — 393 YEARS AGO"
   }
  ],
  "cues": [],
  "dropped": []
 },
 "S5sh7": {
  "input_fingerprint": "c28e651437d306ba",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — Wyoming, United States, 2419; all people are English-speaking Americans unless stated. TIME OF DAY (lock): twilight.\n\nSHOT TEXT (authoritative, Korean): 모니터 화면 한구석에 UNSENT — 393 YEARS AGO 팝업이 뜬 클로즈업\n\nLOCATION (lock): Inside the mountain contact post's underground control room, at the system-log monitor. The message notification glows from one corner of the active screen. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: insert close-up on a detail\n- FRAME LAYOUT: UNSENT notification area in the middle-center of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Monitor notification area (Displaying UNSENT — 393 YEARS AGO) — The content-bearing screen face is seen at a slight oblique angle, with the notification fully readable near its visible corner; used as Primary textual focal point and spatial source for Tony's subsequent attention; Monitor edge (Visible beside the notification area) — The near side edge angles away from camera and defines the direction of the upcoming lateral transition; used as Provides depth and a visual exit toward the next shot.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Cool, low-key interior illumination keeps the notification legible with restrained contrast and no additional emphasized source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The monitor shows the restored 2026 mine log and a corner notification reading “UNSENT — 393 YEARS AGO.” The phone remains wired on unstable external power inside the isolation case.\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 토니(앤서니 로저스) (미국인 남성, 30대 중반, 자연스러운 성인 남성 얼굴, 구체적으로 명시되지 않은 자연색 머리카락). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nWORDS TO RENDER (authoritative — the scene itself calls for these; render each as period-real physical lettering in the native script, exactly as written; add no other readable text anywhere):\n- \"UNSENT — 393 YEARS AGO\"\n\nThe WORDS TO RENDER above are the only readable writing in this image: render those words exactly as given, in the place and era's own language and script, and nothing else legible. Invent no other wording a viewer could read. No caption, subtitle, watermark, logo or overlay. Surfaces that would carry writing may still be present — stage any wording they would carry out of legibility: a hand across, an oblique angle, shallow focus.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — Wyoming, United States, 2419; all people are English-speaking Americans unless stated. TIME OF DAY (lock): twilight.\n\nSHOT TEXT (authoritative, Korean): 모니터 화면 한구석에 UNSENT — 393 YEARS AGO 팝업이 뜬 클로즈업\n\nLOCATION (lock): Inside the mountain contact post's underground control room, at the system-log monitor. The message notification glows from one corner of the active screen. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: insert close-up on a detail\n- FRAME LAYOUT: UNSENT notification area in the middle-center of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Monitor notification area (Displaying UNSENT — 393 YEARS AGO) — The content-bearing screen face is seen at a slight oblique angle, with the notification fully readable near its visible corner; used as Primary textual focal point and spatial source for Tony's subsequent attention; Monitor edge (Visible beside the notification area) — The near side edge angles away from camera and defines the direction of the upcoming lateral transition; used as Provides depth and a visual exit toward the next shot.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Cool, low-key interior illumination keeps the notification legible with restrained contrast and no additional emphasized source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The monitor shows the restored 2026 mine log and a corner notification reading “UNSENT — 393 YEARS AGO.” The phone remains wired on unstable external power inside the isolation case.\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 토니(앤서니 로저스) (미국인 남성, 30대 중반, 자연스러운 성인 남성 얼굴, 구체적으로 명시되지 않은 자연색 머리카락). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nWORDS TO RENDER (authoritative — the scene itself calls for these; render each as period-real physical lettering in the native script, exactly as written; add no other readable text anywhere):\n- \"UNSENT — 393 YEARS AGO\"\n\nThe WORDS TO RENDER above are the only readable writing in this image: render those words exactly as given, in the place and era's own language and script, and nothing else legible. Invent no other wording a viewer could read. No caption, subtitle, watermark, logo or overlay. Surfaces that would carry writing may still be present — stage any wording they would carry out of legibility: a hand across, an oblique angle, shallow focus.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — Wyoming, United States, 2419; all people are English-speaking Americans unless stated. TIME OF DAY (lock): twilight.\n\nSHOT TEXT (authoritative, Korean): 모니터 화면 한구석에 UNSENT — 393 YEARS AGO 팝업이 뜬 클로즈업\n\nLOCATION (lock): Inside the mountain contact post's underground control room, at the system-log monitor. The message notification glows from one corner of the active screen. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: insert close-up on a detail\n- FRAME LAYOUT: UNSENT notification area in the middle-center of the frame, midground.\n- KEY BACKGROUND ELEMENTS: Monitor notification area (Displaying UNSENT — 393 YEARS AGO) — The content-bearing screen face is seen at a slight oblique angle, with the notification fully readable near its visible corner; used as Primary textual focal point and spatial source for Tony's subsequent attention; Monitor edge (Visible beside the notification area) — The near side edge angles away from camera and defines the direction of the upcoming lateral transition; used as Provides depth and a visual exit toward the next shot.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Cool, low-key interior illumination keeps the notification legible with restrained contrast and no additional emphasized source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The monitor shows the restored 2026 mine log and a corner notification reading “UNSENT — 393 YEARS AGO.” The phone remains wired on unstable external power inside the isolation case.\n\nPEOPLE: the SHOT TEXT alone decides who is visible in this shot. People known to appear somewhere in this scene: 토니(앤서니 로저스) (미국인 남성, 30대 중반, 자연스러운 성인 남성 얼굴, 구체적으로 명시되지 않은 자연색 머리카락). That list is scene-level, not a cast list for this frame — it may name someone this shot does not show, and it may omit someone this shot does show. If the shot text names a person who is not on the list, draw that person exactly as the shot text describes them; the list does not override the shot text. Never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person.\n\nWORDS TO RENDER (authoritative — the scene itself calls for these; render each as period-real physical lettering in the native script, exactly as written; add no other readable text anywhere):\n- \"UNSENT — 393 YEARS AGO\"\n\nThe WORDS TO RENDER above are the only readable writing in this image: render those words exactly as given, in the place and era's own language and script, and nothing else legible. Invent no other wording a viewer could read. No caption, subtitle, watermark, logo or overlay. Surfaces that would carry writing may still be present — stage any wording they would carry out of legibility: a hand across, an oblique angle, shallow focus.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "카메라는 모니터 화면을 비스듬히 바라보고 있으며, 중앙의 팝업 텍스트에 정확히 초점을 맞추고 있음.",
    "built_space": "지시된 제어실 환경에 맞는 평면 모니터 화면의 클로즈업이며, 우측 가장자리가 심도를 제공함.",
    "entities": "프롬프트가 명시한 'UNSENT — 393 YEARS AGO'가 엠대시(—)와 함께 정확히 표기되었으며, 나머지 화면 텍스트는 지시대로 얕은 심도로 흐려져 읽기 어렵게 렌더링됨.",
    "hard_violations": [],
    "physics": "모니터와 화면의 빛 반사 및 디스플레이 형태가 물리적 법칙에 맞게 자연스럽게 표현됨."
   },
   {
    "label": "B",
    "direction": "카메라는 모니터 화면을 비스듬히 바라보며 중앙의 팝업 텍스트와 상단 영역을 향하고 있음.",
    "built_space": "두꺼운 형태의 모니터가 책상 위에 배치되어 있으며, 좌측 및 상단 테두리가 화면 공간을 둘러싸고 있음.",
    "entities": "지정된 팝업 텍스트가 표기되었으나, 프롬프트 지시에 어긋나게 'SYSTEM LOG'라는 발명된 텍스트가 화면 좌측 상단에 선명하게 추가됨.",
    "hard_violations": [
     "[gemini-pro] 지정된 문구 외의 다른 텍스트를 선명하게 추가함 ('SYSTEM LOG')",
     "[gpt] 화면에 허용 문구 외의 ‘SYSTEM LOG’ 및 일부 가독 가능한 로그 텍스트가 추가로 노출되었다."
    ],
    "physics": "무거운 형태의 모니터가 책상 표면에 물리적으로 안정감 있게 놓여 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 10,
        "verdict_ko": "프롬프트가 요구한 정확한 텍스트와 구도를 완벽히 구현했으며, 화면 내 다른 텍스트들을 얕은 심도로 흐리게 처리해 지시사항을 철저히 따름."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "요구된 구도는 잡았으나 'SYSTEM LOG'라는 추가 텍스트를 선명하게 노출시켜, 지정된 문구 외에는 읽을 수 없어야 한다는 규칙을 위반함."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "카메라는 모니터 화면을 비스듬히 바라보고 있으며, 중앙의 팝업 텍스트에 정확히 초점을 맞추고 있음.",
        "built_space": "지시된 제어실 환경에 맞는 평면 모니터 화면의 클로즈업이며, 우측 가장자리가 심도를 제공함.",
        "entities": "프롬프트가 명시한 'UNSENT — 393 YEARS AGO'가 엠대시(—)와 함께 정확히 표기되었으며, 나머지 화면 텍스트는 지시대로 얕은 심도로 흐려져 읽기 어렵게 렌더링됨.",
        "hard_violations": [],
        "physics": "모니터와 화면의 빛 반사 및 디스플레이 형태가 물리적 법칙에 맞게 자연스럽게 표현됨."
       },
       {
        "label": "B",
        "direction": "카메라는 모니터 화면을 비스듬히 바라보며 중앙의 팝업 텍스트와 상단 영역을 향하고 있음.",
        "built_space": "두꺼운 형태의 모니터가 책상 위에 배치되어 있으며, 좌측 및 상단 테두리가 화면 공간을 둘러싸고 있음.",
        "entities": "지정된 팝업 텍스트가 표기되었으나, 프롬프트 지시에 어긋나게 'SYSTEM LOG'라는 발명된 텍스트가 화면 좌측 상단에 선명하게 추가됨.",
        "hard_violations": [
         "지정된 문구 외의 다른 텍스트를 선명하게 추가함 ('SYSTEM LOG')"
        ],
        "physics": "무거운 형태의 모니터가 책상 표면에 물리적으로 안정감 있게 놓여 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 10,
        "verdict_ko": "프롬프트가 요구한 정확한 텍스트와 구도를 완벽히 구현했으며, 화면 내 다른 텍스트들을 얕은 심도로 흐리게 처리해 지시사항을 철저히 따름."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "요구된 구도는 잡았으나 'SYSTEM LOG'라는 추가 텍스트를 선명하게 노출시켜, 지정된 문구 외에는 읽을 수 없어야 한다는 규칙을 위반함."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "카메라는 모니터 화면을 비스듬히 바라보고 있으며, 중앙의 팝업 텍스트에 정확히 초점을 맞추고 있음.",
        "built_space": "지시된 제어실 환경에 맞는 평면 모니터 화면의 클로즈업이며, 우측 가장자리가 심도를 제공함.",
        "entities": "프롬프트가 명시한 'UNSENT — 393 YEARS AGO'가 엠대시(—)와 함께 정확히 표기되었으며, 나머지 화면 텍스트는 지시대로 얕은 심도로 흐려져 읽기 어렵게 렌더링됨.",
        "hard_violations": [],
        "physics": "모니터와 화면의 빛 반사 및 디스플레이 형태가 물리적 법칙에 맞게 자연스럽게 표현됨."
       },
       {
        "label": "B",
        "direction": "카메라는 모니터 화면을 비스듬히 바라보며 중앙의 팝업 텍스트와 상단 영역을 향하고 있음.",
        "built_space": "두꺼운 형태의 모니터가 책상 위에 배치되어 있으며, 좌측 및 상단 테두리가 화면 공간을 둘러싸고 있음.",
        "entities": "지정된 팝업 텍스트가 표기되었으나, 프롬프트 지시에 어긋나게 'SYSTEM LOG'라는 발명된 텍스트가 화면 좌측 상단에 선명하게 추가됨.",
        "hard_violations": [
         "지정된 문구 외의 다른 텍스트를 선명하게 추가함 ('SYSTEM LOG')"
        ],
        "physics": "무거운 형태의 모니터가 책상 표면에 물리적으로 안정감 있게 놓여 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "알림 문구와 비스듬한 화면 구도는 맞지만, ‘SYSTEM LOG’를 비롯한 금지된 추가 가독 텍스트가 있어 단일 허용 문구 조건을 위반한다."
       },
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "중앙부의 정확한 ‘UNSENT — 393 YEARS AGO’ 알림, 비스듬히 보이는 화면과 우측 모니터 가장자리, 차갑고 절제된 조명을 가장 충실히 구현했다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "사람의 시선, 무기, 이동체는 없다. 모니터의 기능면이 카메라를 향해 비스듬히 놓였고, 가까운 좌측·상단 테두리에서 화면 면이 우하단 방향으로 멀어진다.",
        "built_space": "고정된 시스템 로그 모니터 1대의 화면과 두꺼운 금속성 베젤이 크게 보이며, 좌측에는 책상 위에 놓인 직사각형 장비 또는 물체 1개가 일부 보인다. 알림은 화면 상단 쪽이면서 프레임 중앙에 놓였다. 화면이나 물체에 광학적으로 불가능한 반사는 없다.",
        "entities": "핵심 알림은 정확히 ‘UNSENT — 393 YEARS AGO’로 표시된다. 사람은 없으며 이는 인서트 클로즈업 지시와 맞는다. 다만 화면 상단의 ‘SYSTEM LOG’와 여러 로그 문장이 추가로 보이고 일부는 읽을 수 있어, 허용된 유일한 문구 조건과 맞지 않는다.",
        "hard_violations": [
         "화면에 허용 문구 외의 ‘SYSTEM LOG’ 및 일부 가독 가능한 로그 텍스트가 추가로 노출되었다."
        ],
        "physics": "모니터는 두꺼운 베젤과 뒤쪽 하우징에 고정되어 있고 책상 또는 콘솔 위에 설치된 것으로 보인다. 좌측 직사각형 물체도 책상 표면에 놓여 있다. 떠 있거나 지지 없는 물체는 없다."
       },
       {
        "label": "B",
        "direction": "사람의 시선, 무기, 이동체는 없다. 화면 기능면은 카메라 쪽을 향한 약한 사선이며, 우측 모니터 가장자리가 뒤로 물러나는 방향을 명확히 만들어 다음 측면 전환의 시각적 출구가 된다.",
        "built_space": "시스템 로그 모니터 1대의 화면과 우측 베젤 1개가 보이는 인서트 클로즈업이다. 알림은 프레임의 중앙에서 약간 오른쪽인 중경에 있고, 화면의 가시 우측 영역 가까이에 배치됐다. 주변 로그는 초점과 크기 때문에 읽기 어렵게 처리되었으며 불가능한 반사는 없다.",
        "entities": "정확한 알림 문구 ‘UNSENT — 393 YEARS AGO’가 완전히 읽힌다. 사람이나 불필요한 전경 소품은 없다. 나머지 화면 내용은 시스템 로그 형태지만 실질적으로 판독되지 않아 핵심 문구만 읽히는 조건을 대체로 지킨다.",
        "hard_violations": [],
        "physics": "화면은 우측의 단단한 베젤과 모니터 하우징에 물리적으로 고정되어 있다. 움직이거나 공중에 뜬 물체가 없고, 지지 없는 요소도 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "알림 문구와 비스듬한 화면 구도는 맞지만, ‘SYSTEM LOG’를 비롯한 금지된 추가 가독 텍스트가 있어 단일 허용 문구 조건을 위반한다."
       },
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "중앙부의 정확한 ‘UNSENT — 393 YEARS AGO’ 알림, 비스듬히 보이는 화면과 우측 모니터 가장자리, 차갑고 절제된 조명을 가장 충실히 구현했다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "사람의 시선, 무기, 이동체는 없다. 모니터의 기능면이 카메라를 향해 비스듬히 놓였고, 가까운 좌측·상단 테두리에서 화면 면이 우하단 방향으로 멀어진다.",
        "built_space": "고정된 시스템 로그 모니터 1대의 화면과 두꺼운 금속성 베젤이 크게 보이며, 좌측에는 책상 위에 놓인 직사각형 장비 또는 물체 1개가 일부 보인다. 알림은 화면 상단 쪽이면서 프레임 중앙에 놓였다. 화면이나 물체에 광학적으로 불가능한 반사는 없다.",
        "entities": "핵심 알림은 정확히 ‘UNSENT — 393 YEARS AGO’로 표시된다. 사람은 없으며 이는 인서트 클로즈업 지시와 맞는다. 다만 화면 상단의 ‘SYSTEM LOG’와 여러 로그 문장이 추가로 보이고 일부는 읽을 수 있어, 허용된 유일한 문구 조건과 맞지 않는다.",
        "hard_violations": [
         "화면에 허용 문구 외의 ‘SYSTEM LOG’ 및 일부 가독 가능한 로그 텍스트가 추가로 노출되었다."
        ],
        "physics": "모니터는 두꺼운 베젤과 뒤쪽 하우징에 고정되어 있고 책상 또는 콘솔 위에 설치된 것으로 보인다. 좌측 직사각형 물체도 책상 표면에 놓여 있다. 떠 있거나 지지 없는 물체는 없다."
       },
       {
        "label": "A",
        "direction": "사람의 시선, 무기, 이동체는 없다. 화면 기능면은 카메라 쪽을 향한 약한 사선이며, 우측 모니터 가장자리가 뒤로 물러나는 방향을 명확히 만들어 다음 측면 전환의 시각적 출구가 된다.",
        "built_space": "시스템 로그 모니터 1대의 화면과 우측 베젤 1개가 보이는 인서트 클로즈업이다. 알림은 프레임의 중앙에서 약간 오른쪽인 중경에 있고, 화면의 가시 우측 영역 가까이에 배치됐다. 주변 로그는 초점과 크기 때문에 읽기 어렵게 처리되었으며 불가능한 반사는 없다.",
        "entities": "정확한 알림 문구 ‘UNSENT — 393 YEARS AGO’가 완전히 읽힌다. 사람이나 불필요한 전경 소품은 없다. 나머지 화면 내용은 시스템 로그 형태지만 실질적으로 판독되지 않아 핵심 문구만 읽히는 조건을 대체로 지킨다.",
        "hard_violations": [],
        "physics": "화면은 우측의 단단한 베젤과 모니터 하우징에 물리적으로 고정되어 있다. 움직이거나 공중에 뜬 물체가 없고, 지지 없는 요소도 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.044
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.794
   },
   "violations": {
    "B": [
     "[gemini-pro] 지정된 문구 외의 다른 텍스트를 선명하게 추가함 ('SYSTEM LOG')",
     "[gpt] 화면에 허용 문구 외의 ‘SYSTEM LOG’ 및 일부 가독 가능한 로그 텍스트가 추가로 노출되었다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 794
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "프롬프트가 요구한 정확한 텍스트와 구도를 완벽히 구현했으며, 화면 내 다른 텍스트들을 얕은 심도로 흐리게 처리해 지시사항을 철저히 따름."
   },
   {
    "label": "B",
    "score": 794,
    "verdict_ko": "요구된 구도는 잡았으나 'SYSTEM LOG'라는 추가 텍스트를 선명하게 노출시켜, 지정된 문구 외에는 읽을 수 없어야 한다는 규칙을 위반함.  ★위반: [gemini-pro] 지정된 문구 외의 다른 텍스트를 선명하게 추가함 ('SYSTEM LOG') / [gpt] 화면에 허용 문구 외의 ‘SYSTEM LOG’ 및 일부 가독 가능한 로그 텍스트가 추가로 노출되었다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/37154490-a4ed-4b37-9092-489c054d8a72/episodes/cc1d3d0b-626e-4f20-bc33-f5f9e2c6ee7e/images/background_chain/L11B02.png",
    "asset_id": "eb7f72c2-429d-4eb2-880a-e3c8b93c34f5",
    "role": "location_plate"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": true,
  "shot_run_uid": "06a9bdc5-00bb-741b-bfd6-1bce1c036629",
  "ref_mode": "플레이트만 (배경 전용)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S5sh7::cine": {
  "applied": true,
  "attempted_at": "2026-09-05T09:10:33.240162+00:00",
  "fingerprint": "fb950aeeb7073d67e84f883e8dfbf9007b226908eaa284d2f5a93281c62087f6",
  "fingerprint_version": 2,
  "provider": "grok",
  "endpoint": "openrouter/chat-completions",
  "model": "x-ai/grok-imagine-image-2.0",
  "pack": "24.202608252115",
  "source_file": "S5sh7_sel.png",
  "source_sha256": "1643279a699faf6f9ee2ec9b5163b9f83d0d56de8a747884370392b0557652fa",
  "file": "S5sh7_cine.png",
  "staged_sha256": "e5bf70bed8107663d1bf815e290b2093ecb1f1ceca039603aad80a4eb8115526",
  "latency_ms": 10581
 },
 "S5sh8::signage": {
  "fp": "6b43d950418633bf",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S5sh8": {
  "input_fingerprint": "fb7de4fb6b9aa2d0",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — Wyoming, United States, 2419; all people are English-speaking Americans unless stated. TIME OF DAY (lock): twilight.\n\nSHOT TEXT (authoritative, Korean): 팝업 창의 불빛을 받으며 입을 꾹 다문 채 고뇌하는 토니의 얼굴\n\nLOCATION (lock): Inside the mountain contact post's underground control room, immediately before the system-log monitor. The message popup casts screen light onto the silent man's face. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: 토니(앤서니 로저스) in the middle-center of the frame, foreground, looks toward off-frame message notification.\n- KEY BACKGROUND ELEMENTS: Monitor edge and partial popup (Popup remains active, with only a narrow portion visible) — The content-bearing face falls near the frame edge; only part of the active notification is visible while Tony receives its light; used as Retains the source of Tony's attention without competing with his expression.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The notification's light passes across Tony's conflicted face within an otherwise cool, muted, low-key interior.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established — the place, its fixed features and wear, each person's clothing and state — persists. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Take the same monitor station, fixed control hardware, rough stone interior, and red ambient lighting from the reference. Exclude the previously enlarged certificate text and retain only the notification glow relevant to this moment.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The “UNSENT — 393 YEARS AGO” notification remains lit on the monitor, casting its light into the room. The phone remains wired inside the transparent isolation case. 토니(앤서니 로저스): He continues holding the broken employee ID badge and pauses without answering.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 토니(앤서니 로저스) (미국인 남성, 30대 중반, 자연스러운 성인 남성 얼굴, 구체적으로 명시되지 않은 자연색 머리카락) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — Wyoming, United States, 2419; all people are English-speaking Americans unless stated. TIME OF DAY (lock): twilight.\n\nSHOT TEXT (authoritative, Korean): 팝업 창의 불빛을 받으며 입을 꾹 다문 채 고뇌하는 토니의 얼굴\n\nLOCATION (lock): Inside the mountain contact post's underground control room, immediately before the system-log monitor. The message popup casts screen light onto the silent man's face. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: 토니(앤서니 로저스) in the middle-center of the frame, foreground, looks toward off-frame message notification.\n- KEY BACKGROUND ELEMENTS: Monitor edge and partial popup (Popup remains active, with only a narrow portion visible) — The content-bearing face falls near the frame edge; only part of the active notification is visible while Tony receives its light; used as Retains the source of Tony's attention without competing with his expression.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The notification's light passes across Tony's conflicted face within an otherwise cool, muted, low-key interior.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established — the place, its fixed features and wear, each person's clothing and state — persists. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Take the same monitor station, fixed control hardware, rough stone interior, and red ambient lighting from the reference. Exclude the previously enlarged certificate text and retain only the notification glow relevant to this moment.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The “UNSENT — 393 YEARS AGO” notification remains lit on the monitor, casting its light into the room. The phone remains wired inside the transparent isolation case. 토니(앤서니 로저스): He continues holding the broken employee ID badge and pauses without answering.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 토니(앤서니 로저스) (미국인 남성, 30대 중반, 자연스러운 성인 남성 얼굴, 구체적으로 명시되지 않은 자연색 머리카락) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — Wyoming, United States, 2419; all people are English-speaking Americans unless stated. TIME OF DAY (lock): twilight.\n\nSHOT TEXT (authoritative, Korean): 팝업 창의 불빛을 받으며 입을 꾹 다문 채 고뇌하는 토니의 얼굴\n\nLOCATION (lock): Inside the mountain contact post's underground control room, immediately before the system-log monitor. The message popup casts screen light onto the silent man's face. The shot takes place here — the attached PREVIOUS SHOT STILL shows this exact place.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: 토니(앤서니 로저스) in the middle-center of the frame, foreground, looks toward off-frame message notification.\n- KEY BACKGROUND ELEMENTS: Monitor edge and partial popup (Popup remains active, with only a narrow portion visible) — The content-bearing face falls near the frame edge; only part of the active notification is visible while Tony receives its light; used as Retains the source of Tony's attention without competing with his expression.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The notification's light passes across Tony's conflicted face within an otherwise cool, muted, low-key interior.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nTHIS SHOT CONTINUES THE PREVIOUS SHOT: everything the attached still established — the place, its fixed features and wear, each person's clothing and state — persists. Any person in it who cannot move stays PRECISELY as photographed (body, pose, contact points, held objects); only the camera changes.\n\nPREVIOUS STILL USAGE (follow exactly — what to take from the attached still and what to exclude; it governs place and objects only, never who is in this shot or how the camera sees them): Take the same monitor station, fixed control hardware, rough stone interior, and red ambient lighting from the reference. Exclude the previously enlarged certificate text and retain only the notification glow relevant to this moment.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The “UNSENT — 393 YEARS AGO” notification remains lit on the monitor, casting its light into the room. The phone remains wired inside the transparent isolation case. 토니(앤서니 로저스): He continues holding the broken employee ID badge and pauses without answering.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 토니(앤서니 로저스) (미국인 남성, 30대 중반, 자연스러운 성인 남성 얼굴, 구체적으로 명시되지 않은 자연색 머리카락) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "토니의 시선은 화면 좌측 가장자리에 위치한 모니터를 향하고 있습니다.",
    "built_space": "거친 석재 벽면, 붉은빛의 배경 조명, 우측의 제어 하드웨어 등 지하 통제실의 배경 요소가 레퍼런스와 일치하게 잘 구현되었습니다.",
    "entities": "토니의 얼굴, 머리, 의상은 레퍼런스와 일치합니다. 좌측 프레임 가장자리에 잘린 모니터와 팝업창의 일부만 보이며, 가슴 쪽에 부서진 신분증이 있습니다.",
    "hard_violations": [],
    "physics": "부서진 신분증은 토니의 가슴 장비에 클립과 끈으로 연결되어 물리적으로 올바르게 지지되고 있으며, 부자연스럽게 떠 있는 물체는 없습니다."
   },
   {
    "label": "B",
    "direction": "토니의 시선은 화면 좌측에 온전히 드러난 모니터를 향하고 있습니다.",
    "built_space": "거친 석재 벽면과 붉은 조명, 뒤쪽의 기계 제어판 등 레퍼런스에서 제시된 통제실 환경의 특징들을 잘 반영했습니다.",
    "entities": "토니의 인물 외형과 의상은 레퍼런스와 동일합니다. 화면 좌측 모니터에 알림창 전체가 노출되어 있고, 전경에서 손이 부서진 신분증을 들고 있습니다.",
    "hard_violations": [
     "[gpt] 읽을 수 있는 ‘UNSENT — 393 YEARS AGO’ 팝업 문구와 화면 글자가 노출되어 ‘어디에도 읽을 수 있는 글이 없어야 한다’는 조건을 위반한다."
    ],
    "physics": "부서진 신분증은 장갑을 낀 손에 단단히 쥐어져 있어 물리적 지지가 명확합니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "모니터 가장자리와 팝업창 일부만 보이게 배치하고 얕은 심도를 적용하라는 구도 지시를 훌륭하게 수행했으나, 신분증을 손으로 쥐고 있지 않은 점이 사소한 감점 요인입니다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "신분증을 들고 있는 액션은 묘사되었으나, 팝업창 전체가 선명하게 드러나 '일부만 보여야 한다'는 핵심 프레이밍 및 가독성 배제 지시를 위반했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "토니의 시선은 화면 좌측 가장자리에 위치한 모니터를 향하고 있습니다.",
        "built_space": "거친 석재 벽면, 붉은빛의 배경 조명, 우측의 제어 하드웨어 등 지하 통제실의 배경 요소가 레퍼런스와 일치하게 잘 구현되었습니다.",
        "entities": "토니의 얼굴, 머리, 의상은 레퍼런스와 일치합니다. 좌측 프레임 가장자리에 잘린 모니터와 팝업창의 일부만 보이며, 가슴 쪽에 부서진 신분증이 있습니다.",
        "hard_violations": [],
        "physics": "부서진 신분증은 토니의 가슴 장비에 클립과 끈으로 연결되어 물리적으로 올바르게 지지되고 있으며, 부자연스럽게 떠 있는 물체는 없습니다."
       },
       {
        "label": "B",
        "direction": "토니의 시선은 화면 좌측에 온전히 드러난 모니터를 향하고 있습니다.",
        "built_space": "거친 석재 벽면과 붉은 조명, 뒤쪽의 기계 제어판 등 레퍼런스에서 제시된 통제실 환경의 특징들을 잘 반영했습니다.",
        "entities": "토니의 인물 외형과 의상은 레퍼런스와 동일합니다. 화면 좌측 모니터에 알림창 전체가 노출되어 있고, 전경에서 손이 부서진 신분증을 들고 있습니다.",
        "hard_violations": [],
        "physics": "부서진 신분증은 장갑을 낀 손에 단단히 쥐어져 있어 물리적 지지가 명확합니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "모니터 가장자리와 팝업창 일부만 보이게 배치하고 얕은 심도를 적용하라는 구도 지시를 훌륭하게 수행했으나, 신분증을 손으로 쥐고 있지 않은 점이 사소한 감점 요인입니다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "신분증을 들고 있는 액션은 묘사되었으나, 팝업창 전체가 선명하게 드러나 '일부만 보여야 한다'는 핵심 프레이밍 및 가독성 배제 지시를 위반했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "토니의 시선은 화면 좌측 가장자리에 위치한 모니터를 향하고 있습니다.",
        "built_space": "거친 석재 벽면, 붉은빛의 배경 조명, 우측의 제어 하드웨어 등 지하 통제실의 배경 요소가 레퍼런스와 일치하게 잘 구현되었습니다.",
        "entities": "토니의 얼굴, 머리, 의상은 레퍼런스와 일치합니다. 좌측 프레임 가장자리에 잘린 모니터와 팝업창의 일부만 보이며, 가슴 쪽에 부서진 신분증이 있습니다.",
        "hard_violations": [],
        "physics": "부서진 신분증은 토니의 가슴 장비에 클립과 끈으로 연결되어 물리적으로 올바르게 지지되고 있으며, 부자연스럽게 떠 있는 물체는 없습니다."
       },
       {
        "label": "B",
        "direction": "토니의 시선은 화면 좌측에 온전히 드러난 모니터를 향하고 있습니다.",
        "built_space": "거친 석재 벽면과 붉은 조명, 뒤쪽의 기계 제어판 등 레퍼런스에서 제시된 통제실 환경의 특징들을 잘 반영했습니다.",
        "entities": "토니의 인물 외형과 의상은 레퍼런스와 동일합니다. 화면 좌측 모니터에 알림창 전체가 노출되어 있고, 전경에서 손이 부서진 신분증을 들고 있습니다.",
        "hard_violations": [],
        "physics": "부서진 신분증은 장갑을 낀 손에 단단히 쥐어져 있어 물리적 지지가 명확합니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "토니를 중앙에 둔 얼굴 중심 클로즈업과 화면 가장자리의 흐릿하고 부분적인 팝업, 모니터를 향한 시선과 굳게 다문 고뇌의 표정이 핵심 구도를 가장 충실히 따른다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "토니의 시선과 표정 및 배지 파지는 맞지만, 구도가 더 넓고 토니가 우측으로 치우쳤으며 팝업 문구가 완전히 노출되어 읽히므로 명시된 프레이밍과 무가독성 조건을 위반한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "토니의 두 눈과 얼굴은 프레임 왼쪽의 모니터 및 활성 알림을 정확히 향한다. 카메라나 다른 곳을 보는 것이 아니며 시선의 목표가 화면에 보인다.",
        "built_space": "지하 통제실 안에 모니터 1대와 활성 팝업 1개가 왼쪽에 보이고, 뒤에는 거친 암벽과 낡은 금속 제어반 및 붉은 조명이 있다. 토니는 모니터 바로 앞 제어석 위치에 있으므로 공간 관계는 성립한다. 다만 모니터 화면과 팝업이 좁은 가장자리 조각이 아니라 상당히 넓게 노출되고, 토니도 중앙보다 오른쪽에 놓였다. 불가능한 반사는 없다.",
        "entities": "등장 인물은 30대 중반의 미국인 남성 토니 한 명뿐이며 얼굴·짧은 자연색 머리·체격과 오염된 전술복은 인물 참고와 대체로 일치한다. 왼쪽에 시스템 로그 모니터와 붉은 알림이 있고, 전경의 장갑 낀 손에는 파손된 직원 ID 배지가 잡혀 있다. 그러나 알림의 영문 문구와 주변 로그 글자가 읽힐 정도로 드러난다.",
        "hard_violations": [
         "읽을 수 있는 ‘UNSENT — 393 YEARS AGO’ 팝업 문구와 화면 글자가 노출되어 ‘어디에도 읽을 수 있는 글이 없어야 한다’는 조건을 위반한다."
        ],
        "physics": "토니의 상체는 제어석 앞에서 자연스럽게 세워져 있고 프레임 아래로 이어져 지지 관계에 이상이 없다. 파손된 ID 배지는 전경의 장갑 낀 손이 직접 움켜쥐어 지지하며, 공중에 떠 있는 물체나 비정상적인 신체는 없다."
       },
       {
        "label": "B",
        "direction": "토니는 눈과 얼굴을 프레임 왼쪽의 모니터 가장자리와 부분적으로 보이는 활성 알림 쪽으로 향한다. 시선은 알림에 정확히 도달하며 입을 굳게 다문 채 갈등하는 표정이다.",
        "built_space": "왼쪽 가장자리에 모니터 1대와 잘린 활성 팝업 1개가 보이고, 뒤에는 거친 암벽과 붉은빛을 받는 고정 제어 하드웨어가 있다. 토니는 모니터 바로 앞 중앙에 위치하며 설비와 인물의 크기 관계도 자연스럽다. 모니터는 전경 가장자리에서 크게 보이는 위치 관계가 가능하고 불가능한 반사는 없다.",
        "entities": "보이는 사람은 토니 한 명뿐이며 30대 중반 미국인 남성, 짧은 갈색 계열 머리, 자연스러운 성인 남성 얼굴과 낡은 전술복이 참고 인물과 대체로 일치한다. 시스템 로그 모니터와 붉은 팝업은 얕은 초점과 가장자리 절단으로 내용이 명확히 읽히지 않는다. 파손된 ID 배지는 프레임 하단에 일부만 보이며 손은 올바른 클로즈업 때문에 화면 밖이다.",
        "hard_violations": [],
        "physics": "토니의 상체는 모니터 앞에 자연스럽게 서 있거나 앉은 자세로 프레임 아래까지 이어져 있으며 부유하지 않는다. 하단의 배지는 일부만 보여 직접 쥔 손은 확인되지 않지만, 프레임 밖 손 또는 목·가슴 쪽 연결부로 이어지는 위치이고 화면 안에서 지지 없이 떠 있는 것으로 보이지 않는다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "토니를 중앙에 둔 얼굴 중심 클로즈업과 화면 가장자리의 흐릿하고 부분적인 팝업, 모니터를 향한 시선과 굳게 다문 고뇌의 표정이 핵심 구도를 가장 충실히 따른다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "토니의 시선과 표정 및 배지 파지는 맞지만, 구도가 더 넓고 토니가 우측으로 치우쳤으며 팝업 문구가 완전히 노출되어 읽히므로 명시된 프레이밍과 무가독성 조건을 위반한다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "토니의 두 눈과 얼굴은 프레임 왼쪽의 모니터 및 활성 알림을 정확히 향한다. 카메라나 다른 곳을 보는 것이 아니며 시선의 목표가 화면에 보인다.",
        "built_space": "지하 통제실 안에 모니터 1대와 활성 팝업 1개가 왼쪽에 보이고, 뒤에는 거친 암벽과 낡은 금속 제어반 및 붉은 조명이 있다. 토니는 모니터 바로 앞 제어석 위치에 있으므로 공간 관계는 성립한다. 다만 모니터 화면과 팝업이 좁은 가장자리 조각이 아니라 상당히 넓게 노출되고, 토니도 중앙보다 오른쪽에 놓였다. 불가능한 반사는 없다.",
        "entities": "등장 인물은 30대 중반의 미국인 남성 토니 한 명뿐이며 얼굴·짧은 자연색 머리·체격과 오염된 전술복은 인물 참고와 대체로 일치한다. 왼쪽에 시스템 로그 모니터와 붉은 알림이 있고, 전경의 장갑 낀 손에는 파손된 직원 ID 배지가 잡혀 있다. 그러나 알림의 영문 문구와 주변 로그 글자가 읽힐 정도로 드러난다.",
        "hard_violations": [
         "읽을 수 있는 ‘UNSENT — 393 YEARS AGO’ 팝업 문구와 화면 글자가 노출되어 ‘어디에도 읽을 수 있는 글이 없어야 한다’는 조건을 위반한다."
        ],
        "physics": "토니의 상체는 제어석 앞에서 자연스럽게 세워져 있고 프레임 아래로 이어져 지지 관계에 이상이 없다. 파손된 ID 배지는 전경의 장갑 낀 손이 직접 움켜쥐어 지지하며, 공중에 떠 있는 물체나 비정상적인 신체는 없다."
       },
       {
        "label": "A",
        "direction": "토니는 눈과 얼굴을 프레임 왼쪽의 모니터 가장자리와 부분적으로 보이는 활성 알림 쪽으로 향한다. 시선은 알림에 정확히 도달하며 입을 굳게 다문 채 갈등하는 표정이다.",
        "built_space": "왼쪽 가장자리에 모니터 1대와 잘린 활성 팝업 1개가 보이고, 뒤에는 거친 암벽과 붉은빛을 받는 고정 제어 하드웨어가 있다. 토니는 모니터 바로 앞 중앙에 위치하며 설비와 인물의 크기 관계도 자연스럽다. 모니터는 전경 가장자리에서 크게 보이는 위치 관계가 가능하고 불가능한 반사는 없다.",
        "entities": "보이는 사람은 토니 한 명뿐이며 30대 중반 미국인 남성, 짧은 갈색 계열 머리, 자연스러운 성인 남성 얼굴과 낡은 전술복이 참고 인물과 대체로 일치한다. 시스템 로그 모니터와 붉은 팝업은 얕은 초점과 가장자리 절단으로 내용이 명확히 읽히지 않는다. 파손된 ID 배지는 프레임 하단에 일부만 보이며 손은 올바른 클로즈업 때문에 화면 밖이다.",
        "hard_violations": [],
        "physics": "토니의 상체는 모니터 앞에 자연스럽게 서 있거나 앉은 자세로 프레임 아래까지 이어져 있으며 부유하지 않는다. 하단의 배지는 일부만 보여 직접 쥔 손은 확인되지 않지만, 프레임 밖 손 또는 목·가슴 쪽 연결부로 이어지는 위치이고 화면 안에서 지지 없이 떠 있는 것으로 보이지 않는다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.125
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.875
   },
   "violations": {
    "B": [
     "[gpt] 읽을 수 있는 ‘UNSENT — 393 YEARS AGO’ 팝업 문구와 화면 글자가 노출되어 ‘어디에도 읽을 수 있는 글이 없어야 한다’는 조건을 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 875
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "모니터 가장자리와 팝업창 일부만 보이게 배치하고 얕은 심도를 적용하라는 구도 지시를 훌륭하게 수행했으나, 신분증을 손으로 쥐고 있지 않은 점이 사소한 감점 요인입니다."
   },
   {
    "label": "B",
    "score": 875,
    "verdict_ko": "신분증을 들고 있는 액션은 묘사되었으나, 팝업창 전체가 선명하게 드러나 '일부만 보여야 한다'는 핵심 프레이밍 및 가독성 배제 지시를 위반했습니다.  ★위반: [gpt] 읽을 수 있는 ‘UNSENT — 393 YEARS AGO’ 팝업 문구와 화면 글자가 노출되어 ‘어디에도 읽을 수 있는 글이 없어야 한다’는 조건을 위반한다."
   }
  ],
  "refs": [
   {
    "label": "PREVIOUS SHOT STILL — a visually related earlier shot of this same place: the location's look, materials, fixed features, lighting mood and each person's clothing are LOCKED to this photo; never copy its camera framing. If this photo shows a character who cannot move (dead or unconscious), that character's exact pose, position and orientation are ALSO LOCKED — treat their body as a fixed prop of the set that only the camera moves around.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/37154490-a4ed-4b37-9092-489c054d8a72/images/cc1d3d0b-626e-4f20-bc33-f5f9e2c6ee7e/scene/recipe/S5sh7_sel.png",
    "asset_id": "f55557a1-7836-4834-bb66-d9de441f646a",
    "role": "prev_still"
   },
   {
    "label": "CHARACTER REFERENCE — 토니(앤서니 로저스): the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:929851>",
    "asset_id": "1d3c6fe6-553b-4c00-b68f-025be6d786c1",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": true,
  "shot_run_uid": "06a9bdc9-3f49-7e88-9472-c15922f7f549",
  "ref_mode": "prev+엔티티",
  "share_plan": {
   "ref_plan": "prev",
   "prev_anchor_tag": "S5sh7"
  }
 },
 "S5sh8::cine": {
  "applied": true,
  "attempted_at": "2026-09-05T09:11:40.332397+00:00",
  "fingerprint": "36431a4a8d6b5ab56db95dc06904e5cf7e8b9eaa22b01e6174d8eb29ed66b82f",
  "fingerprint_version": 2,
  "provider": "grok",
  "endpoint": "openrouter/chat-completions",
  "model": "x-ai/grok-imagine-image-2.0",
  "pack": "24.202608252115",
  "source_file": "S5sh8_sel.png",
  "source_sha256": "8a5d2d563ddf2d64dd2e0c219b560217d754091948073f498bf7593890b55e9f",
  "file": "S5sh8_cine.png",
  "staged_sha256": "c6106f10cb518e12e3d8d2887dd3e9edb457c511ecc6e3cd8467ea25aef3a94c",
  "latency_ms": 16858
 },
 "S5sh11::signage": {
  "fp": "6c6be76348ebf7fe",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S5sh11": {
  "input_fingerprint": "7627d1390bb4e4c5",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — Wyoming, United States, 2419; all people are English-speaking Americans unless stated. TIME OF DAY (lock): twilight.\n\nSHOT TEXT (authoritative, Korean): 천장의 갈라진 틈새 사이로 굵은 돌가루가 허공에 쏟아져 내리는 찰나\n\nLOCATION (lock): Inside the mountain contact post's underground control room, beneath a cracked section of the stone ceiling. Dust pours down into the room after an impact outside. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Cracked ceiling gap (Open along a fresh crack while stone dust falls through it) — The underside of the ceiling and the opening within it face the steeply upward camera; used as Primary impact point held across the upper portion of the frame; Falling stone dust (Descending in its heaviest stream); used as Creates the central vertical movement and marks the instant of structural danger.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Low-key ambient interior illumination gives the falling stone dust enough separation to register against the ceiling without introducing a new visible source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The Mara message remains preserved in the wired phone system, and the contact station stays sealed under red light. Fresh cracks are visible overhead as coarse stone dust pours down. Coarse stone dust falls visibly from newly cracked gaps in the ceiling.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — Wyoming, United States, 2419; all people are English-speaking Americans unless stated. TIME OF DAY (lock): twilight.\n\nSHOT TEXT (authoritative, Korean): 천장의 갈라진 틈새 사이로 굵은 돌가루가 허공에 쏟아져 내리는 찰나\n\nLOCATION (lock): Inside the mountain contact post's underground control room, beneath a cracked section of the stone ceiling. Dust pours down into the room after an impact outside. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Cracked ceiling gap (Open along a fresh crack while stone dust falls through it) — The underside of the ceiling and the opening within it face the steeply upward camera; used as Primary impact point held across the upper portion of the frame; Falling stone dust (Descending in its heaviest stream); used as Creates the central vertical movement and marks the instant of structural danger.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Low-key ambient interior illumination gives the falling stone dust enough separation to register against the ceiling without introducing a new visible source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The Mara message remains preserved in the wired phone system, and the contact station stays sealed under red light. Fresh cracks are visible overhead as coarse stone dust pours down. Coarse stone dust falls visibly from newly cracked gaps in the ceiling.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — Wyoming, United States, 2419; all people are English-speaking Americans unless stated. TIME OF DAY (lock): twilight.\n\nSHOT TEXT (authoritative, Korean): 천장의 갈라진 틈새 사이로 굵은 돌가루가 허공에 쏟아져 내리는 찰나\n\nLOCATION (lock): Inside the mountain contact post's underground control room, beneath a cracked section of the stone ceiling. Dust pours down into the room after an impact outside. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- KEY BACKGROUND ELEMENTS: Cracked ceiling gap (Open along a fresh crack while stone dust falls through it) — The underside of the ceiling and the opening within it face the steeply upward camera; used as Primary impact point held across the upper portion of the frame; Falling stone dust (Descending in its heaviest stream); used as Creates the central vertical movement and marks the instant of structural danger.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Low-key ambient interior illumination gives the falling stone dust enough separation to register against the ceiling without introducing a new visible source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The Mara message remains preserved in the wired phone system, and the contact station stays sealed under red light. Fresh cracks are visible overhead as coarse stone dust pours down. Coarse stone dust falls visibly from newly cracked gaps in the ceiling.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "카메라는 가파르게 위를 향하여 천장의 갈라진 틈을 바라보고 있으며, 돌가루와 파편이 아래쪽의 장비 위로 곧게 쏟아져 내림.",
    "built_space": "붉은 조명이 비추는 콘크리트 천장에 큰 균열이 생겨 있으며, 하단에는 통제실의 서버 랙으로 보이는 장비 일부가 위치함.",
    "entities": "천장의 갈라진 틈, 쏟아지는 굵은 돌가루와 파편이 명확히 묘사됨. 인물은 없음.",
    "hard_violations": [],
    "physics": "돌가루와 파편들이 중력에 의해 자연스럽게 낙하하고 있음."
   },
   {
    "label": "B",
    "direction": "카메라는 수평 방향으로 방 전체를 조망하고 있으며, 천장의 틈에서 돌가루가 바닥으로 떨어지고 있음.",
    "built_space": "레퍼런스의 통제실 내부 요소들을 포함하고 있으나, 천장 높이가 물리적으로 불가능할 정도로 낮아져 문의 상단 부분을 가리고 있음.",
    "entities": "통제실 내부의 책상, 모니터, 문, 원통형 구조물 등이 보이며 천장에서 먼지와 돌가루가 떨어짐. 인물은 없음.",
    "hard_violations": [
     "[gemini-pro] 지시된 카메라 앵글(위로 향함)과 프레이밍(클로즈업)을 완전히 무시함",
     "[gemini-pro] 천장의 높이가 비정상적으로 낮아져 레퍼런스 공간의 물리적 구조를 훼손함 (문 상단이 천장에 잘림)"
    ],
    "physics": "먼지와 파편들이 중력에 의해 바닥으로 떨어지고 있음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "지시된 클로즈업 프레이밍과 위를 올려다보는 카메라 앵글을 정확히 구현하여 천장의 균열과 쏟아지는 돌가루를 생생하게 포착했습니다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "레퍼런스 이미지의 카메라 구도를 그대로 베끼지 말라는 지시를 어기고 와이드 샷으로 연출했으며, 천장이 비정상적으로 낮아져 공간 구조가 왜곡되었습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "카메라는 가파르게 위를 향하여 천장의 갈라진 틈을 바라보고 있으며, 돌가루와 파편이 아래쪽의 장비 위로 곧게 쏟아져 내림.",
        "built_space": "붉은 조명이 비추는 콘크리트 천장에 큰 균열이 생겨 있으며, 하단에는 통제실의 서버 랙으로 보이는 장비 일부가 위치함.",
        "entities": "천장의 갈라진 틈, 쏟아지는 굵은 돌가루와 파편이 명확히 묘사됨. 인물은 없음.",
        "hard_violations": [],
        "physics": "돌가루와 파편들이 중력에 의해 자연스럽게 낙하하고 있음."
       },
       {
        "label": "B",
        "direction": "카메라는 수평 방향으로 방 전체를 조망하고 있으며, 천장의 틈에서 돌가루가 바닥으로 떨어지고 있음.",
        "built_space": "레퍼런스의 통제실 내부 요소들을 포함하고 있으나, 천장 높이가 물리적으로 불가능할 정도로 낮아져 문의 상단 부분을 가리고 있음.",
        "entities": "통제실 내부의 책상, 모니터, 문, 원통형 구조물 등이 보이며 천장에서 먼지와 돌가루가 떨어짐. 인물은 없음.",
        "hard_violations": [
         "지시된 카메라 앵글(위로 향함)과 프레이밍(클로즈업)을 완전히 무시함",
         "천장의 높이가 비정상적으로 낮아져 레퍼런스 공간의 물리적 구조를 훼손함 (문 상단이 천장에 잘림)"
        ],
        "physics": "먼지와 파편들이 중력에 의해 바닥으로 떨어지고 있음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "지시된 클로즈업 프레이밍과 위를 올려다보는 카메라 앵글을 정확히 구현하여 천장의 균열과 쏟아지는 돌가루를 생생하게 포착했습니다."
       },
       {
        "label": "B",
        "score": 2,
        "verdict_ko": "레퍼런스 이미지의 카메라 구도를 그대로 베끼지 말라는 지시를 어기고 와이드 샷으로 연출했으며, 천장이 비정상적으로 낮아져 공간 구조가 왜곡되었습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "카메라는 가파르게 위를 향하여 천장의 갈라진 틈을 바라보고 있으며, 돌가루와 파편이 아래쪽의 장비 위로 곧게 쏟아져 내림.",
        "built_space": "붉은 조명이 비추는 콘크리트 천장에 큰 균열이 생겨 있으며, 하단에는 통제실의 서버 랙으로 보이는 장비 일부가 위치함.",
        "entities": "천장의 갈라진 틈, 쏟아지는 굵은 돌가루와 파편이 명확히 묘사됨. 인물은 없음.",
        "hard_violations": [],
        "physics": "돌가루와 파편들이 중력에 의해 자연스럽게 낙하하고 있음."
       },
       {
        "label": "B",
        "direction": "카메라는 수평 방향으로 방 전체를 조망하고 있으며, 천장의 틈에서 돌가루가 바닥으로 떨어지고 있음.",
        "built_space": "레퍼런스의 통제실 내부 요소들을 포함하고 있으나, 천장 높이가 물리적으로 불가능할 정도로 낮아져 문의 상단 부분을 가리고 있음.",
        "entities": "통제실 내부의 책상, 모니터, 문, 원통형 구조물 등이 보이며 천장에서 먼지와 돌가루가 떨어짐. 인물은 없음.",
        "hard_violations": [
         "지시된 카메라 앵글(위로 향함)과 프레이밍(클로즈업)을 완전히 무시함",
         "천장의 높이가 비정상적으로 낮아져 레퍼런스 공간의 물리적 구조를 훼손함 (문 상단이 천장에 잘림)"
        ],
        "physics": "먼지와 파편들이 중력에 의해 바닥으로 떨어지고 있음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "가파른 상향 클로즈업으로 천장 균열과 중앙의 가장 굵은 낙진 순간을 정확히 포착해, 장소 식별력이 다소 약해도 핵심 촬영 지시를 가장 충실히 실현한다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "낙진과 기준 통제실은 잘 재현했지만 카메라가 방 전체를 넓게 내려다보는 구도여서 요구된 천장 중심의 가파른 상향 클로즈업과 크게 어긋난다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "굵은 돌가루와 파편이 화면 상단의 긴 천장 틈에서 중앙 하단의 실내 바닥 방향으로 수직 낙하한다. 여행 중인 물체의 목표는 균열 아래의 방 중앙이며, 사람의 시선이나 사용 중인 무기 조준은 없다. 원형 설비에 기대어진 총기 같은 물체는 사용되지 않고 위쪽으로 기울어져 있다.",
        "built_space": "지하 통제실을 넓게 보여 주며, 왼쪽에 봉인문 1개와 암벽 벽체, 뒤쪽에 콘솔 벽과 의자 1개, 오른쪽에 작업대 및 투명 보관함, 중앙에 원통형 설비 1개가 보인다. 기준 사진의 재료와 주요 배치는 상당히 잘 이어지지만, 천장과 균열은 화면 위쪽의 일부만 차지하고 카메라는 가파르게 올려다보기보다 높은 위치에서 방 안쪽을 거의 수평 또는 약간 하향으로 본다. 반사면이나 거울은 없어 불가능한 반사는 없다.",
        "entities": "새로 벌어진 천장 균열, 굵은 돌가루, 붉은 저조도 통제실, 봉인된 시설이 모두 보이며 사람·얼굴·신체는 없다. 유선 전화의 메시지 자체는 이 프레임에서 확인되지 않지만 올바른 프레이밍이 가린 항목은 아니다. 눈에 확실히 읽히는 글자는 없다.",
        "hard_violations": [],
        "physics": "돌가루와 여러 크기의 파편은 천장 균열에서 분리되어 중력으로 아래로 떨어지며, 출발점과 예상 낙하지점이 모두 성립한다. 공중의 파편은 낙하 동작으로 설명되며 지지 없이 정지해 떠 있는 물체는 없다."
       },
       {
        "label": "B",
        "direction": "돌가루와 작은 돌들이 화면 상단의 갈라진 개구부에서 화면 중앙을 지나 아래쪽 장비 방향으로 곧게 낙하한다. 낙하의 출발점은 신선한 천장 균열이고 목표 경로는 바로 아래 통제실 내부다. 사람의 시선, 조준 중인 무기 또는 다른 방향성 소품은 없다.",
        "built_space": "가파르게 위를 올려다본 클로즈업으로, 콘크리트 천장과 길게 벌어진 균열 1개가 상부 대부분을 차지하고 하단에는 통제실 장비 랙 상단 1개만 잘려 보인다. 기준 장소의 붉은 비상 조명, 노후 콘크리트와 전자 장비 재질은 이어지지만, 넓은 방의 고정 설비들은 의도적인 클로즈업 때문에 제외되어 장소 식별력은 A보다 약하다. 반사면이나 광학적으로 불가능한 반사는 없다.",
        "entities": "핵심 대상인 갈라진 천장 틈과 가장 무거운 중앙 돌가루 흐름이 명확하다. 붉은 조명 아래의 지하 통제실 장비 일부가 보이고 사람·얼굴·신체는 전혀 없다. 전화 메시지는 보이지 않으며 읽을 수 있는 글자나 오버레이도 없다.",
        "hard_violations": [],
        "physics": "굵은 가루와 파편은 균열 가장자리에서 떨어져 중력 방향으로 하강하고 있으며, 분리 지점과 아래쪽 착지 경로가 분명하다. 일부 돌이 공중에 있지만 모두 낙하 흐름 안에 있어 물리적으로 설명되며, 지지 없이 정지한 물체는 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "가파른 상향 클로즈업으로 천장 균열과 중앙의 가장 굵은 낙진 순간을 정확히 포착해, 장소 식별력이 다소 약해도 핵심 촬영 지시를 가장 충실히 실현한다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "낙진과 기준 통제실은 잘 재현했지만 카메라가 방 전체를 넓게 내려다보는 구도여서 요구된 천장 중심의 가파른 상향 클로즈업과 크게 어긋난다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "굵은 돌가루와 파편이 화면 상단의 긴 천장 틈에서 중앙 하단의 실내 바닥 방향으로 수직 낙하한다. 여행 중인 물체의 목표는 균열 아래의 방 중앙이며, 사람의 시선이나 사용 중인 무기 조준은 없다. 원형 설비에 기대어진 총기 같은 물체는 사용되지 않고 위쪽으로 기울어져 있다.",
        "built_space": "지하 통제실을 넓게 보여 주며, 왼쪽에 봉인문 1개와 암벽 벽체, 뒤쪽에 콘솔 벽과 의자 1개, 오른쪽에 작업대 및 투명 보관함, 중앙에 원통형 설비 1개가 보인다. 기준 사진의 재료와 주요 배치는 상당히 잘 이어지지만, 천장과 균열은 화면 위쪽의 일부만 차지하고 카메라는 가파르게 올려다보기보다 높은 위치에서 방 안쪽을 거의 수평 또는 약간 하향으로 본다. 반사면이나 거울은 없어 불가능한 반사는 없다.",
        "entities": "새로 벌어진 천장 균열, 굵은 돌가루, 붉은 저조도 통제실, 봉인된 시설이 모두 보이며 사람·얼굴·신체는 없다. 유선 전화의 메시지 자체는 이 프레임에서 확인되지 않지만 올바른 프레이밍이 가린 항목은 아니다. 눈에 확실히 읽히는 글자는 없다.",
        "hard_violations": [],
        "physics": "돌가루와 여러 크기의 파편은 천장 균열에서 분리되어 중력으로 아래로 떨어지며, 출발점과 예상 낙하지점이 모두 성립한다. 공중의 파편은 낙하 동작으로 설명되며 지지 없이 정지해 떠 있는 물체는 없다."
       },
       {
        "label": "A",
        "direction": "돌가루와 작은 돌들이 화면 상단의 갈라진 개구부에서 화면 중앙을 지나 아래쪽 장비 방향으로 곧게 낙하한다. 낙하의 출발점은 신선한 천장 균열이고 목표 경로는 바로 아래 통제실 내부다. 사람의 시선, 조준 중인 무기 또는 다른 방향성 소품은 없다.",
        "built_space": "가파르게 위를 올려다본 클로즈업으로, 콘크리트 천장과 길게 벌어진 균열 1개가 상부 대부분을 차지하고 하단에는 통제실 장비 랙 상단 1개만 잘려 보인다. 기준 장소의 붉은 비상 조명, 노후 콘크리트와 전자 장비 재질은 이어지지만, 넓은 방의 고정 설비들은 의도적인 클로즈업 때문에 제외되어 장소 식별력은 A보다 약하다. 반사면이나 광학적으로 불가능한 반사는 없다.",
        "entities": "핵심 대상인 갈라진 천장 틈과 가장 무거운 중앙 돌가루 흐름이 명확하다. 붉은 조명 아래의 지하 통제실 장비 일부가 보이고 사람·얼굴·신체는 전혀 없다. 전화 메시지는 보이지 않으며 읽을 수 있는 글자나 오버레이도 없다.",
        "hard_violations": [],
        "physics": "굵은 가루와 파편은 균열 가장자리에서 떨어져 중력 방향으로 하강하고 있으며, 분리 지점과 아래쪽 착지 경로가 분명하다. 일부 돌이 공중에 있지만 모두 낙하 흐름 안에 있어 물리적으로 설명되며, 지지 없이 정지한 물체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt"
   ],
   "normalized": {
    "A": 2.0,
    "B": 0.889
   },
   "adjusted": {
    "A": 2.0,
    "B": 0.639
   },
   "violations": {
    "B": [
     "[gemini-pro] 지시된 카메라 앵글(위로 향함)과 프레이밍(클로즈업)을 완전히 무시함",
     "[gemini-pro] 천장의 높이가 비정상적으로 낮아져 레퍼런스 공간의 물리적 구조를 훼손함 (문 상단이 천장에 잘림)"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 639
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "지시된 클로즈업 프레이밍과 위를 올려다보는 카메라 앵글을 정확히 구현하여 천장의 균열과 쏟아지는 돌가루를 생생하게 포착했습니다."
   },
   {
    "label": "B",
    "score": 639,
    "verdict_ko": "레퍼런스 이미지의 카메라 구도를 그대로 베끼지 말라는 지시를 어기고 와이드 샷으로 연출했으며, 천장이 비정상적으로 낮아져 공간 구조가 왜곡되었습니다.  ★위반: [gemini-pro] 지시된 카메라 앵글(위로 향함)과 프레이밍(클로즈업)을 완전히 무시함 / [gemini-pro] 천장의 높이가 비정상적으로 낮아져 레퍼런스 공간의 물리적 구조를 훼손함 (문 상단이 천장에 잘림)"
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/37154490-a4ed-4b37-9092-489c054d8a72/episodes/cc1d3d0b-626e-4f20-bc33-f5f9e2c6ee7e/images/background_chain/L11B02.png",
    "asset_id": "eb7f72c2-429d-4eb2-880a-e3c8b93c34f5",
    "role": "location_plate"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": true,
  "shot_run_uid": "06a9bdcd-d531-7d92-986f-5dabc7178388",
  "ref_mode": "플레이트만 (배경 전용)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S5sh11::cine": {
  "applied": true,
  "attempted_at": "2026-09-05T09:12:53.954736+00:00",
  "fingerprint": "755999157cca89d9af7117de062ff16ef40411190cfe67f5df90cc0dab19d701",
  "fingerprint_version": 2,
  "provider": "grok",
  "endpoint": "openrouter/chat-completions",
  "model": "x-ai/grok-imagine-image-2.0",
  "pack": "24.202608252115",
  "source_file": "S5sh11_sel.png",
  "source_sha256": "07d222cd11dc6ceb13e42914f7678816a849bc7c6d930f31b5166f666d8c6762",
  "file": "S5sh11_cine.png",
  "staged_sha256": "b38af22881fc922495a596351126d470c1bfb6b29aafb28b014542b7cdd89eb0",
  "latency_ms": 12327
 },
 "S6sh11::signage": {
  "fp": "087d1275dcfc07d9",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S6sh11": {
  "input_fingerprint": "9c0d321720693ca0",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — Wyoming, United States, 2419; all people are English-speaking Americans unless stated. TIME OF DAY (lock): twilight.\n\nSHOT TEXT (authoritative, Korean): 닫힌 바위문 한가운데에 눈부시게 빛나는 원형의 선이 그어진 찰나\n\nLOCATION (lock): Inside the mountain contact post, directly before the sealed rock entrance door. A brilliant circular line forms across the door as its material begins disappearing from the edge. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: closed rock door and circular line in the middle-center of the frame, midground; narrow tunnel entrance edge in the middle-left of the frame, foreground.\n- KEY BACKGROUND ELEMENTS: Closed rock door with circular line (Closed, with a brilliant circular line visible at its center) — The door's front face is viewed obliquely from a low three-quarter angle, keeping the centered circular line clearly visible; used as Primary focal plane and immediate threat marker; Narrow tunnel entrance (Open and only wide enough for one person to crawl through) — One side of the entrance faces the camera beside its low position while the passage recedes away from the door; used as Provides a restrained foreground edge and establishes the escape route for the following reversal.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The brilliant circular line supplies a severe local highlight against the scene's otherwise cool, muted, low-key illumination.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The concealed wall is open, exposing a one-person crawl tunnel. The phone has been disconnected and shut down at 1% battery, while a brilliant circular line glows at the center of the closed rock door and material vanishes from its edge. A brilliant circular trace appears on the inner face of the rock door, with visible matter disappearing from the rim inward.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — Wyoming, United States, 2419; all people are English-speaking Americans unless stated. TIME OF DAY (lock): twilight.\n\nSHOT TEXT (authoritative, Korean): 닫힌 바위문 한가운데에 눈부시게 빛나는 원형의 선이 그어진 찰나\n\nLOCATION (lock): Inside the mountain contact post, directly before the sealed rock entrance door. A brilliant circular line forms across the door as its material begins disappearing from the edge. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: closed rock door and circular line in the middle-center of the frame, midground; narrow tunnel entrance edge in the middle-left of the frame, foreground.\n- KEY BACKGROUND ELEMENTS: Closed rock door with circular line (Closed, with a brilliant circular line visible at its center) — The door's front face is viewed obliquely from a low three-quarter angle, keeping the centered circular line clearly visible; used as Primary focal plane and immediate threat marker; Narrow tunnel entrance (Open and only wide enough for one person to crawl through) — One side of the entrance faces the camera beside its low position while the passage recedes away from the door; used as Provides a restrained foreground edge and establishes the escape route for the following reversal.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The brilliant circular line supplies a severe local highlight against the scene's otherwise cool, muted, low-key illumination.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The concealed wall is open, exposing a one-person crawl tunnel. The phone has been disconnected and shut down at 1% battery, while a brilliant circular line glows at the center of the closed rock door and material vanishes from its edge. A brilliant circular trace appears on the inner face of the rock door, with visible matter disappearing from the rim inward.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — Wyoming, United States, 2419; all people are English-speaking Americans unless stated. TIME OF DAY (lock): twilight.\n\nSHOT TEXT (authoritative, Korean): 닫힌 바위문 한가운데에 눈부시게 빛나는 원형의 선이 그어진 찰나\n\nLOCATION (lock): Inside the mountain contact post, directly before the sealed rock entrance door. A brilliant circular line forms across the door as its material begins disappearing from the edge. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: closed rock door and circular line in the middle-center of the frame, midground; narrow tunnel entrance edge in the middle-left of the frame, foreground.\n- KEY BACKGROUND ELEMENTS: Closed rock door with circular line (Closed, with a brilliant circular line visible at its center) — The door's front face is viewed obliquely from a low three-quarter angle, keeping the centered circular line clearly visible; used as Primary focal plane and immediate threat marker; Narrow tunnel entrance (Open and only wide enough for one person to crawl through) — One side of the entrance faces the camera beside its low position while the passage recedes away from the door; used as Provides a restrained foreground edge and establishes the escape route for the following reversal.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The brilliant circular line supplies a severe local highlight against the scene's otherwise cool, muted, low-key illumination.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The concealed wall is open, exposing a one-person crawl tunnel. The phone has been disconnected and shut down at 1% battery, while a brilliant circular line glows at the center of the closed rock door and material vanishes from its edge. A brilliant circular trace appears on the inner face of the rock door, with visible matter disappearing from the rim inward.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "사람이나 무기, 조준되는 대상이 화면에 존재하지 않음.",
    "built_space": "통제실 중앙에 바위문이 있고 왼쪽에 안으로 이어지는 터널이 배치됨. 좁은 크기가 아닌 넓은 개구부 형태로 묘사되었으나, 전체 조명은 지시된 차갑고 어두운 톤으로 잘 수정됨.",
    "entities": "닫힌 바위문 표면에 빛나는 얇은 원형 선이 묘사됨.",
    "hard_violations": [
     "[gpt] 문 왼쪽 표지판에 읽을 수 있는 영문이 노출되어 ‘읽을 수 있는 글자 금지’를 위반한다."
    ],
    "physics": "빛나는 선이 문의 질감과 구조 위에 자연스럽게 맺혀 있으며, 물리적으로 어긋나거나 공중에 뜬 객체는 없음."
   },
   {
    "label": "B",
    "direction": "사람이나 무기, 조준되는 대상이 화면에 존재하지 않음.",
    "built_space": "터널 또는 환기구 내부에서 바깥의 바위문과 통제실을 내다보는 전경 구도로 렌더링됨. 실내 조명은 지시사항을 무시하고 레퍼런스의 붉은빛을 그대로 유지함.",
    "entities": "바위문에 얇은 선 대신 두꺼운 연기와 불꽃을 뿜는 고리가 형성되어 있음.",
    "hard_violations": [
     "[gpt] 오른쪽 제어 화면에 ‘FALSE EXIT’ 등 읽을 수 있는 글자가 노출되어 ‘읽을 수 있는 글자 금지’를 위반한다."
    ],
    "physics": "불꽃 고리가 문 표면에서 발생하고 있으며, 지지되지 않고 떠 있는 물체나 물리적 오류는 발견되지 않음."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "차갑고 절제된 조명과 바위문에 그어진 빛나는 원형 선을 잘 표현했으나, 좌측 터널이 기어갈 수 있는 좁은 크기가 아니며 텍스트 비가독성 지시를 어겼습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "카메라가 터널 내부로 들어가 구도 지시를 벗어났고, 조명이 레퍼런스의 붉은 톤에 머물러 있으며 모니터의 텍스트가 선명하게 읽힙니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "사람이나 무기, 조준되는 대상이 화면에 존재하지 않음.",
        "built_space": "통제실 중앙에 바위문이 있고 왼쪽에 안으로 이어지는 터널이 배치됨. 좁은 크기가 아닌 넓은 개구부 형태로 묘사되었으나, 전체 조명은 지시된 차갑고 어두운 톤으로 잘 수정됨.",
        "entities": "닫힌 바위문 표면에 빛나는 얇은 원형 선이 묘사됨.",
        "hard_violations": [],
        "physics": "빛나는 선이 문의 질감과 구조 위에 자연스럽게 맺혀 있으며, 물리적으로 어긋나거나 공중에 뜬 객체는 없음."
       },
       {
        "label": "B",
        "direction": "사람이나 무기, 조준되는 대상이 화면에 존재하지 않음.",
        "built_space": "터널 또는 환기구 내부에서 바깥의 바위문과 통제실을 내다보는 전경 구도로 렌더링됨. 실내 조명은 지시사항을 무시하고 레퍼런스의 붉은빛을 그대로 유지함.",
        "entities": "바위문에 얇은 선 대신 두꺼운 연기와 불꽃을 뿜는 고리가 형성되어 있음.",
        "hard_violations": [],
        "physics": "불꽃 고리가 문 표면에서 발생하고 있으며, 지지되지 않고 떠 있는 물체나 물리적 오류는 발견되지 않음."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "차갑고 절제된 조명과 바위문에 그어진 빛나는 원형 선을 잘 표현했으나, 좌측 터널이 기어갈 수 있는 좁은 크기가 아니며 텍스트 비가독성 지시를 어겼습니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "카메라가 터널 내부로 들어가 구도 지시를 벗어났고, 조명이 레퍼런스의 붉은 톤에 머물러 있으며 모니터의 텍스트가 선명하게 읽힙니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "사람이나 무기, 조준되는 대상이 화면에 존재하지 않음.",
        "built_space": "통제실 중앙에 바위문이 있고 왼쪽에 안으로 이어지는 터널이 배치됨. 좁은 크기가 아닌 넓은 개구부 형태로 묘사되었으나, 전체 조명은 지시된 차갑고 어두운 톤으로 잘 수정됨.",
        "entities": "닫힌 바위문 표면에 빛나는 얇은 원형 선이 묘사됨.",
        "hard_violations": [],
        "physics": "빛나는 선이 문의 질감과 구조 위에 자연스럽게 맺혀 있으며, 물리적으로 어긋나거나 공중에 뜬 객체는 없음."
       },
       {
        "label": "B",
        "direction": "사람이나 무기, 조준되는 대상이 화면에 존재하지 않음.",
        "built_space": "터널 또는 환기구 내부에서 바깥의 바위문과 통제실을 내다보는 전경 구도로 렌더링됨. 실내 조명은 지시사항을 무시하고 레퍼런스의 붉은빛을 그대로 유지함.",
        "entities": "바위문에 얇은 선 대신 두꺼운 연기와 불꽃을 뿜는 고리가 형성되어 있음.",
        "hard_violations": [],
        "physics": "불꽃 고리가 문 표면에서 발생하고 있으며, 지지되지 않고 떠 있는 물체나 물리적 오류는 발견되지 않음."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "문과 원형 광선을 화면 중앙에 두고 왼쪽 통로가 뒤로 이어지는 핵심 배치는 가장 정확하지만, 통로가 기어가야 할 크기가 아닌 보행 가능한 출입구이며 표지 글자도 읽힌다."
       },
       {
        "label": "A",
        "score": 4,
        "verdict_ko": "낮은 사선 시점과 문 표면의 밝은 원형 흔적은 좋지만, 카메라가 통로 안에 들어가 검은 테두리가 화면을 에워싸고 문도 중앙보다 왼쪽에 치우쳐 지정 구도를 놓쳤으며 화면 글자도 읽힌다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "사람, 시선, 무기 또는 겨냥된 물체는 없다. 원형 광선은 닫힌 문의 정면에 붙어 있고 특정 외부 표적을 향하지 않는다. 카메라는 어두운 통로 내부에서 통제실과 문 쪽을 내다보는 방향이다.",
        "built_space": "닫힌 암반 출입문 1개가 중경의 중앙 왼쪽에 있고, 그 표면 중앙에 원형 광선 1개가 있다. 암실 같은 통로 개구부 1개의 가장자리가 왼쪽뿐 아니라 위·아래까지 크게 둘러싸며 카메라가 통로 안에 놓인 형태다. 문 오른쪽에는 장비 캐비닛 여러 대와 제어 화면 1개, 의자 1개가 보인다. 요구된 ‘중간 왼쪽의 제한된 전경 가장자리’보다 개구부가 프레임 전체의 마스크처럼 강하고, 통로가 문 옆에서 반대 방향으로 후퇴하는 구조도 분명하지 않다.",
        "entities": "사람과 신체 일부는 없다. 닫힌 문은 암반에 매입된 금속 전면의 봉인문으로 장소 사진과 대체로 일치한다. 광선은 내부가 문 재질로 남는 밝은 고리이며, 연기와 입자가 가장자리 소실을 암시한다. 휴대전화는 보이지 않는다. 오른쪽 화면에는 ‘FALSE EXIT’ 등 식별 가능한 영문이 남아 있다.",
        "hard_violations": [
         "오른쪽 제어 화면에 ‘FALSE EXIT’ 등 읽을 수 있는 글자가 노출되어 ‘읽을 수 있는 글자 금지’를 위반한다."
        ],
        "physics": "문과 장비는 벽과 바닥에 고정되어 있다. 원형 광선은 문 표면을 따라 형성된 발광 흔적으로 읽히며 공중에 떠 있지 않는다. 입자와 연기는 문 가장자리의 소실 반응에서 발생하는 것으로 보인다. 지지되지 않은 사람이나 물체는 없다."
       },
       {
        "label": "B",
        "direction": "사람, 시선, 무기 또는 겨냥된 물체는 없다. 왼쪽 통로는 화면 안쪽 왼편으로 후퇴하고, 원형 광선은 중앙의 닫힌 문 표면에 형성되어 있다.",
        "built_space": "중경 중앙에 닫힌 암반 출입문 1개와 그 중심의 원형 광선 1개가 있다. 왼쪽 전경에는 열린 통로 입구 1개가 있고 통로가 문에서 멀어지는 방향으로 후퇴한다. 오른쪽에는 장비 캐비닛과 제어실 일부, 원통형 탁자 일부가 보인다. 문과 통로의 화면 배치는 요구에 가깝지만, 왼쪽 입구는 성인이 서서 지나갈 수 있을 정도로 높아 ‘한 사람이 기어서만 통과할 수 있는 좁은 터널’의 규격과 맞지 않는다.",
        "entities": "사람과 신체 일부는 없다. 장소 사진과 같은 암반 벽, 매입형 금속 봉인문, 제어 장비가 확인된다. 원형 광선은 내부에 문 표면이 그대로 보이는 선형 고리이며, 불꽃 입자가 재질 소실을 암시한다. 휴대전화는 보이지 않는다. 문 왼쪽 표지판에 여러 줄의 영문이 판독 가능한 상태로 남아 있다.",
        "hard_violations": [
         "문 왼쪽 표지판에 읽을 수 있는 영문이 노출되어 ‘읽을 수 있는 글자 금지’를 위반한다."
        ],
        "physics": "문은 암반 벽에 고정되고 장비와 탁자는 바닥에 지지된다. 빛나는 고리는 문 표면에 붙은 흔적이며, 불꽃과 입자는 문 표면의 소실 지점에서 방출되는 것으로 보인다. 공중에 떠 있거나 지지되지 않은 사람·물체는 없다."
       }
      ],
      "all_candidates_fail": true
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "문과 원형 광선을 화면 중앙에 두고 왼쪽 통로가 뒤로 이어지는 핵심 배치는 가장 정확하지만, 통로가 기어가야 할 크기가 아닌 보행 가능한 출입구이며 표지 글자도 읽힌다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "낮은 사선 시점과 문 표면의 밝은 원형 흔적은 좋지만, 카메라가 통로 안에 들어가 검은 테두리가 화면을 에워싸고 문도 중앙보다 왼쪽에 치우쳐 지정 구도를 놓쳤으며 화면 글자도 읽힌다."
       }
      ],
      "all_candidates_fail": true,
      "readings": [
       {
        "label": "B",
        "direction": "사람, 시선, 무기 또는 겨냥된 물체는 없다. 원형 광선은 닫힌 문의 정면에 붙어 있고 특정 외부 표적을 향하지 않는다. 카메라는 어두운 통로 내부에서 통제실과 문 쪽을 내다보는 방향이다.",
        "built_space": "닫힌 암반 출입문 1개가 중경의 중앙 왼쪽에 있고, 그 표면 중앙에 원형 광선 1개가 있다. 암실 같은 통로 개구부 1개의 가장자리가 왼쪽뿐 아니라 위·아래까지 크게 둘러싸며 카메라가 통로 안에 놓인 형태다. 문 오른쪽에는 장비 캐비닛 여러 대와 제어 화면 1개, 의자 1개가 보인다. 요구된 ‘중간 왼쪽의 제한된 전경 가장자리’보다 개구부가 프레임 전체의 마스크처럼 강하고, 통로가 문 옆에서 반대 방향으로 후퇴하는 구조도 분명하지 않다.",
        "entities": "사람과 신체 일부는 없다. 닫힌 문은 암반에 매입된 금속 전면의 봉인문으로 장소 사진과 대체로 일치한다. 광선은 내부가 문 재질로 남는 밝은 고리이며, 연기와 입자가 가장자리 소실을 암시한다. 휴대전화는 보이지 않는다. 오른쪽 화면에는 ‘FALSE EXIT’ 등 식별 가능한 영문이 남아 있다.",
        "hard_violations": [
         "오른쪽 제어 화면에 ‘FALSE EXIT’ 등 읽을 수 있는 글자가 노출되어 ‘읽을 수 있는 글자 금지’를 위반한다."
        ],
        "physics": "문과 장비는 벽과 바닥에 고정되어 있다. 원형 광선은 문 표면을 따라 형성된 발광 흔적으로 읽히며 공중에 떠 있지 않는다. 입자와 연기는 문 가장자리의 소실 반응에서 발생하는 것으로 보인다. 지지되지 않은 사람이나 물체는 없다."
       },
       {
        "label": "A",
        "direction": "사람, 시선, 무기 또는 겨냥된 물체는 없다. 왼쪽 통로는 화면 안쪽 왼편으로 후퇴하고, 원형 광선은 중앙의 닫힌 문 표면에 형성되어 있다.",
        "built_space": "중경 중앙에 닫힌 암반 출입문 1개와 그 중심의 원형 광선 1개가 있다. 왼쪽 전경에는 열린 통로 입구 1개가 있고 통로가 문에서 멀어지는 방향으로 후퇴한다. 오른쪽에는 장비 캐비닛과 제어실 일부, 원통형 탁자 일부가 보인다. 문과 통로의 화면 배치는 요구에 가깝지만, 왼쪽 입구는 성인이 서서 지나갈 수 있을 정도로 높아 ‘한 사람이 기어서만 통과할 수 있는 좁은 터널’의 규격과 맞지 않는다.",
        "entities": "사람과 신체 일부는 없다. 장소 사진과 같은 암반 벽, 매입형 금속 봉인문, 제어 장비가 확인된다. 원형 광선은 내부에 문 표면이 그대로 보이는 선형 고리이며, 불꽃 입자가 재질 소실을 암시한다. 휴대전화는 보이지 않는다. 문 왼쪽 표지판에 여러 줄의 영문이 판독 가능한 상태로 남아 있다.",
        "hard_violations": [
         "문 왼쪽 표지판에 읽을 수 있는 영문이 노출되어 ‘읽을 수 있는 글자 금지’를 위반한다."
        ],
        "physics": "문은 암반 벽에 고정되고 장비와 탁자는 바닥에 지지된다. 빛나는 고리는 문 표면에 붙은 흔적이며, 불꽃과 입자는 문 표면의 소실 지점에서 방출되는 것으로 보인다. 공중에 떠 있거나 지지되지 않은 사람·물체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.467
   },
   "adjusted": {
    "A": 1.75,
    "B": 1.217
   },
   "violations": {
    "B": [
     "[gpt] 오른쪽 제어 화면에 ‘FALSE EXIT’ 등 읽을 수 있는 글자가 노출되어 ‘읽을 수 있는 글자 금지’를 위반한다."
    ],
    "A": [
     "[gpt] 문 왼쪽 표지판에 읽을 수 있는 영문이 노출되어 ‘읽을 수 있는 글자 금지’를 위반한다."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 1750,
   "B": 1217
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1750,
    "verdict_ko": "차갑고 절제된 조명과 바위문에 그어진 빛나는 원형 선을 잘 표현했으나, 좌측 터널이 기어갈 수 있는 좁은 크기가 아니며 텍스트 비가독성 지시를 어겼습니다.  ★위반: [gpt] 문 왼쪽 표지판에 읽을 수 있는 영문이 노출되어 ‘읽을 수 있는 글자 금지’를 위반한다."
   },
   {
    "label": "B",
    "score": 1217,
    "verdict_ko": "카메라가 터널 내부로 들어가 구도 지시를 벗어났고, 조명이 레퍼런스의 붉은 톤에 머물러 있으며 모니터의 텍스트가 선명하게 읽힙니다.  ★위반: [gpt] 오른쪽 제어 화면에 ‘FALSE EXIT’ 등 읽을 수 있는 글자가 노출되어 ‘읽을 수 있는 글자 금지’를 위반한다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/37154490-a4ed-4b37-9092-489c054d8a72/episodes/cc1d3d0b-626e-4f20-bc33-f5f9e2c6ee7e/images/background_chain/L11B02.png",
    "asset_id": "eb7f72c2-429d-4eb2-880a-e3c8b93c34f5",
    "role": "location_plate"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": true,
  "shot_run_uid": "06a9bdd2-26cf-712c-99e1-b11e3f2af0a9",
  "ref_mode": "플레이트만 (배경 전용)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S6sh11::cine": {
  "applied": true,
  "attempted_at": "2026-09-05T09:14:14.548132+00:00",
  "fingerprint": "4f9be7050fbf45aaf6667c6c0e3bdc01989d9860deaa7fb53fe4aa368465ba68",
  "fingerprint_version": 2,
  "provider": "grok",
  "endpoint": "openrouter/chat-completions",
  "model": "x-ai/grok-imagine-image-2.0",
  "pack": "24.202608252115",
  "source_file": "S6sh11_sel.png",
  "source_sha256": "58c6e789b48e3107e4534094cbfe195ffa61a930c12f74952d303b1607d3de62",
  "file": "S6sh11_cine.png",
  "staged_sha256": "97920b2492a5f62b30fe60363ff6cc7ee234c43f99349f620fa84cf53d2ef2b0",
  "latency_ms": 13615
 },
 "S6sh13::signage": {
  "fp": "15aeb4279c200df4",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S6sh13::bgfirst_bg": {
  "input_fingerprint": "7e10b3d830aa4c48",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 좁은 터널 안으로 몸을 구부린 채 기어 들어가는 윌마의 뒷모습\n\nLOCATION (lock): Inside the one-person crawl tunnel hidden behind a wall of the mountain contact post. The passage is so tight that occupants must bend low and crawl.\n\nTIME OF DAY (lock): twilight.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: 윌마 디어링 in the middle-center of the frame, midground, moves toward deeper tunnel passage; deeper tunnel passage in the upper-center of the frame, background.\n- KEY BACKGROUND ELEMENTS: Narrow escape tunnel (Open and restricted to single-person crawling) — The passage recedes directly away from camera along Wilma's path, with its close boundaries visible around her; used as Compresses the frame around Wilma and supplies the movement axis.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained, low-key illumination appropriate to the tunnel maintains a cool, muted image without specifying an additional source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 좁은 터널 안으로 몸을 구부린 채 기어 들어가는 윌마의 뒷모습\n\nLOCATION (lock): Inside the one-person crawl tunnel hidden behind a wall of the mountain contact post. The passage is so tight that occupants must bend low and crawl.\n\nTIME OF DAY (lock): twilight.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: 윌마 디어링 in the middle-center of the frame, midground, moves toward deeper tunnel passage; deeper tunnel passage in the upper-center of the frame, background.\n- KEY BACKGROUND ELEMENTS: Narrow escape tunnel (Open and restricted to single-person crawling) — The passage recedes directly away from camera along Wilma's path, with its close boundaries visible around her; used as Compresses the frame around Wilma and supplies the movement axis.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained, low-key illumination appropriate to the tunnel maintains a cool, muted image without specifying an additional source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/37154490-a4ed-4b37-9092-489c054d8a72/images/cc1d3d0b-626e-4f20-bc33-f5f9e2c6ee7e/scene/recipe/S6sh13__bgfirst_bg.png",
  "asset_id": "f813b3ae-a882-4493-8ed2-fc7fc34367b2",
  "input_asset_ids": [
   "dd5f575c-605d-4e9e-99b6-db87eee785b5",
   "a755db6c-33f5-4899-9c38-164f6c8fa798"
  ]
 },
 "S6sh13": {
  "input_fingerprint": "45c4ce9345e5ab82",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — Wyoming, United States, 2419; all people are English-speaking Americans unless stated. TIME OF DAY (lock): twilight.\n\nSHOT TEXT (authoritative, Korean): 좁은 터널 안으로 몸을 구부린 채 기어 들어가는 윌마의 뒷모습\n\nLOCATION (lock): Inside the one-person crawl tunnel hidden behind a wall of the mountain contact post. The passage is so tight that occupants must bend low and crawl. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: 윌마 디어링 in the middle-center of the frame, midground, moves toward deeper tunnel passage; deeper tunnel passage in the upper-center of the frame, background.\n- KEY BACKGROUND ELEMENTS: Narrow escape tunnel (Open and restricted to single-person crawling) — The passage recedes directly away from camera along Wilma's path, with its close boundaries visible around her; used as Compresses the frame around Wilma and supplies the movement axis.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained, low-key illumination appropriate to the tunnel maintains a cool, muted image without specifying an additional source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The narrow crawl tunnel remains exposed behind the opened wall. The phone screen is off, and the glowing circular breach continues consuming the rock door from its edges. 윌마 디어링: She bends low and crawls into the one-person tunnel. The luminous breach remains visible on the rock door as its material disappears from the perimeter.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 윌마 디어링 (미국인 여성, 20대 후반, 자연스러운 성인 여성 얼굴, 구체적으로 명시되지 않은 자연색 머리카락) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — Wyoming, United States, 2419; all people are English-speaking Americans unless stated. TIME OF DAY (lock): twilight.\n\nSHOT TEXT (authoritative, Korean): 좁은 터널 안으로 몸을 구부린 채 기어 들어가는 윌마의 뒷모습\n\nLOCATION (lock): Inside the one-person crawl tunnel hidden behind a wall of the mountain contact post. The passage is so tight that occupants must bend low and crawl. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: 윌마 디어링 in the middle-center of the frame, midground, moves toward deeper tunnel passage; deeper tunnel passage in the upper-center of the frame, background.\n- KEY BACKGROUND ELEMENTS: Narrow escape tunnel (Open and restricted to single-person crawling) — The passage recedes directly away from camera along Wilma's path, with its close boundaries visible around her; used as Compresses the frame around Wilma and supplies the movement axis.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained, low-key illumination appropriate to the tunnel maintains a cool, muted image without specifying an additional source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The narrow crawl tunnel remains exposed behind the opened wall. The phone screen is off, and the glowing circular breach continues consuming the rock door from its edges. 윌마 디어링: She bends low and crawls into the one-person tunnel. The luminous breach remains visible on the rock door as its material disappears from the perimeter.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 윌마 디어링 (미국인 여성, 20대 후반, 자연스러운 성인 여성 얼굴, 구체적으로 명시되지 않은 자연색 머리카락) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — Wyoming, United States, 2419; all people are English-speaking Americans unless stated. TIME OF DAY (lock): twilight.\n\nSHOT TEXT (authoritative, Korean): 좁은 터널 안으로 몸을 구부린 채 기어 들어가는 윌마의 뒷모습\n\nLOCATION (lock): Inside the one-person crawl tunnel hidden behind a wall of the mountain contact post. The passage is so tight that occupants must bend low and crawl. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: medium shot\n- FRAME LAYOUT: 윌마 디어링 in the middle-center of the frame, midground, moves toward deeper tunnel passage; deeper tunnel passage in the upper-center of the frame, background.\n- KEY BACKGROUND ELEMENTS: Narrow escape tunnel (Open and restricted to single-person crawling) — The passage recedes directly away from camera along Wilma's path, with its close boundaries visible around her; used as Compresses the frame around Wilma and supplies the movement axis.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Restrained, low-key illumination appropriate to the tunnel maintains a cool, muted image without specifying an additional source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The narrow crawl tunnel remains exposed behind the opened wall. The phone screen is off, and the glowing circular breach continues consuming the rock door from its edges. 윌마 디어링: She bends low and crawls into the one-person tunnel. The luminous breach remains visible on the rock door as its material disappears from the perimeter.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 윌마 디어링 (미국인 여성, 20대 후반, 자연스러운 성인 여성 얼굴, 구체적으로 명시되지 않은 자연색 머리카락) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/37154490-a4ed-4b37-9092-489c054d8a72/images/cc1d3d0b-626e-4f20-bc33-f5f9e2c6ee7e/scene/recipe/S6sh13__bgfirst_bg.png",
     "asset_id": "f813b3ae-a882-4493-8ed2-fc7fc34367b2",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/37154490-a4ed-4b37-9092-489c054d8a72/images/cc1d3d0b-626e-4f20-bc33-f5f9e2c6ee7e/conti/conti_S6sh13.png",
     "asset_id": "dd5f575c-605d-4e9e-99b6-db87eee785b5",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 윌마 디어링: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:766962>",
     "asset_id": "4982a15f-bf33-4ea0-b9c8-11e549bdf723",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/37154490-a4ed-4b37-9092-489c054d8a72/episodes/cc1d3d0b-626e-4f20-bc33-f5f9e2c6ee7e/images/background_chain/L11B01.png",
     "asset_id": "a755db6c-33f5-4899-9c38-164f6c8fa798",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 윌마 디어링: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:766962>",
     "asset_id": "4982a15f-bf33-4ea0-b9c8-11e549bdf723",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "카메라는 좁은 터널 안쪽을 똑바로 향하고 있으며, 윌마는 터널 깊은 곳을 향해 정면으로 기어 들어가고 있다.",
    "built_space": "입구 주변은 암석으로 되어 있고, 내부는 좁고 매끄러운 네모난 금속 터널이다. 화면 우측에는 레퍼런스 사진과 정확히 일치하는 붉은색 레버와 잠금장치가 달린 두꺼운 금속 해치문이 바깥쪽으로 열려 있다.",
    "entities": "뒷모습이 보이는 윌마 디어링. 녹색 점프수트와 전술화 등 캐릭터 레퍼런스의 복장과 체형이 정확히 일치한다.",
    "hard_violations": [],
    "physics": "무릎과 왼손으로 터널 바닥을 짚어 체중을 지탱하고 있으며, 오른발은 입구 턱에 자연스럽게 걸쳐져 있어 실제 기어가는 자세를 안정적으로 보여준다."
   },
   {
    "label": "B",
    "direction": "카메라는 바위 터널 안쪽을 향하고 있으며, 윌마는 화면 안쪽 멀리 있는 붉게 빛나는 원형 틈을 향해 기어가고 있다.",
    "built_space": "울퉁불퉁하고 불규칙한 좁은 바위 터널이다. 화면 좌측에 금속 프레임과 문이 일부 보이지만 레퍼런스의 우측 해치문과 구조 및 위치가 완전히 다르다. 터널 끝 배경에 붉게 빛나는 원형 구멍이 뚫려 있다.",
    "entities": "뒷모습의 윌마 디어링. 녹색 점프수트와 부츠 등 지정된 복장을 입고 있다.",
    "hard_violations": [],
    "physics": "양쪽 무릎과 손으로 바위 바닥을 짚고 엎드려 있으며, 바닥에 체중이 자연스럽게 실려 있다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "레퍼런스 이미지 우측에 위치한 작은 금속 해치문을 '1인용 터널'로 정확히 식별하고, 문의 디테일과 잠금장치까지 완벽하게 재현하여 로케이션 일치도에서 압도적으로 우수합니다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "터널 안쪽에 빛나는 틈을 묘사하려 했으나, 레퍼런스의 해치문 디자인과 위치(우측)를 완전히 무시하고 임의의 바위 터널과 좌측 문을 생성하여 로케이션 설정에 실패했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "카메라는 좁은 터널 안쪽을 똑바로 향하고 있으며, 윌마는 터널 깊은 곳을 향해 정면으로 기어 들어가고 있다.",
        "built_space": "입구 주변은 암석으로 되어 있고, 내부는 좁고 매끄러운 네모난 금속 터널이다. 화면 우측에는 레퍼런스 사진과 정확히 일치하는 붉은색 레버와 잠금장치가 달린 두꺼운 금속 해치문이 바깥쪽으로 열려 있다.",
        "entities": "뒷모습이 보이는 윌마 디어링. 녹색 점프수트와 전술화 등 캐릭터 레퍼런스의 복장과 체형이 정확히 일치한다.",
        "hard_violations": [],
        "physics": "무릎과 왼손으로 터널 바닥을 짚어 체중을 지탱하고 있으며, 오른발은 입구 턱에 자연스럽게 걸쳐져 있어 실제 기어가는 자세를 안정적으로 보여준다."
       },
       {
        "label": "B",
        "direction": "카메라는 바위 터널 안쪽을 향하고 있으며, 윌마는 화면 안쪽 멀리 있는 붉게 빛나는 원형 틈을 향해 기어가고 있다.",
        "built_space": "울퉁불퉁하고 불규칙한 좁은 바위 터널이다. 화면 좌측에 금속 프레임과 문이 일부 보이지만 레퍼런스의 우측 해치문과 구조 및 위치가 완전히 다르다. 터널 끝 배경에 붉게 빛나는 원형 구멍이 뚫려 있다.",
        "entities": "뒷모습의 윌마 디어링. 녹색 점프수트와 부츠 등 지정된 복장을 입고 있다.",
        "hard_violations": [],
        "physics": "양쪽 무릎과 손으로 바위 바닥을 짚고 엎드려 있으며, 바닥에 체중이 자연스럽게 실려 있다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "레퍼런스 이미지 우측에 위치한 작은 금속 해치문을 '1인용 터널'로 정확히 식별하고, 문의 디테일과 잠금장치까지 완벽하게 재현하여 로케이션 일치도에서 압도적으로 우수합니다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "터널 안쪽에 빛나는 틈을 묘사하려 했으나, 레퍼런스의 해치문 디자인과 위치(우측)를 완전히 무시하고 임의의 바위 터널과 좌측 문을 생성하여 로케이션 설정에 실패했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "카메라는 좁은 터널 안쪽을 똑바로 향하고 있으며, 윌마는 터널 깊은 곳을 향해 정면으로 기어 들어가고 있다.",
        "built_space": "입구 주변은 암석으로 되어 있고, 내부는 좁고 매끄러운 네모난 금속 터널이다. 화면 우측에는 레퍼런스 사진과 정확히 일치하는 붉은색 레버와 잠금장치가 달린 두꺼운 금속 해치문이 바깥쪽으로 열려 있다.",
        "entities": "뒷모습이 보이는 윌마 디어링. 녹색 점프수트와 전술화 등 캐릭터 레퍼런스의 복장과 체형이 정확히 일치한다.",
        "hard_violations": [],
        "physics": "무릎과 왼손으로 터널 바닥을 짚어 체중을 지탱하고 있으며, 오른발은 입구 턱에 자연스럽게 걸쳐져 있어 실제 기어가는 자세를 안정적으로 보여준다."
       },
       {
        "label": "B",
        "direction": "카메라는 바위 터널 안쪽을 향하고 있으며, 윌마는 화면 안쪽 멀리 있는 붉게 빛나는 원형 틈을 향해 기어가고 있다.",
        "built_space": "울퉁불퉁하고 불규칙한 좁은 바위 터널이다. 화면 좌측에 금속 프레임과 문이 일부 보이지만 레퍼런스의 우측 해치문과 구조 및 위치가 완전히 다르다. 터널 끝 배경에 붉게 빛나는 원형 구멍이 뚫려 있다.",
        "entities": "뒷모습의 윌마 디어링. 녹색 점프수트와 부츠 등 지정된 복장을 입고 있다.",
        "hard_violations": [],
        "physics": "양쪽 무릎과 손으로 바위 바닥을 짚고 엎드려 있으며, 바닥에 체중이 자연스럽게 실려 있다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "터널 내부의 중간 쇼트에서 윌마의 등과 낮은 포복 동작, 상단 중앙으로 이어지는 진행축, 발광 원형 균열을 모두 구현해 핵심 쇼트 지시를 가장 충실히 따른다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "포복 방향과 인물 배치는 맞지만 카메라가 터널 바깥 문턱에 가깝고, 반드시 지속되어야 할 발광 원형 균열이 보이지 않아 A보다 충실도가 낮다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "윌마는 카메라에 등을 보이고 몸과 머리를 화면 상단 중앙의 더 깊은 터널 및 붉은 원형 균열 쪽으로 향한다. 터널의 소실축도 윌마의 이동 방향과 정확히 일치한다. 무기나 별도의 지시 물체는 없다.",
        "built_space": "카메라는 좁은 암석 터널 안에 있으며, 좌우 암벽과 낮은 천장·바닥이 윌마를 밀착해 둘러싼다. 양쪽에 금속 보강 프레임이 보이고 통로는 한 사람이 기어갈 폭으로 중앙 후방까지 이어진다. 후방 끝에는 하나의 발광 원형 균열이 있는 문 같은 면이 보인다. 참고 장소의 외부 문틀과 우측 고정 설비는 이 내부 구도에서는 보이지 않지만, 터널 내부라는 장소 잠금과는 양립한다.",
        "entities": "등을 보이는 인물은 한 명뿐이며, 20대 후반 미국인 여성 윌마로 읽힌다. 묶은 자연색 금발과 체격, 짙은 녹색 작업복·전술 벨트·장갑·부츠가 인물 참고와 대체로 맞는다. 전화기는 보이지 않아 화면이 꺼졌는지는 확인할 수 없다. 붉게 빛나는 원형 균열은 후방 암석 문 표면에 물리적으로 놓인 발광 가장자리로 보이며 요구된 균열을 나타낸다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "윌마의 양 무릎과 왼손이 바닥을 지지하며, 다른 손은 몸과 원근에 가려진 채 앞쪽에 짚은 것으로 자연스럽게 이어진다. 부츠와 무릎의 굽힘, 앞으로 기운 골반과 등은 실제 네발 포복 중의 체중 이동으로 성립한다. 떠 있거나 지지되지 않은 신체나 물체는 없다."
       },
       {
        "label": "B",
        "direction": "윌마는 카메라에 등을 보이고 화면 상단 중앙의 직선형 금속 터널 안쪽을 향한다. 머리·몸통과 통로의 소실축이 같은 목표인 깊은 통로에 정렬되어 있어 이동 방향은 맞는다. 무기나 별도의 지시 물체는 없다.",
        "built_space": "정면에 하나의 큰 사각 금속 터널 입구와 연속된 내부 프레임들이 있고, 우측 암벽에는 참고 사진과 유사한 하나의 작은 열린 금속 해치, 하나의 격자 설비, 붉은 판, 배관이 보인다. 윌마는 큰 입구의 문턱을 넘어가는 위치이고 카메라는 터널 내부라기보다 바깥 통로 또는 입구 문턱 쪽에 놓였다. 터널은 한 사람만 기어갈 높이지만 A보다 폭이 넓고 매끈한 덕트 형태이며, 발광 균열이나 사라지는 암석 문 가장자리는 없다.",
        "entities": "인물은 윌마 한 명뿐이고 여성 체격과 짙은 녹색 작업복·부츠는 대체로 맞는다. 다만 참고 인물의 묶은 금발 대신 어깨 부근까지 내려온 머리로 보여 머리 재현이 덜 정확하다. 전화기는 보이지 않아 꺼진 화면을 확인할 수 없고, 요구된 발광 원형 균열도 없다. 우측의 작은 해치와 설비는 장소 참고와 닮았으며 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "왼손과 양 무릎 또는 정강이가 바닥에 하중을 전달하고, 한 부츠는 다음 포복 동작을 위해 뒤에서 들린 상태로 관절 연결과 추진 동작이 성립한다. 큰 문턱을 넘어 안쪽으로 기어가는 자세는 물리적으로 가능하며, 지지 없이 떠 있는 신체나 물체는 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "터널 내부의 중간 쇼트에서 윌마의 등과 낮은 포복 동작, 상단 중앙으로 이어지는 진행축, 발광 원형 균열을 모두 구현해 핵심 쇼트 지시를 가장 충실히 따른다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "포복 방향과 인물 배치는 맞지만 카메라가 터널 바깥 문턱에 가깝고, 반드시 지속되어야 할 발광 원형 균열이 보이지 않아 A보다 충실도가 낮다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "윌마는 카메라에 등을 보이고 몸과 머리를 화면 상단 중앙의 더 깊은 터널 및 붉은 원형 균열 쪽으로 향한다. 터널의 소실축도 윌마의 이동 방향과 정확히 일치한다. 무기나 별도의 지시 물체는 없다.",
        "built_space": "카메라는 좁은 암석 터널 안에 있으며, 좌우 암벽과 낮은 천장·바닥이 윌마를 밀착해 둘러싼다. 양쪽에 금속 보강 프레임이 보이고 통로는 한 사람이 기어갈 폭으로 중앙 후방까지 이어진다. 후방 끝에는 하나의 발광 원형 균열이 있는 문 같은 면이 보인다. 참고 장소의 외부 문틀과 우측 고정 설비는 이 내부 구도에서는 보이지 않지만, 터널 내부라는 장소 잠금과는 양립한다.",
        "entities": "등을 보이는 인물은 한 명뿐이며, 20대 후반 미국인 여성 윌마로 읽힌다. 묶은 자연색 금발과 체격, 짙은 녹색 작업복·전술 벨트·장갑·부츠가 인물 참고와 대체로 맞는다. 전화기는 보이지 않아 화면이 꺼졌는지는 확인할 수 없다. 붉게 빛나는 원형 균열은 후방 암석 문 표면에 물리적으로 놓인 발광 가장자리로 보이며 요구된 균열을 나타낸다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "윌마의 양 무릎과 왼손이 바닥을 지지하며, 다른 손은 몸과 원근에 가려진 채 앞쪽에 짚은 것으로 자연스럽게 이어진다. 부츠와 무릎의 굽힘, 앞으로 기운 골반과 등은 실제 네발 포복 중의 체중 이동으로 성립한다. 떠 있거나 지지되지 않은 신체나 물체는 없다."
       },
       {
        "label": "A",
        "direction": "윌마는 카메라에 등을 보이고 화면 상단 중앙의 직선형 금속 터널 안쪽을 향한다. 머리·몸통과 통로의 소실축이 같은 목표인 깊은 통로에 정렬되어 있어 이동 방향은 맞는다. 무기나 별도의 지시 물체는 없다.",
        "built_space": "정면에 하나의 큰 사각 금속 터널 입구와 연속된 내부 프레임들이 있고, 우측 암벽에는 참고 사진과 유사한 하나의 작은 열린 금속 해치, 하나의 격자 설비, 붉은 판, 배관이 보인다. 윌마는 큰 입구의 문턱을 넘어가는 위치이고 카메라는 터널 내부라기보다 바깥 통로 또는 입구 문턱 쪽에 놓였다. 터널은 한 사람만 기어갈 높이지만 A보다 폭이 넓고 매끈한 덕트 형태이며, 발광 균열이나 사라지는 암석 문 가장자리는 없다.",
        "entities": "인물은 윌마 한 명뿐이고 여성 체격과 짙은 녹색 작업복·부츠는 대체로 맞는다. 다만 참고 인물의 묶은 금발 대신 어깨 부근까지 내려온 머리로 보여 머리 재현이 덜 정확하다. 전화기는 보이지 않아 꺼진 화면을 확인할 수 없고, 요구된 발광 원형 균열도 없다. 우측의 작은 해치와 설비는 장소 참고와 닮았으며 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "왼손과 양 무릎 또는 정강이가 바닥에 하중을 전달하고, 한 부츠는 다음 포복 동작을 위해 뒤에서 들린 상태로 관절 연결과 추진 동작이 성립한다. 큰 문턱을 넘어 안쪽으로 기어가는 자세는 물리적으로 가능하며, 지지 없이 떠 있는 신체나 물체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt"
   ],
   "normalized": {
    "A": 1.75,
    "B": 1.556
   },
   "adjusted": {
    "A": 1.75,
    "B": 1.556
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1750,
   "B": 1556
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1750,
    "verdict_ko": "레퍼런스 이미지 우측에 위치한 작은 금속 해치문을 '1인용 터널'로 정확히 식별하고, 문의 디테일과 잠금장치까지 완벽하게 재현하여 로케이션 일치도에서 압도적으로 우수합니다."
   },
   {
    "label": "B",
    "score": 1556,
    "verdict_ko": "터널 안쪽에 빛나는 틈을 묘사하려 했으나, 레퍼런스의 해치문 디자인과 위치(우측)를 완전히 무시하고 임의의 바위 터널과 좌측 문을 생성하여 로케이션 설정에 실패했습니다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/37154490-a4ed-4b37-9092-489c054d8a72/episodes/cc1d3d0b-626e-4f20-bc33-f5f9e2c6ee7e/images/background_chain/L11B01.png",
    "asset_id": "a755db6c-33f5-4899-9c38-164f6c8fa798",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 윌마 디어링: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:766962>",
    "asset_id": "4982a15f-bf33-4ea0-b9c8-11e549bdf723",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": true,
  "shot_run_uid": "06a9bdd7-44c6-71f9-a75d-8005b3e9d80f",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/37154490-a4ed-4b37-9092-489c054d8a72/images/cc1d3d0b-626e-4f20-bc33-f5f9e2c6ee7e/scene/recipe/S6sh13__bgfirst_bg.png",
   "bg_asset_id": "f813b3ae-a882-4493-8ed2-fc7fc34367b2",
   "bg_record_key": "S6sh13::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S6sh13::cine": {
  "applied": true,
  "attempted_at": "2026-09-05T09:17:11.925943+00:00",
  "fingerprint": "69da3938e56bc1a0a4a7302d39ef747c32ba4221c6c3221718593291c127deba",
  "fingerprint_version": 2,
  "provider": "grok",
  "endpoint": "openrouter/chat-completions",
  "model": "x-ai/grok-imagine-image-2.0",
  "pack": "24.202608252115",
  "source_file": "S6sh13_sel.png",
  "source_sha256": "a3324f9cac3eed4df22dc6bc6a0dce183267181f8f600e986947fd2a3670c1f0",
  "file": "S6sh13_cine.png",
  "staged_sha256": "aa5ff4156e4875b13d5c7db7a14b0cc4c2cdc0a1b0201dd7f20b6b9d388e5fdf",
  "latency_ms": 13192
 },
 "S6sh17::signage": {
  "fp": "6e32cb901ab75bd4",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S6sh17::bgfirst_bg": {
  "input_fingerprint": "b47226cd369a184d",
  "prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 모니터의 백색 빛을 받으며 터널 안쪽을 향해 고개를 돌린 앨런의 얼굴\n\nLOCATION (lock): Inside the mountain contact post's underground control room, beside the concealed crawl-tunnel opening. White light from the erased monitors falls across the last man's face as he turns toward the tunnel.\n\nTIME OF DAY (lock): twilight.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: 앨런 in the middle-center of the frame, foreground, looks toward off-frame tunnel entrance.\n- KEY BACKGROUND ELEMENTS: Erased monitor screens (Showing blank white after record deletion) — The content-bearing faces are visible behind Alan's outer shoulder, displaying only white and remaining outside his direct eyeline; used as Backlights the consequence of Alan's choice without becoming his gaze target; Record-deletion lever (Pulled) — The lever's handle is held in its activated position beside the console; used as Retains physical evidence of the completed action near the lower frame edge; Tunnel entrance (Open) — The entrance lies beyond Alan's turn and outside the close frame; used as Off-frame destination motivating Alan's head turn.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: White monitor light shapes Alan's face within the cool, muted, low-key contact-station interior.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "effective_prompt": "Create the EMPTY BACKGROUND PLATE for one film shot — NO PEOPLE, no figures, no body parts, no sketch lines, no arrows anywhere.\n\"Empty\" means no people only: KEEP the location's inherent occupants and stock that define the place — animals in an animal shelter, pen or farm, goods in a market, moored boats in a harbour — unless the shot text explicitly removes them.\nThe FIRST attached image is a thin-line storyboard sketch: use ONLY its camera angle, horizon, perspective and the placement/size of buildings and set masses — ignore the sketched people and arrows entirely. The SECOND attached image (LOCATION PHOTOGRAPH) is the real place: take its architecture, materials, signage and fixed features, and RE-PROJECT them into the sketch's camera. If the photograph's camera differs from the sketch's, the sketch's camera wins.\nHUMAN-SCALE CALIBRATION: derive every structure's true size from human-scale elements — a door ≈ 2m, a window ≈ 1–1.5m wide, one storey ≈ 2.5–3m; never inflate a small structure or shrink a large one.\n\nSHOT TEXT this background must serve (Korean): 모니터의 백색 빛을 받으며 터널 안쪽을 향해 고개를 돌린 앨런의 얼굴\n\nLOCATION (lock): Inside the mountain contact post's underground control room, beside the concealed crawl-tunnel opening. White light from the erased monitors falls across the last man's face as he turns toward the tunnel.\n\nTIME OF DAY (lock): twilight.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: 앨런 in the middle-center of the frame, foreground, looks toward off-frame tunnel entrance.\n- KEY BACKGROUND ELEMENTS: Erased monitor screens (Showing blank white after record deletion) — The content-bearing faces are visible behind Alan's outer shoulder, displaying only white and remaining outside his direct eyeline; used as Backlights the consequence of Alan's choice without becoming his gaze target; Record-deletion lever (Pulled) — The lever's handle is held in its activated position beside the console; used as Retains physical evidence of the completed action near the lower frame edge; Tunnel entrance (Open) — The entrance lies beyond Alan's turn and outside the close frame; used as Off-frame destination motivating Alan's head turn.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: White monitor light shapes Alan's face within the cool, muted, low-key contact-station interior.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nRender ONE photorealistic empty location photograph, 16:9, that this shot can be staged inside later. No readable writing anywhere: surfaces that would carry writing may be present, but stage any wording out of legibility — an oblique angle, distance, shallow focus. No captions, watermarks or overlay text.",
  "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/37154490-a4ed-4b37-9092-489c054d8a72/images/cc1d3d0b-626e-4f20-bc33-f5f9e2c6ee7e/scene/recipe/S6sh17__bgfirst_bg.png",
  "asset_id": "5da98a53-d964-431c-9995-36ad516f43fc",
  "input_asset_ids": [
   "91f81319-717f-4970-bd79-b5801e9d909a",
   "a755db6c-33f5-4899-9c38-164f6c8fa798"
  ]
 },
 "S6sh17": {
  "input_fingerprint": "a042719b7f7c1683",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — Wyoming, United States, 2419; all people are English-speaking Americans unless stated. TIME OF DAY (lock): twilight.\n\nSHOT TEXT (authoritative, Korean): 모니터의 백색 빛을 받으며 터널 안쪽을 향해 고개를 돌린 앨런의 얼굴\n\nLOCATION (lock): Inside the mountain contact post's underground control room, beside the concealed crawl-tunnel opening. White light from the erased monitors falls across the last man's face as he turns toward the tunnel. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: 앨런 in the middle-center of the frame, foreground, looks toward off-frame tunnel entrance.\n- KEY BACKGROUND ELEMENTS: Erased monitor screens (Showing blank white after record deletion) — The content-bearing faces are visible behind Alan's outer shoulder, displaying only white and remaining outside his direct eyeline; used as Backlights the consequence of Alan's choice without becoming his gaze target; Record-deletion lever (Pulled) — The lever's handle is held in its activated position beside the console; used as Retains physical evidence of the completed action near the lower frame edge; Tunnel entrance (Open) — The entrance lies beyond Alan's turn and outside the close frame; used as Off-frame destination motivating Alan's head turn.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: White monitor light shapes Alan's face within the cool, muted, low-key contact-station interior.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The crawl tunnel remains open, the phone is powered off, and every contact-station screen has been wiped to solid white. The damaged rock door retains the bright circular breach. 앨런: He stays at the tunnel entrance with the lowered long gun after pulling the record-deletion lever, then turns toward the tunnel. The brilliant circular breach remains visible on the inner face of the rock door.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앨런 (미국계 성인 남성, 자연스러운 인간 얼굴, 자연 색상의 머리카락, 구체적 얼굴형은 명시되지 않음) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Stage the shot. The FIRST attached image (SHOT BACKGROUND) is the finished empty background of this shot — keep it EXACTLY: its camera, perspective, architecture, lighting and every fixed feature stay untouched. The SECOND attached image (LAYOUT SKETCH) tells you ONLY where the people go: each sketched person's position, screen size, pose and the gaze/motion arrows. Ignore the sketch's background lines. The CHARACTER REFERENCE photographs show the real people.\nPlace the real people into the background at exactly the sketched positions, sizes and poses, following the arrow directions. No sketch lines or arrows may remain.\n\nCreate ONE FINAL photorealistic live-action film still of the moment below — Wyoming, United States, 2419; all people are English-speaking Americans unless stated. TIME OF DAY (lock): twilight.\n\nSHOT TEXT (authoritative, Korean): 모니터의 백색 빛을 받으며 터널 안쪽을 향해 고개를 돌린 앨런의 얼굴\n\nLOCATION (lock): Inside the mountain contact post's underground control room, beside the concealed crawl-tunnel opening. White light from the erased monitors falls across the last man's face as he turns toward the tunnel. The shot takes place here — the FIRST attached image (SHOT BACKGROUND) is this exact place, already built: its ground, structures, horizon, materials and lighting are the finished truth of this location and must not be redesigned or replaced. No location photograph is attached — read the place from that image alone, and add no scenery, structure, vehicle or fixture that it does not already show. This lock governs the place only; the figures in the shot follow the staging and pose instructions.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: 앨런 in the middle-center of the frame, foreground, looks toward off-frame tunnel entrance.\n- KEY BACKGROUND ELEMENTS: Erased monitor screens (Showing blank white after record deletion) — The content-bearing faces are visible behind Alan's outer shoulder, displaying only white and remaining outside his direct eyeline; used as Backlights the consequence of Alan's choice without becoming his gaze target; Record-deletion lever (Pulled) — The lever's handle is held in its activated position beside the console; used as Retains physical evidence of the completed action near the lower frame edge; Tunnel entrance (Open) — The entrance lies beyond Alan's turn and outside the close frame; used as Off-frame destination motivating Alan's head turn.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: White monitor light shapes Alan's face within the cool, muted, low-key contact-station interior.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The crawl tunnel remains open, the phone is powered off, and every contact-station screen has been wiped to solid white. The damaged rock door retains the bright circular breach. 앨런: He stays at the tunnel entrance with the lowered long gun after pulling the record-deletion lever, then turns toward the tunnel. The brilliant circular breach remains visible on the inner face of the rock door.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앨런 (미국계 성인 남성, 자연스러운 인간 얼굴, 자연 색상의 머리카락, 구체적 얼굴형은 명시되지 않음) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — Wyoming, United States, 2419; all people are English-speaking Americans unless stated. TIME OF DAY (lock): twilight.\n\nSHOT TEXT (authoritative, Korean): 모니터의 백색 빛을 받으며 터널 안쪽을 향해 고개를 돌린 앨런의 얼굴\n\nLOCATION (lock): Inside the mountain contact post's underground control room, beside the concealed crawl-tunnel opening. White light from the erased monitors falls across the last man's face as he turns toward the tunnel. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: close-up\n- FRAME LAYOUT: 앨런 in the middle-center of the frame, foreground, looks toward off-frame tunnel entrance.\n- KEY BACKGROUND ELEMENTS: Erased monitor screens (Showing blank white after record deletion) — The content-bearing faces are visible behind Alan's outer shoulder, displaying only white and remaining outside his direct eyeline; used as Backlights the consequence of Alan's choice without becoming his gaze target; Record-deletion lever (Pulled) — The lever's handle is held in its activated position beside the console; used as Retains physical evidence of the completed action near the lower frame edge; Tunnel entrance (Open) — The entrance lies beyond Alan's turn and outside the close frame; used as Off-frame destination motivating Alan's head turn.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: White monitor light shapes Alan's face within the cool, muted, low-key contact-station interior.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nEVERY CHARACTER IS A HUMAN BEING: unless the story explicitly\nfeatures non-human or virtual beings (as in science fiction or\nfantasy), every character — however indirectly, vaguely or\nfiguratively the text describes them — IS a real human. When the\ntext gives no direct visual description of a person it puts in\nthis shot, imagine that description and still show them as a\nconcrete, fully-formed human being:\nrender the human form as fully as the framing shows — never reduce\na person to a shape, blob, solid silhouette or abstract mass.\n\nEXPRESSIONS ARE ACTED, NEVER ANATOMICAL: when the text describes a\nperson's eyes, face or presence emotionally or figuratively (vacant,\nhollow, dazed, burning, lifeless gaze and the like), it describes an\nACTOR'S PERFORMANCE captured by a real camera — realize it ONLY\nthrough gaze direction, focus, eyelids, facial muscles, stillness\nand posture. Every living person's eyes remain anatomically normal\nhuman eyes with a natural iris and pupil, natural sclera and normal\nproportions; NEVER whiten, blank out, cloud over, glow, enlarge or\notherwise alter eyeballs, skin or anatomy — unless the story\nexplicitly declares that being non-human or supernatural in form.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nNATURAL PERFORMANCE (default only — every explicit direction above\nwins): when no pose contract, immobility contract or shot-text\ndirection says otherwise, people read as alive in mid-moment —\nbelievable weight shift, hands naturally positioned for the action\nthe SHOT TEXT already gives them, gaze on the target the SHOT TEXT\nimplies; avoid a stiff attention stance (feet together, arms hanging\nstraight down) and a blank stare into the lens. If the SHOT TEXT or\nany POSE/IMMOBILE contract above stages stillness, death, sleep,\nunconsciousness, restraint, drill or an explicit direct-to-camera\nlook, follow THAT exactly — this clause never overrides it.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nBODY & SUPPORT (default only — every explicit direction above wins): a body relates to what holds it. A seated person sits the way the seat is built to be used — hips on the seat, back toward the backrest, legs falling naturally toward the floor; people sharing adjacent seats each occupy their own seat, side by side. Hands, hips and feet keep believable contact with whatever they rest on. When the text stages someone frozen, stunned or holding still, the body stops mid-action exactly where the moment caught it — weight already committed to one side, hands where the interrupted movement left them — rather than resetting into a symmetric at-attention stance. If the SHOT TEXT or any contract above stages a specific arrangement, follow THAT exactly.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The crawl tunnel remains open, the phone is powered off, and every contact-station screen has been wiped to solid white. The damaged rock door retains the bright circular breach. 앨런: He stays at the tunnel entrance with the lowered long gun after pulling the record-deletion lever, then turns toward the tunnel. The brilliant circular breach remains visible on the inner face of the rock door.\n\nPEOPLE: the SHOT TEXT alone decides whether any person is visible in this shot. IF a person appears, they must be one of: 앨런 (미국계 성인 남성, 자연스러운 인간 얼굴, 자연 색상의 머리카락, 구체적 얼굴형은 명시되지 않음) — never anyone else, and never add a person the shot text does not show. IF only part of a person is in frame (a hand, arm, foot, back, silhouette), that body part belongs to the specific person the shot text names — its sex, age, build, skin and grooming must unmistakably match that person's profile above.\n\nCHARACTER REFERENCE ROLE (follow exactly): the attached CHARACTER REFERENCE images establish identity only — face, hair, build and clothing. The pose, gaze direction, camera angle and framing inside those reference images belong to the reference photos, not to this shot; never copy them. Pose and gaze in this shot follow only the CAMERA and action text of this prompt. When this prompt stages a character's face as covered or hidden — by a costume head, mask, helmet, hood, or a body turned away — that staging wins: keep the covering exactly as described and never pull the reference face into view from under it; the reference then guides only what stays visible, such as build and clothing.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "roll_refs": {
   "A": [
    {
     "label": "SHOT BACKGROUND",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/37154490-a4ed-4b37-9092-489c054d8a72/images/cc1d3d0b-626e-4f20-bc33-f5f9e2c6ee7e/scene/recipe/S6sh17__bgfirst_bg.png",
     "asset_id": "5da98a53-d964-431c-9995-36ad516f43fc",
     "role": "bgfirst_bg"
    },
    {
     "label": "LAYOUT SKETCH (people placement only)",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/37154490-a4ed-4b37-9092-489c054d8a72/images/cc1d3d0b-626e-4f20-bc33-f5f9e2c6ee7e/conti/conti_S6sh17.png",
     "asset_id": "91f81319-717f-4970-bd79-b5801e9d909a",
     "role": "conti_light"
    },
    {
     "label": "CHARACTER REFERENCE — 앨런: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1001252>",
     "asset_id": "b426ff26-eb7d-4b66-a4a7-de7c648b9df3",
     "role": "character_ref"
    }
   ],
   "B": [
    {
     "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
     "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/37154490-a4ed-4b37-9092-489c054d8a72/episodes/cc1d3d0b-626e-4f20-bc33-f5f9e2c6ee7e/images/background_chain/L11B01.png",
     "asset_id": "a755db6c-33f5-4899-9c38-164f6c8fa798",
     "role": "location_plate"
    },
    {
     "label": "CHARACTER REFERENCE — 앨런: the exact person appearing in this shot; match face, hair and build exactly.",
     "path": "<bytes:1001252>",
     "asset_id": "b426ff26-eb7d-4b66-a4a7-de7c648b9df3",
     "role": "character_ref"
    }
   ]
  },
  "shot_run_spend_attempt_count": 1,
  "judge_flip_skipped": "judge_fn_owns_order",
  "readings": [
   {
    "label": "A",
    "direction": "앨런은 화면 오른쪽 프레임 밖(암시된 터널 방향)을 향해 고개를 돌려 시선을 두고 있으며, 좌측 하단에 일부 보이는 장총의 총구는 아래를 향하고 있습니다.",
    "built_space": "카메라는 터널 입구 부근에서 통제실 안쪽을 바라보는 구도입니다. 배경에 창문 3개, 모니터 2대, 중앙의 둥근 탁자 등 레퍼런스의 공간적 특징이 정확히 배치되었습니다. 앨런은 입구 바로 옆에 서 있고, 우측에는 암석 문의 일부가, 좌측 하단 프레임 경계에는 지시된 콘솔이 자연스럽게 자리 잡고 있습니다.",
    "entities": "앨런의 인상착의(비니, 수염, 재킷)는 캐릭터 레퍼런스와 정확히 일치합니다. 프롬프트가 요구한 백색 화면의 모니터들, 하단의 레버, 아래로 향한 장총, 그리고 우측 암석 표면의 빛나는 둥근 파공이 모두 화면에 빠짐없이 존재합니다.",
    "hard_violations": [],
    "physics": "앨런은 입구 쪽에 지면을 딛고 안정적으로 서 있으며, 하단의 장총은 몸 쪽에 지지된 상태로 자연스럽게 늘어져 있어 떠 있거나 물리 법칙에 어긋나는 요소가 없습니다."
   },
   {
    "label": "B",
    "direction": "앨런은 오른쪽 뒤편 프레임 내에 보이는 터널 입구를 향해 고개를 돌려 시선을 두고 있으며, 오른손은 데스크 위의 레버를 향해 뻗어 쥐고 있습니다.",
    "built_space": "카메라는 통제실 내부에서 우측 벽면과 출입구를 바라보는 시점을 취하고 있습니다. 하지만 앨런이 '터널 입구'가 아닌 방 안쪽의 대형 모니터가 있는 데스크 앞에 서 있어 프롬프트가 지시한 공간 내 동선 및 배치와 완전히 어긋납니다.",
    "entities": "앨런의 외양과 의상은 레퍼런스와 일치하며 백색 모니터와 레버가 등장합니다. 하지만 프롬프트에 명시된 필수 요소인 '내려놓은 장총'과 '암석 문의 빛나는 파공'이 전혀 묘사되지 않았습니다.",
    "hard_violations": [
     "[gemini-pro] 앨런이 터널 입구가 아닌 방 안쪽 데스크에 배치되어 인물의 공간적 위치 지시를 위반함 (Person placed where the staging does not put them)",
     "[gemini-pro] 클로즈업이 아닌 미디엄 샷으로 프레이밍되었고 프레임 밖에 있어야 할 터널 입구가 프레임 내부에 노출됨 (Staging and framing negative constraint violation)"
    ],
    "physics": "앨런은 데스크 앞에서 바닥을 딛고 자연스럽게 서서 레버를 조작하고 있으며, 공중에 떠 있거나 지지되지 않은 물리적 오류는 보이지 않습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "클로즈업 프레이밍, 프레임 밖 터널을 향한 시선, 백색 모니터 조명, 장총 및 암석 문의 빛나는 파공 등 샷 텍스트와 모든 세부 지시를 훌륭하게 구현한 훌륭한 결과물입니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "클로즈업이 아닌 미디엄 샷으로 렌더링되었고 터널 입구가 프레임 내에 직접 노출되었으며, 인물의 위치가 잘못 배치되고 장총과 파공이 누락되어 핵심 지시사항들을 다수 위반했습니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "앨런은 화면 오른쪽 프레임 밖(암시된 터널 방향)을 향해 고개를 돌려 시선을 두고 있으며, 좌측 하단에 일부 보이는 장총의 총구는 아래를 향하고 있습니다.",
        "built_space": "카메라는 터널 입구 부근에서 통제실 안쪽을 바라보는 구도입니다. 배경에 창문 3개, 모니터 2대, 중앙의 둥근 탁자 등 레퍼런스의 공간적 특징이 정확히 배치되었습니다. 앨런은 입구 바로 옆에 서 있고, 우측에는 암석 문의 일부가, 좌측 하단 프레임 경계에는 지시된 콘솔이 자연스럽게 자리 잡고 있습니다.",
        "entities": "앨런의 인상착의(비니, 수염, 재킷)는 캐릭터 레퍼런스와 정확히 일치합니다. 프롬프트가 요구한 백색 화면의 모니터들, 하단의 레버, 아래로 향한 장총, 그리고 우측 암석 표면의 빛나는 둥근 파공이 모두 화면에 빠짐없이 존재합니다.",
        "hard_violations": [],
        "physics": "앨런은 입구 쪽에 지면을 딛고 안정적으로 서 있으며, 하단의 장총은 몸 쪽에 지지된 상태로 자연스럽게 늘어져 있어 떠 있거나 물리 법칙에 어긋나는 요소가 없습니다."
       },
       {
        "label": "B",
        "direction": "앨런은 오른쪽 뒤편 프레임 내에 보이는 터널 입구를 향해 고개를 돌려 시선을 두고 있으며, 오른손은 데스크 위의 레버를 향해 뻗어 쥐고 있습니다.",
        "built_space": "카메라는 통제실 내부에서 우측 벽면과 출입구를 바라보는 시점을 취하고 있습니다. 하지만 앨런이 '터널 입구'가 아닌 방 안쪽의 대형 모니터가 있는 데스크 앞에 서 있어 프롬프트가 지시한 공간 내 동선 및 배치와 완전히 어긋납니다.",
        "entities": "앨런의 외양과 의상은 레퍼런스와 일치하며 백색 모니터와 레버가 등장합니다. 하지만 프롬프트에 명시된 필수 요소인 '내려놓은 장총'과 '암석 문의 빛나는 파공'이 전혀 묘사되지 않았습니다.",
        "hard_violations": [
         "앨런이 터널 입구가 아닌 방 안쪽 데스크에 배치되어 인물의 공간적 위치 지시를 위반함 (Person placed where the staging does not put them)",
         "클로즈업이 아닌 미디엄 샷으로 프레이밍되었고 프레임 밖에 있어야 할 터널 입구가 프레임 내부에 노출됨 (Staging and framing negative constraint violation)"
        ],
        "physics": "앨런은 데스크 앞에서 바닥을 딛고 자연스럽게 서서 레버를 조작하고 있으며, 공중에 떠 있거나 지지되지 않은 물리적 오류는 보이지 않습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "클로즈업 프레이밍, 프레임 밖 터널을 향한 시선, 백색 모니터 조명, 장총 및 암석 문의 빛나는 파공 등 샷 텍스트와 모든 세부 지시를 훌륭하게 구현한 훌륭한 결과물입니다."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "클로즈업이 아닌 미디엄 샷으로 렌더링되었고 터널 입구가 프레임 내에 직접 노출되었으며, 인물의 위치가 잘못 배치되고 장총과 파공이 누락되어 핵심 지시사항들을 다수 위반했습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "앨런은 화면 오른쪽 프레임 밖(암시된 터널 방향)을 향해 고개를 돌려 시선을 두고 있으며, 좌측 하단에 일부 보이는 장총의 총구는 아래를 향하고 있습니다.",
        "built_space": "카메라는 터널 입구 부근에서 통제실 안쪽을 바라보는 구도입니다. 배경에 창문 3개, 모니터 2대, 중앙의 둥근 탁자 등 레퍼런스의 공간적 특징이 정확히 배치되었습니다. 앨런은 입구 바로 옆에 서 있고, 우측에는 암석 문의 일부가, 좌측 하단 프레임 경계에는 지시된 콘솔이 자연스럽게 자리 잡고 있습니다.",
        "entities": "앨런의 인상착의(비니, 수염, 재킷)는 캐릭터 레퍼런스와 정확히 일치합니다. 프롬프트가 요구한 백색 화면의 모니터들, 하단의 레버, 아래로 향한 장총, 그리고 우측 암석 표면의 빛나는 둥근 파공이 모두 화면에 빠짐없이 존재합니다.",
        "hard_violations": [],
        "physics": "앨런은 입구 쪽에 지면을 딛고 안정적으로 서 있으며, 하단의 장총은 몸 쪽에 지지된 상태로 자연스럽게 늘어져 있어 떠 있거나 물리 법칙에 어긋나는 요소가 없습니다."
       },
       {
        "label": "B",
        "direction": "앨런은 오른쪽 뒤편 프레임 내에 보이는 터널 입구를 향해 고개를 돌려 시선을 두고 있으며, 오른손은 데스크 위의 레버를 향해 뻗어 쥐고 있습니다.",
        "built_space": "카메라는 통제실 내부에서 우측 벽면과 출입구를 바라보는 시점을 취하고 있습니다. 하지만 앨런이 '터널 입구'가 아닌 방 안쪽의 대형 모니터가 있는 데스크 앞에 서 있어 프롬프트가 지시한 공간 내 동선 및 배치와 완전히 어긋납니다.",
        "entities": "앨런의 외양과 의상은 레퍼런스와 일치하며 백색 모니터와 레버가 등장합니다. 하지만 프롬프트에 명시된 필수 요소인 '내려놓은 장총'과 '암석 문의 빛나는 파공'이 전혀 묘사되지 않았습니다.",
        "hard_violations": [
         "앨런이 터널 입구가 아닌 방 안쪽 데스크에 배치되어 인물의 공간적 위치 지시를 위반함 (Person placed where the staging does not put them)",
         "클로즈업이 아닌 미디엄 샷으로 프레이밍되었고 프레임 밖에 있어야 할 터널 입구가 프레임 내부에 노출됨 (Staging and framing negative constraint violation)"
        ],
        "physics": "앨런은 데스크 앞에서 바닥을 딛고 자연스럽게 서서 레버를 조작하고 있으며, 공중에 떠 있거나 지지되지 않은 물리적 오류는 보이지 않습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "요구된 얼굴 중심 클로즈업, 백색 모니터 역광, 레버를 당긴 직후 터널 쪽으로 고개와 시선을 돌리는 순간을 가장 충실히 구현했지만 터널 입구가 프레임 밖이 아니라 오른쪽에 보인다."
       },
       {
        "label": "B",
        "score": 5,
        "verdict_ko": "장소와 백색 화면, 내려진 장총 및 레버 상태는 잘 보존했으나 클로즈업이 아닌 넓은 미디엄 숏이며 앨런의 시선이 오른쪽 터널·해치가 아니라 왼쪽 통제실 안쪽을 향한다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "앨런은 몸을 콘솔 쪽에 둔 채 머리와 눈을 화면 오른쪽의 어두운 암반 터널 입구로 돌리고 있어 실제로 보이는 터널을 향한다. 무기 총구나 다른 지시 물체는 보이지 않는다.",
        "built_space": "암반과 금속 프레임으로 둘러싸인 통제실 내부이며, 앨런 뒤 왼쪽에 백색 화면 세 개가 보이고 오른쪽에는 열린 터널 통로 하나가 있다. 콘솔 하단 가장자리에는 레버 하나가 있으며 앨런의 왼손이 손잡이를 잡은 채 당겨진 위치를 유지한다. 다만 기준 장소의 작은 측면 해치 구조보다는 큰 보행용 터널로 바뀌었고, 본래 프레임 밖이어야 할 터널 입구가 가까운 배경에 노출된다. 반사상 모순은 없다.",
        "entities": "앨런 한 명만 등장하며 성인 미국계 남성, 회색 비니, 짧은 수염, 회색 작업 재킷과 얼굴 인상이 인물 기준과 잘 맞는다. 지워진 모니터들은 글자 없는 백색 화면이며 레버도 확인된다. 장총과 원형 breach는 이 클로즈업에서 보이지 않지만 올바른 프레이밍이 배제한 항목이므로 감점 근거가 아니다. 보이는 터널은 열려 있으나 기준 장소의 은폐형 크롤 터널과 형태가 다소 다르다.",
        "hard_violations": [],
        "physics": "앨런은 콘솔에 몸을 기대고 오른팔을 아래쪽에 둔 안정된 자세이며, 왼손이 레버 손잡이를 실제로 잡고 있다. 머리 회전과 시선 전환도 몸통의 비틀림으로 자연스럽게 지지된다. 떠 있거나 지지되지 않은 인물·물체는 없다."
       },
       {
        "label": "B",
        "direction": "앨런의 얼굴과 눈은 화면 왼쪽, 즉 통제실 내부를 향한다. 열린 크롤 터널로 해석되는 작은 해치와 밝은 원형 breach는 그의 오른쪽 뒤에 있으므로 시선이 목표에 닿지 않는다. 내려진 장총은 아래쪽을 향해 공격 목표를 겨누지 않는다.",
        "built_space": "기준 사진과 가까운 콘크리트 통제실, 후면 콘솔, 의자 두 개, 중앙 원형 테이블 하나, 여러 백색 화면, 오른쪽 암반 벽의 작은 금속 해치 하나가 보인다. 전경 왼쪽에는 당겨진 레버 하나가 있고 앨런은 오른쪽 해치 옆 출입 프레임에 서 있다. 다만 밝은 원형 breach가 금속문 자체가 아니라 해치 위 암벽에 난 구멍처럼 배치되었다. 광학적으로 불가능한 반사는 없다.",
        "entities": "앨런 한 명만 있으며 성인 남성, 비니, 수염, 회색 작업복은 기준 인물과 대체로 일치한다. 모든 모니터는 백색으로 지워졌고 내려진 장총과 작동된 레버도 보인다. 작은 금속 해치와 밝은 원형 구멍은 존재하지만 원형 breach의 위치가 ‘암석 문의 안쪽 면’과 맞지 않는다. 읽을 수 있는 문구나 추가 인물은 없다.",
        "hard_violations": [],
        "physics": "앨런은 바닥에 서 있는 수직 자세로 보이며 몸이 출입 프레임 가까이에 안정적으로 놓여 있다. 장총은 몸 앞에서 아래로 늘어져 있고 신체나 스트랩에 의해 지지되는 형태이며, 레버는 기계 베이스에 고정되어 있다. 떠 있거나 아무 지지도 받지 않는 대상은 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "요구된 얼굴 중심 클로즈업, 백색 모니터 역광, 레버를 당긴 직후 터널 쪽으로 고개와 시선을 돌리는 순간을 가장 충실히 구현했지만 터널 입구가 프레임 밖이 아니라 오른쪽에 보인다."
       },
       {
        "label": "A",
        "score": 5,
        "verdict_ko": "장소와 백색 화면, 내려진 장총 및 레버 상태는 잘 보존했으나 클로즈업이 아닌 넓은 미디엄 숏이며 앨런의 시선이 오른쪽 터널·해치가 아니라 왼쪽 통제실 안쪽을 향한다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "앨런은 몸을 콘솔 쪽에 둔 채 머리와 눈을 화면 오른쪽의 어두운 암반 터널 입구로 돌리고 있어 실제로 보이는 터널을 향한다. 무기 총구나 다른 지시 물체는 보이지 않는다.",
        "built_space": "암반과 금속 프레임으로 둘러싸인 통제실 내부이며, 앨런 뒤 왼쪽에 백색 화면 세 개가 보이고 오른쪽에는 열린 터널 통로 하나가 있다. 콘솔 하단 가장자리에는 레버 하나가 있으며 앨런의 왼손이 손잡이를 잡은 채 당겨진 위치를 유지한다. 다만 기준 장소의 작은 측면 해치 구조보다는 큰 보행용 터널로 바뀌었고, 본래 프레임 밖이어야 할 터널 입구가 가까운 배경에 노출된다. 반사상 모순은 없다.",
        "entities": "앨런 한 명만 등장하며 성인 미국계 남성, 회색 비니, 짧은 수염, 회색 작업 재킷과 얼굴 인상이 인물 기준과 잘 맞는다. 지워진 모니터들은 글자 없는 백색 화면이며 레버도 확인된다. 장총과 원형 breach는 이 클로즈업에서 보이지 않지만 올바른 프레이밍이 배제한 항목이므로 감점 근거가 아니다. 보이는 터널은 열려 있으나 기준 장소의 은폐형 크롤 터널과 형태가 다소 다르다.",
        "hard_violations": [],
        "physics": "앨런은 콘솔에 몸을 기대고 오른팔을 아래쪽에 둔 안정된 자세이며, 왼손이 레버 손잡이를 실제로 잡고 있다. 머리 회전과 시선 전환도 몸통의 비틀림으로 자연스럽게 지지된다. 떠 있거나 지지되지 않은 인물·물체는 없다."
       },
       {
        "label": "A",
        "direction": "앨런의 얼굴과 눈은 화면 왼쪽, 즉 통제실 내부를 향한다. 열린 크롤 터널로 해석되는 작은 해치와 밝은 원형 breach는 그의 오른쪽 뒤에 있으므로 시선이 목표에 닿지 않는다. 내려진 장총은 아래쪽을 향해 공격 목표를 겨누지 않는다.",
        "built_space": "기준 사진과 가까운 콘크리트 통제실, 후면 콘솔, 의자 두 개, 중앙 원형 테이블 하나, 여러 백색 화면, 오른쪽 암반 벽의 작은 금속 해치 하나가 보인다. 전경 왼쪽에는 당겨진 레버 하나가 있고 앨런은 오른쪽 해치 옆 출입 프레임에 서 있다. 다만 밝은 원형 breach가 금속문 자체가 아니라 해치 위 암벽에 난 구멍처럼 배치되었다. 광학적으로 불가능한 반사는 없다.",
        "entities": "앨런 한 명만 있으며 성인 남성, 비니, 수염, 회색 작업복은 기준 인물과 대체로 일치한다. 모든 모니터는 백색으로 지워졌고 내려진 장총과 작동된 레버도 보인다. 작은 금속 해치와 밝은 원형 구멍은 존재하지만 원형 breach의 위치가 ‘암석 문의 안쪽 면’과 맞지 않는다. 읽을 수 있는 문구나 추가 인물은 없다.",
        "hard_violations": [],
        "physics": "앨런은 바닥에 서 있는 수직 자세로 보이며 몸이 출입 프레임 가까이에 안정적으로 놓여 있다. 장총은 몸 앞에서 아래로 늘어져 있고 신체나 스트랩에 의해 지지되는 형태이며, 레버는 기계 베이스에 고정되어 있다. 떠 있거나 아무 지지도 받지 않는 대상은 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt"
   ],
   "normalized": {
    "A": 1.625,
    "B": 1.333
   },
   "adjusted": {
    "A": 1.625,
    "B": 1.083
   },
   "violations": {
    "B": [
     "[gemini-pro] 앨런이 터널 입구가 아닌 방 안쪽 데스크에 배치되어 인물의 공간적 위치 지시를 위반함 (Person placed where the staging does not put them)",
     "[gemini-pro] 클로즈업이 아닌 미디엄 샷으로 프레이밍되었고 프레임 밖에 있어야 할 터널 입구가 프레임 내부에 노출됨 (Staging and framing negative constraint violation)"
    ]
   },
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1625,
   "B": 1083
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1625,
    "verdict_ko": "클로즈업 프레이밍, 프레임 밖 터널을 향한 시선, 백색 모니터 조명, 장총 및 암석 문의 빛나는 파공 등 샷 텍스트와 모든 세부 지시를 훌륭하게 구현한 훌륭한 결과물입니다."
   },
   {
    "label": "B",
    "score": 1083,
    "verdict_ko": "클로즈업이 아닌 미디엄 샷으로 렌더링되었고 터널 입구가 프레임 내에 직접 노출되었으며, 인물의 위치가 잘못 배치되고 장총과 파공이 누락되어 핵심 지시사항들을 다수 위반했습니다.  ★위반: [gemini-pro] 앨런이 터널 입구가 아닌 방 안쪽 데스크에 배치되어 인물의 공간적 위치 지시를 위반함 (Person placed where the staging does not put them) / [gemini-pro] 클로즈업이 아닌 미디엄 샷으로 프레이밍되었고 프레임 밖에 있어야 할 터널 입구가 프레임 내부에 노출됨 (Staging and framing negative constraint violation)"
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/37154490-a4ed-4b37-9092-489c054d8a72/episodes/cc1d3d0b-626e-4f20-bc33-f5f9e2c6ee7e/images/background_chain/L11B01.png",
    "asset_id": "a755db6c-33f5-4899-9c38-164f6c8fa798",
    "role": "location_plate"
   },
   {
    "label": "CHARACTER REFERENCE — 앨런: the exact person appearing in this shot; match face, hair and build exactly.",
    "path": "<bytes:1001252>",
    "asset_id": "b426ff26-eb7d-4b66-a4a7-de7c648b9df3",
    "role": "character_ref"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": true,
  "shot_run_uid": "06a9bde2-542f-7ec5-99a9-f6cad6b4716b",
  "bgfirst": {
   "bg_path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/37154490-a4ed-4b37-9092-489c054d8a72/images/cc1d3d0b-626e-4f20-bc33-f5f9e2c6ee7e/scene/recipe/S6sh17__bgfirst_bg.png",
   "bg_asset_id": "5da98a53-d964-431c-9995-36ad516f43fc",
   "bg_record_key": "S6sh17::bgfirst_bg",
   "chain_winner": true,
   "authority": "plate"
  },
  "ref_mode": "재투영 배경+콘티+엔티티 (2택1: 체인 승)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S6sh17::cine": {
  "applied": true,
  "attempted_at": "2026-09-05T09:20:55.517270+00:00",
  "fingerprint": "88a15cf7c49db741998baafbe59799f8cc658907d92c3263ae996b2529d81806",
  "fingerprint_version": 2,
  "provider": "grok",
  "endpoint": "openrouter/chat-completions",
  "model": "x-ai/grok-imagine-image-2.0",
  "pack": "24.202608252115",
  "source_file": "S6sh17_sel.png",
  "source_sha256": "2422743eb2f7734526ffb334f2581e5b2c2e058100a01673807cbec7f316dd48",
  "file": "S6sh17_cine.png",
  "staged_sha256": "b17aebb967918c41835174c1cf5b21f45590299198e8b4a7c956aa84d5b62337",
  "latency_ms": 18343
 },
 "S7sh3::signage": {
  "fp": "5dc1cf557be8ea3f",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S7sh3": {
  "input_fingerprint": "5ed5f6fa3e7c7024",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — Wyoming, United States, 2419; all people are English-speaking Americans unless stated. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 어두운 산 정상의 능선을 따라 수많은 붉은 불빛들이 밀집한 원경\n\nLOCATION (lock): Outside on the nighttime forested mountainside near the concealed rock exit, looking toward the dark summit ridge. Numerous red lights cluster along the distant crest. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: ridge with converging red lights in the upper-center of the frame, background.\n- KEY BACKGROUND ELEMENTS: Mountain ridge (Dark and empty of the escaping group) — Its crest is seen from the distant elevated vantage, extending laterally through the upper frame; used as Carries the concentrated lights and confirms that the decoy drew attention uphill; Numerous red lights (Densely converged along the ridge); used as Primary distant focal pattern and restrained warning-red accent; Empty hillside distance (Separating the camera from the ridge) — The slope recedes upward toward the lit crest; used as Establishes scale and confirms the lights are far from the group.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Night darkness and the concentrated warning-red lights create a cool, muted long view with stark but restrained points of contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): It is night around the wet rock fissure below the secret exit, and the contact station is collapsing quietly behind it. Numerous red lights cluster along the empty mountaintop ridge.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — Wyoming, United States, 2419; all people are English-speaking Americans unless stated. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 어두운 산 정상의 능선을 따라 수많은 붉은 불빛들이 밀집한 원경\n\nLOCATION (lock): Outside on the nighttime forested mountainside near the concealed rock exit, looking toward the dark summit ridge. Numerous red lights cluster along the distant crest. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: ridge with converging red lights in the upper-center of the frame, background.\n- KEY BACKGROUND ELEMENTS: Mountain ridge (Dark and empty of the escaping group) — Its crest is seen from the distant elevated vantage, extending laterally through the upper frame; used as Carries the concentrated lights and confirms that the decoy drew attention uphill; Numerous red lights (Densely converged along the ridge); used as Primary distant focal pattern and restrained warning-red accent; Empty hillside distance (Separating the camera from the ridge) — The slope recedes upward toward the lit crest; used as Establishes scale and confirms the lights are far from the group.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Night darkness and the concentrated warning-red lights create a cool, muted long view with stark but restrained points of contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): It is night around the wet rock fissure below the secret exit, and the contact station is collapsing quietly behind it. Numerous red lights cluster along the empty mountaintop ridge.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — Wyoming, United States, 2419; all people are English-speaking Americans unless stated. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 어두운 산 정상의 능선을 따라 수많은 붉은 불빛들이 밀집한 원경\n\nLOCATION (lock): Outside on the nighttime forested mountainside near the concealed rock exit, looking toward the dark summit ridge. Numerous red lights cluster along the distant crest. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: ridge with converging red lights in the upper-center of the frame, background.\n- KEY BACKGROUND ELEMENTS: Mountain ridge (Dark and empty of the escaping group) — Its crest is seen from the distant elevated vantage, extending laterally through the upper frame; used as Carries the concentrated lights and confirms that the decoy drew attention uphill; Numerous red lights (Densely converged along the ridge); used as Primary distant focal pattern and restrained warning-red accent; Empty hillside distance (Separating the camera from the ridge) — The slope recedes upward toward the lit crest; used as Establishes scale and confirms the lights are far from the group.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Night darkness and the concentrated warning-red lights create a cool, muted long view with stark but restrained points of contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): It is night around the wet rock fissure below the secret exit, and the contact station is collapsing quietly behind it. Numerous red lights cluster along the empty mountaintop ridge.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "카메라는 어두운 숲을 넘어 먼 산 정상의 능선을 향함.",
    "built_space": "인공 구조물 없음. 바위 전경과 숲, 산 능선으로 이루어진 자연 지형.",
    "entities": "산 능선에 밀집된 수많은 붉은 불빛들 확인됨. 지시대로 인물은 없음.",
    "hard_violations": [],
    "physics": "바위와 나무들이 지형에 자연스럽게 고정되어 있음."
   },
   {
    "label": "B",
    "direction": "카메라는 바위 출구 벽면과 먼 능선을 동시에 향함.",
    "built_space": "바위벽에 삽입된 금속 문 1개와 지면에 놓인 장치 1개.",
    "entities": "먼 능선에 붉은 불빛들이 위치함. 인물 없음.",
    "hard_violations": [],
    "physics": "문과 바닥의 장치가 지면에 단단히 지지됨."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "요구된 원경 구도와 산등성이의 붉은 불빛을 정확히 구현했으며, 레퍼런스 구도를 복사하지 말라는 지시를 훌륭히 준수함."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "레퍼런스 이미지의 구도를 절대 복사하지 말라는 지시를 정면으로 위반하여 동일한 프레임을 그대로 사용함."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "카메라는 어두운 숲을 넘어 먼 산 정상의 능선을 향함.",
        "built_space": "인공 구조물 없음. 바위 전경과 숲, 산 능선으로 이루어진 자연 지형.",
        "entities": "산 능선에 밀집된 수많은 붉은 불빛들 확인됨. 지시대로 인물은 없음.",
        "hard_violations": [],
        "physics": "바위와 나무들이 지형에 자연스럽게 고정되어 있음."
       },
       {
        "label": "B",
        "direction": "카메라는 바위 출구 벽면과 먼 능선을 동시에 향함.",
        "built_space": "바위벽에 삽입된 금속 문 1개와 지면에 놓인 장치 1개.",
        "entities": "먼 능선에 붉은 불빛들이 위치함. 인물 없음.",
        "hard_violations": [],
        "physics": "문과 바닥의 장치가 지면에 단단히 지지됨."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "요구된 원경 구도와 산등성이의 붉은 불빛을 정확히 구현했으며, 레퍼런스 구도를 복사하지 말라는 지시를 훌륭히 준수함."
       },
       {
        "label": "B",
        "score": 3,
        "verdict_ko": "레퍼런스 이미지의 구도를 절대 복사하지 말라는 지시를 정면으로 위반하여 동일한 프레임을 그대로 사용함."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "카메라는 어두운 숲을 넘어 먼 산 정상의 능선을 향함.",
        "built_space": "인공 구조물 없음. 바위 전경과 숲, 산 능선으로 이루어진 자연 지형.",
        "entities": "산 능선에 밀집된 수많은 붉은 불빛들 확인됨. 지시대로 인물은 없음.",
        "hard_violations": [],
        "physics": "바위와 나무들이 지형에 자연스럽게 고정되어 있음."
       },
       {
        "label": "B",
        "direction": "카메라는 바위 출구 벽면과 먼 능선을 동시에 향함.",
        "built_space": "바위벽에 삽입된 금속 문 1개와 지면에 놓인 장치 1개.",
        "entities": "먼 능선에 붉은 불빛들이 위치함. 인물 없음.",
        "hard_violations": [],
        "physics": "문과 바닥의 장치가 지면에 단단히 지지됨."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "능선과 그 위에 밀집한 수많은 붉은 불빛을 상단 중앙의 원경으로 정확히 잡아, 숲과 빈 산비탈이 만드는 거리감까지 쇼트 지시와 가장 잘 일치한다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "붉은 불빛과 정확한 장소 요소는 충실하지만, 암반 출구와 가방이 전경을 지배해 능선의 원경을 핵심으로 삼으라는 권위적 프레이밍보다 넓고 산만하게 재구성됐다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "사람, 시선, 무기 또는 조준 물체는 없다. 붉은 불빛들은 좌측 상단에서 상단 중앙으로 이어지는 먼 능선 위에 모여 있으며, 빗줄기는 아래로 떨어진다. 우측 숲의 녹색 불빛도 보이지만 특정 대상을 향하지 않는다.",
        "built_space": "좌측 전경 암벽에 금속 프레임을 두른 은폐 출구 1개가 있고 내부는 돌무더기로 막혀 있다. 출구 앞 젖은 지면에 가방형 물체 1개가 놓여 있다. 구조와 습윤 암반은 위치 사진과 매우 가깝지만, 출구가 화면의 큰 비중을 차지해 원경 능선 중심 구도보다 전경 장소 소개에 가깝다.",
        "entities": "사람이나 얼굴은 없다. 젖은 암반, 침엽수림, 막힌 암반 출구, 전경 가방, 먼 산등성이, 다수의 붉은 점광원, 우측 녹색 점광원이 보인다. 붉은 불빛은 충분히 많고 능선에 밀집했으나 화면 중심보다 다소 왼쪽에 치우쳤다.",
        "hard_violations": [],
        "physics": "가방은 젖은 지면에 완전히 놓여 지지된다. 돌무더기는 출구 바닥과 서로 맞물려 지지되고, 붉은 불빛은 먼 능선 표면을 따라 배치되어 떠 있는 물체로 보이지 않는다. 빗줄기는 중력 방향으로 낙하하며 지지 없는 인물이나 물체는 없다."
       },
       {
        "label": "B",
        "direction": "사람, 시선, 무기 또는 조준 물체는 없다. 다수의 붉은 불빛은 화면 상단 중앙의 능선 마루를 따라 가로로 집중되어 있고, 빗줄기는 아래로 떨어진다.",
        "built_space": "화면 안에 인공 구조물이나 고정 설비는 없다. 하단의 젖은 암반과 침엽수림 너머로 비어 있는 넓은 산비탈이 후퇴하고, 상단을 횡단하는 먼 능선으로 이어진다. 출구를 보이지 않는 각도이지만 카메라 각도는 열려 있으며, 요구된 능선 원경에 맞는 동일 산악 지형과 재질을 유지한다.",
        "entities": "사람이나 얼굴은 없다. 젖은 바위, 침엽수림, 빈 산비탈, 어두운 정상 능선과 그 마루에 밀집한 수많은 붉은 불빛이 모두 명확히 보인다. 읽을 수 있는 글이나 그래픽 표시는 없다.",
        "hard_violations": [],
        "physics": "붉은 불빛은 능선 마루와 바로 아래 지면에 붙어 배치된 원거리 광원으로 읽혀 공중에 뜬 물체처럼 보이지 않는다. 암석과 나무는 지면에 고정되어 있고 빗줄기는 아래로 낙하한다. 지지 없는 인물이나 이동 물체는 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "능선과 그 위에 밀집한 수많은 붉은 불빛을 상단 중앙의 원경으로 정확히 잡아, 숲과 빈 산비탈이 만드는 거리감까지 쇼트 지시와 가장 잘 일치한다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "붉은 불빛과 정확한 장소 요소는 충실하지만, 암반 출구와 가방이 전경을 지배해 능선의 원경을 핵심으로 삼으라는 권위적 프레이밍보다 넓고 산만하게 재구성됐다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "사람, 시선, 무기 또는 조준 물체는 없다. 붉은 불빛들은 좌측 상단에서 상단 중앙으로 이어지는 먼 능선 위에 모여 있으며, 빗줄기는 아래로 떨어진다. 우측 숲의 녹색 불빛도 보이지만 특정 대상을 향하지 않는다.",
        "built_space": "좌측 전경 암벽에 금속 프레임을 두른 은폐 출구 1개가 있고 내부는 돌무더기로 막혀 있다. 출구 앞 젖은 지면에 가방형 물체 1개가 놓여 있다. 구조와 습윤 암반은 위치 사진과 매우 가깝지만, 출구가 화면의 큰 비중을 차지해 원경 능선 중심 구도보다 전경 장소 소개에 가깝다.",
        "entities": "사람이나 얼굴은 없다. 젖은 암반, 침엽수림, 막힌 암반 출구, 전경 가방, 먼 산등성이, 다수의 붉은 점광원, 우측 녹색 점광원이 보인다. 붉은 불빛은 충분히 많고 능선에 밀집했으나 화면 중심보다 다소 왼쪽에 치우쳤다.",
        "hard_violations": [],
        "physics": "가방은 젖은 지면에 완전히 놓여 지지된다. 돌무더기는 출구 바닥과 서로 맞물려 지지되고, 붉은 불빛은 먼 능선 표면을 따라 배치되어 떠 있는 물체로 보이지 않는다. 빗줄기는 중력 방향으로 낙하하며 지지 없는 인물이나 물체는 없다."
       },
       {
        "label": "A",
        "direction": "사람, 시선, 무기 또는 조준 물체는 없다. 다수의 붉은 불빛은 화면 상단 중앙의 능선 마루를 따라 가로로 집중되어 있고, 빗줄기는 아래로 떨어진다.",
        "built_space": "화면 안에 인공 구조물이나 고정 설비는 없다. 하단의 젖은 암반과 침엽수림 너머로 비어 있는 넓은 산비탈이 후퇴하고, 상단을 횡단하는 먼 능선으로 이어진다. 출구를 보이지 않는 각도이지만 카메라 각도는 열려 있으며, 요구된 능선 원경에 맞는 동일 산악 지형과 재질을 유지한다.",
        "entities": "사람이나 얼굴은 없다. 젖은 바위, 침엽수림, 빈 산비탈, 어두운 정상 능선과 그 마루에 밀집한 수많은 붉은 불빛이 모두 명확히 보인다. 읽을 수 있는 글이나 그래픽 표시는 없다.",
        "hard_violations": [],
        "physics": "붉은 불빛은 능선 마루와 바로 아래 지면에 붙어 배치된 원거리 광원으로 읽혀 공중에 뜬 물체처럼 보이지 않는다. 암석과 나무는 지면에 고정되어 있고 빗줄기는 아래로 낙하한다. 지지 없는 인물이나 이동 물체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt": "A"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt"
   ],
   "normalized": {
    "A": 2.0,
    "B": 1.095
   },
   "adjusted": {
    "A": 2.0,
    "B": 1.095
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt": "A"
   },
   "agreed": true
  },
  "totals": {
   "A": 2000,
   "B": 1095
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 2000,
    "verdict_ko": "요구된 원경 구도와 산등성이의 붉은 불빛을 정확히 구현했으며, 레퍼런스 구도를 복사하지 말라는 지시를 훌륭히 준수함."
   },
   {
    "label": "B",
    "score": 1095,
    "verdict_ko": "레퍼런스 이미지의 구도를 절대 복사하지 말라는 지시를 정면으로 위반하여 동일한 프레임을 그대로 사용함."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/37154490-a4ed-4b37-9092-489c054d8a72/episodes/cc1d3d0b-626e-4f20-bc33-f5f9e2c6ee7e/images/background_chain/L12B01.png",
    "asset_id": "2546bba5-27f9-48bb-9915-75d5afdc77c5",
    "role": "location_plate"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": true,
  "shot_run_uid": "06a9bdf0-a01b-7e7f-814b-b1f5ca5ecc23",
  "ref_mode": "플레이트만 (배경 전용)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S7sh3::cine": {
  "applied": true,
  "attempted_at": "2026-09-05T09:22:13.416978+00:00",
  "fingerprint": "1476bbefa1e35c83e76aa3aa860de511255d89399b9a4166014b22c7063b7596",
  "fingerprint_version": 2,
  "provider": "grok",
  "endpoint": "openrouter/chat-completions",
  "model": "x-ai/grok-imagine-image-2.0",
  "pack": "24.202608252115",
  "source_file": "S7sh3_sel.png",
  "source_sha256": "7c999f9e4d14c3c863a6d56a5a0ff9c7f70f2429fdc781f7777b50e829abad11",
  "file": "S7sh3_cine.png",
  "staged_sha256": "d5a6c9404c8c3c94d4c0f6cb27d9898a1ac69c1bc91249e1014909eb1aa1baa6",
  "latency_ms": 15203
 },
 "S7sh10::signage": {
  "fp": "123c6b73ba1eadb6",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S7sh10": {
  "input_fingerprint": "3e00faf2359e3338",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — Wyoming, United States, 2419; all people are English-speaking Americans unless stated. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 송신기 표면에 음각으로 새겨진 세 갈래 나무 문장의 초근접 클로즈업\n\nLOCATION (lock): Outside beside the wet rock fissure serving as the hidden mountainside exit. The tiny transmitter rests in a person's palm under the dark forest night. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: insert close-up on a detail\n- KEY BACKGROUND ELEMENTS: Tiny transmitter surface (Removed from the belt and bearing an engraved three-branched tree emblem) — The emblem-bearing face is angled steeply across camera, with the recessed three-branched tree mark visible at the center; used as Isolated evidence detail linking the transmitter to the Wyoming community; Engraved three-branched tree emblem (Clearly visible as a recessed mark) — The marked face is presented obliquely, allowing the three branches and engraved depth to remain legible; used as Primary macro focal detail.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Night-appropriate low-key illumination renders the engraved relief with cool, muted contrast and no newly specified source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The jumper belt’s inner seam has been cut open. A rice-grain-sized transmitter is exposed, its surface engraved with the Wyoming community’s three-branched tree emblem.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — Wyoming, United States, 2419; all people are English-speaking Americans unless stated. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 송신기 표면에 음각으로 새겨진 세 갈래 나무 문장의 초근접 클로즈업\n\nLOCATION (lock): Outside beside the wet rock fissure serving as the hidden mountainside exit. The tiny transmitter rests in a person's palm under the dark forest night. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: insert close-up on a detail\n- KEY BACKGROUND ELEMENTS: Tiny transmitter surface (Removed from the belt and bearing an engraved three-branched tree emblem) — The emblem-bearing face is angled steeply across camera, with the recessed three-branched tree mark visible at the center; used as Isolated evidence detail linking the transmitter to the Wyoming community; Engraved three-branched tree emblem (Clearly visible as a recessed mark) — The marked face is presented obliquely, allowing the three branches and engraved depth to remain legible; used as Primary macro focal detail.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Night-appropriate low-key illumination renders the engraved relief with cool, muted contrast and no newly specified source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The jumper belt’s inner seam has been cut open. A rice-grain-sized transmitter is exposed, its surface engraved with the Wyoming community’s three-branched tree emblem.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — Wyoming, United States, 2419; all people are English-speaking Americans unless stated. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 송신기 표면에 음각으로 새겨진 세 갈래 나무 문장의 초근접 클로즈업\n\nLOCATION (lock): Outside beside the wet rock fissure serving as the hidden mountainside exit. The tiny transmitter rests in a person's palm under the dark forest night. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: insert close-up on a detail\n- KEY BACKGROUND ELEMENTS: Tiny transmitter surface (Removed from the belt and bearing an engraved three-branched tree emblem) — The emblem-bearing face is angled steeply across camera, with the recessed three-branched tree mark visible at the center; used as Isolated evidence detail linking the transmitter to the Wyoming community; Engraved three-branched tree emblem (Clearly visible as a recessed mark) — The marked face is presented obliquely, allowing the three branches and engraved depth to remain legible; used as Primary macro focal detail.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: Night-appropriate low-key illumination renders the engraved relief with cool, muted contrast and no newly specified source.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The jumper belt’s inner seam has been cut open. A rice-grain-sized transmitter is exposed, its surface engraved with the Wyoming community’s three-branched tree emblem.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "카메라는 장갑 낀 손의 손바닥 위에 놓인 작은 금속 송신기를 비스듬히 내려다보고 있습니다.",
    "built_space": "어둡고 축축한 야외 공간으로, 레퍼런스의 바위 틈새와 유사한 젖고 거친 바위 표면이 배경에 아웃포커스되어 있습니다.",
    "entities": "장갑을 낀 사람의 손(일부 맨살 노출), 표면에 나무 문장이 음각된 알약 크기의 작은 금속 송신기. 프롬프트에 명시된 '세 갈래(three-branched)'가 아닌, 잔가지가 많은 나무 형태가 새겨져 있습니다.",
    "hard_violations": [],
    "physics": "금속 물체는 손바닥 중앙에 중력을 받아 안정적으로 놓여 있으며 손이 이를 잘 지탱하고 있습니다."
   },
   {
    "label": "B",
    "direction": "카메라는 흙이 묻은 손의 네 손가락 위에 가로로 얹혀 있는 직사각형의 금속 물체를 위에서 내려다보고 있습니다.",
    "built_space": "어두운 야외 지면으로, 젖은 바위 모서리와 땅에 떨어진 솔잎, 흙이 배경으로 보입니다.",
    "entities": "흙이 묻고 거친 사람의 손, 윗면에 기하학적인 나무 문장이 새겨진 직사각형의 금속 블록. '세 갈래'가 아닌 7개의 끝을 가진 나뭇가지 문장이며, 프롬프트의 '쌀알 크기(rice-grain-sized)' 지시와는 완전히 다르게 매우 큽니다.",
    "hard_violations": [],
    "physics": "크고 무거워 보이는 금속 물체가 손가락들 위에 놓여 지탱되고 있습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "'손바닥 위에 놓인 작은 송신기'라는 프롬프트의 구도와 크기를 비교적 잘 구현했고 젖은 바위 배경이 레퍼런스와 잘 어울리나, 음각된 문장이 지시된 '세 갈래'가 아닌 점이 감점 요소입니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "송신기가 '쌀알 크기'라는 명확한 지시를 무시하고 손가락들을 덮을 만큼 거대하게 묘사되었으며, 문장의 형태(세 갈래 아님)와 놓인 위치(손바닥이 아닌 손가락 위)도 프롬프트와 어긋납니다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "카메라는 장갑 낀 손의 손바닥 위에 놓인 작은 금속 송신기를 비스듬히 내려다보고 있습니다.",
        "built_space": "어둡고 축축한 야외 공간으로, 레퍼런스의 바위 틈새와 유사한 젖고 거친 바위 표면이 배경에 아웃포커스되어 있습니다.",
        "entities": "장갑을 낀 사람의 손(일부 맨살 노출), 표면에 나무 문장이 음각된 알약 크기의 작은 금속 송신기. 프롬프트에 명시된 '세 갈래(three-branched)'가 아닌, 잔가지가 많은 나무 형태가 새겨져 있습니다.",
        "hard_violations": [],
        "physics": "금속 물체는 손바닥 중앙에 중력을 받아 안정적으로 놓여 있으며 손이 이를 잘 지탱하고 있습니다."
       },
       {
        "label": "B",
        "direction": "카메라는 흙이 묻은 손의 네 손가락 위에 가로로 얹혀 있는 직사각형의 금속 물체를 위에서 내려다보고 있습니다.",
        "built_space": "어두운 야외 지면으로, 젖은 바위 모서리와 땅에 떨어진 솔잎, 흙이 배경으로 보입니다.",
        "entities": "흙이 묻고 거친 사람의 손, 윗면에 기하학적인 나무 문장이 새겨진 직사각형의 금속 블록. '세 갈래'가 아닌 7개의 끝을 가진 나뭇가지 문장이며, 프롬프트의 '쌀알 크기(rice-grain-sized)' 지시와는 완전히 다르게 매우 큽니다.",
        "hard_violations": [],
        "physics": "크고 무거워 보이는 금속 물체가 손가락들 위에 놓여 지탱되고 있습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "'손바닥 위에 놓인 작은 송신기'라는 프롬프트의 구도와 크기를 비교적 잘 구현했고 젖은 바위 배경이 레퍼런스와 잘 어울리나, 음각된 문장이 지시된 '세 갈래'가 아닌 점이 감점 요소입니다."
       },
       {
        "label": "B",
        "score": 4,
        "verdict_ko": "송신기가 '쌀알 크기'라는 명확한 지시를 무시하고 손가락들을 덮을 만큼 거대하게 묘사되었으며, 문장의 형태(세 갈래 아님)와 놓인 위치(손바닥이 아닌 손가락 위)도 프롬프트와 어긋납니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "A",
        "direction": "카메라는 장갑 낀 손의 손바닥 위에 놓인 작은 금속 송신기를 비스듬히 내려다보고 있습니다.",
        "built_space": "어둡고 축축한 야외 공간으로, 레퍼런스의 바위 틈새와 유사한 젖고 거친 바위 표면이 배경에 아웃포커스되어 있습니다.",
        "entities": "장갑을 낀 사람의 손(일부 맨살 노출), 표면에 나무 문장이 음각된 알약 크기의 작은 금속 송신기. 프롬프트에 명시된 '세 갈래(three-branched)'가 아닌, 잔가지가 많은 나무 형태가 새겨져 있습니다.",
        "hard_violations": [],
        "physics": "금속 물체는 손바닥 중앙에 중력을 받아 안정적으로 놓여 있으며 손이 이를 잘 지탱하고 있습니다."
       },
       {
        "label": "B",
        "direction": "카메라는 흙이 묻은 손의 네 손가락 위에 가로로 얹혀 있는 직사각형의 금속 물체를 위에서 내려다보고 있습니다.",
        "built_space": "어두운 야외 지면으로, 젖은 바위 모서리와 땅에 떨어진 솔잎, 흙이 배경으로 보입니다.",
        "entities": "흙이 묻고 거친 사람의 손, 윗면에 기하학적인 나무 문장이 새겨진 직사각형의 금속 블록. '세 갈래'가 아닌 7개의 끝을 가진 나뭇가지 문장이며, 프롬프트의 '쌀알 크기(rice-grain-sized)' 지시와는 완전히 다르게 매우 큽니다.",
        "hard_violations": [],
        "physics": "크고 무거워 보이는 금속 물체가 손가락들 위에 놓여 지탱되고 있습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 8,
        "verdict_ko": "송신기 표면과 중앙의 음각 나무 문장을 초근접으로 크게 포착하고 비스듬한 면의 각인 깊이도 선명해 권위 있는 숏 구성을 가장 충실히 구현했지만, 송신기가 쌀알 크기보다 현저히 크고 문장의 가지 수도 다소 모호하다."
       },
       {
        "label": "B",
        "score": 6,
        "verdict_ko": "젖은 손바닥 위 송신기와 야간의 젖은 암석 환경은 잘 맞지만, 프레임이 손과 소매까지 넓게 보여 핵심인 표면 음각의 초근접 인서트가 아니며 나무 문장도 세 갈래보다 많은 가지로 읽힌다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "시선, 무기, 이동체는 없다. 송신기의 문장 면은 카메라를 향하되 화면 대각선으로 비스듬히 놓여 있고, 중앙의 음각 표식이 직접 보인다.",
        "built_space": "초근접 프레임이라 숨겨진 출구의 문이나 균열 구조물은 보이지 않는다. 손 아래와 좌측 배경에 젖은 검은 암석 및 솔잎만 보이며, 고정 설비나 광학적으로 문제 되는 반사는 없다.",
        "entities": "사람의 얼굴이나 신체는 없고 물체를 받치는 한 손만 등장한다. 손바닥에는 마모된 직사각형 금속 송신기 하나가 있으며 중앙에 실제로 파인 듯한 나무형 음각이 있다. 다만 손가락과 비교하면 쌀알 크기가 아니라 수 센티미터급이고, 표식은 정확히 세 갈래라기보다 여러 세부 가지가 있는 도안으로 읽힌다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "송신기는 펼친 손가락과 손바닥 위에 완전히 접촉해 지지되며 떠 있지 않는다. 손의 자세와 물체의 안정된 놓임은 물체를 제시하는 실제 동작으로 가능하다."
       },
       {
        "label": "B",
        "direction": "시선, 무기, 이동체는 없다. 타원형 송신기의 문장 면이 카메라 쪽을 향하고 약간 비스듬히 놓여 있어 각인은 보이지만, 화면에서 작아 주된 초근접 표면으로 읽히지는 않는다.",
        "built_space": "배경에는 젖은 불규칙 암석이 보이지만 출구의 문틀이나 암벽 균열 같은 고정 구조는 프레임 밖이다. 손목 쪽에는 검은 소매 또는 장갑 가장자리 하나가 보이며, 불가능한 반사는 없다.",
        "entities": "얼굴이나 인물 전신 없이 젖은 손과 손목 일부만 나온다. 손바닥에는 작은 타원형 금속 송신기 하나가 있고 나무형 음각이 보인다. 그러나 표식은 명백히 세 갈래보다 많은 가지를 지니며, 송신기도 쌀알보다는 상당히 크다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "송신기는 젖은 손바닥 중앙에 닿아 지지되고 있으므로 부유하지 않는다. 펼친 손이 물체를 받치는 자세는 물리적으로 가능하고 빗물도 피부와 암석 표면에 자연스럽게 붙어 있다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 8,
        "verdict_ko": "송신기 표면과 중앙의 음각 나무 문장을 초근접으로 크게 포착하고 비스듬한 면의 각인 깊이도 선명해 권위 있는 숏 구성을 가장 충실히 구현했지만, 송신기가 쌀알 크기보다 현저히 크고 문장의 가지 수도 다소 모호하다."
       },
       {
        "label": "A",
        "score": 6,
        "verdict_ko": "젖은 손바닥 위 송신기와 야간의 젖은 암석 환경은 잘 맞지만, 프레임이 손과 소매까지 넓게 보여 핵심인 표면 음각의 초근접 인서트가 아니며 나무 문장도 세 갈래보다 많은 가지로 읽힌다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "시선, 무기, 이동체는 없다. 송신기의 문장 면은 카메라를 향하되 화면 대각선으로 비스듬히 놓여 있고, 중앙의 음각 표식이 직접 보인다.",
        "built_space": "초근접 프레임이라 숨겨진 출구의 문이나 균열 구조물은 보이지 않는다. 손 아래와 좌측 배경에 젖은 검은 암석 및 솔잎만 보이며, 고정 설비나 광학적으로 문제 되는 반사는 없다.",
        "entities": "사람의 얼굴이나 신체는 없고 물체를 받치는 한 손만 등장한다. 손바닥에는 마모된 직사각형 금속 송신기 하나가 있으며 중앙에 실제로 파인 듯한 나무형 음각이 있다. 다만 손가락과 비교하면 쌀알 크기가 아니라 수 센티미터급이고, 표식은 정확히 세 갈래라기보다 여러 세부 가지가 있는 도안으로 읽힌다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "송신기는 펼친 손가락과 손바닥 위에 완전히 접촉해 지지되며 떠 있지 않는다. 손의 자세와 물체의 안정된 놓임은 물체를 제시하는 실제 동작으로 가능하다."
       },
       {
        "label": "A",
        "direction": "시선, 무기, 이동체는 없다. 타원형 송신기의 문장 면이 카메라 쪽을 향하고 약간 비스듬히 놓여 있어 각인은 보이지만, 화면에서 작아 주된 초근접 표면으로 읽히지는 않는다.",
        "built_space": "배경에는 젖은 불규칙 암석이 보이지만 출구의 문틀이나 암벽 균열 같은 고정 구조는 프레임 밖이다. 손목 쪽에는 검은 소매 또는 장갑 가장자리 하나가 보이며, 불가능한 반사는 없다.",
        "entities": "얼굴이나 인물 전신 없이 젖은 손과 손목 일부만 나온다. 손바닥에는 작은 타원형 금속 송신기 하나가 있고 나무형 음각이 보인다. 그러나 표식은 명백히 세 갈래보다 많은 가지를 지니며, 송신기도 쌀알보다는 상당히 크다. 읽을 수 있는 글자는 없다.",
        "hard_violations": [],
        "physics": "송신기는 젖은 손바닥 중앙에 닿아 지지되고 있으므로 부유하지 않는다. 펼친 손이 물체를 받치는 자세는 물리적으로 가능하고 빗물도 피부와 암석 표면에 자연스럽게 붙어 있다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt"
   ],
   "slot_winner_match": false,
   "slot_winner": {
    "gemini-pro": "A",
    "gpt": "B"
   },
   "route": "cross_slot_combined"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt"
   ],
   "normalized": {
    "A": 1.75,
    "B": 1.571
   },
   "adjusted": {
    "A": 1.75,
    "B": 1.571
   },
   "violations": {},
   "per_model_winner": {
    "gemini-pro": "A",
    "gpt": "B"
   },
   "agreed": false
  },
  "totals": {
   "A": 1750,
   "B": 1571
  },
  "selected": "A",
  "ranking": [
   "A",
   "B"
  ],
  "verdicts": [
   {
    "label": "A",
    "score": 1750,
    "verdict_ko": "'손바닥 위에 놓인 작은 송신기'라는 프롬프트의 구도와 크기를 비교적 잘 구현했고 젖은 바위 배경이 레퍼런스와 잘 어울리나, 음각된 문장이 지시된 '세 갈래'가 아닌 점이 감점 요소입니다."
   },
   {
    "label": "B",
    "score": 1571,
    "verdict_ko": "송신기가 '쌀알 크기'라는 명확한 지시를 무시하고 손가락들을 덮을 만큼 거대하게 묘사되었으며, 문장의 형태(세 갈래 아님)와 놓인 위치(손바닥이 아닌 손가락 위)도 프롬프트와 어긋납니다."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/37154490-a4ed-4b37-9092-489c054d8a72/episodes/cc1d3d0b-626e-4f20-bc33-f5f9e2c6ee7e/images/background_chain/L12B01.png",
    "asset_id": "2546bba5-27f9-48bb-9915-75d5afdc77c5",
    "role": "location_plate"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": true,
  "shot_run_uid": "06a9bdf5-4c21-762f-9ee1-b0c03ae24df8",
  "ref_mode": "플레이트만 (배경 전용)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S7sh10::cine": {
  "applied": true,
  "attempted_at": "2026-09-05T09:23:34.503441+00:00",
  "fingerprint": "9dc2535bdbb2e7f49d011fac4073cc16f0a5995bf6fb4f24975fa6494c2fef75",
  "fingerprint_version": 2,
  "provider": "grok",
  "endpoint": "openrouter/chat-completions",
  "model": "x-ai/grok-imagine-image-2.0",
  "pack": "24.202608252115",
  "source_file": "S7sh10_sel.png",
  "source_sha256": "67fa94b0d622ac6aa4cf0d522b795ce961d5b92f93565d720358b97420ec7f84",
  "file": "S7sh10_cine.png",
  "staged_sha256": "86d1ee2956851e61ea892d1e9e3e9ca493dee9d9917bbd088c484ac1f242404b",
  "latency_ms": 16177
 },
 "S7sh13::signage": {
  "fp": "e2a43404caf2c2ca",
  "inscriptions": [],
  "cues": [],
  "dropped": []
 },
 "S7sh13": {
  "input_fingerprint": "340671f4011c9148",
  "prompt": "Create ONE FINAL photorealistic live-action film still of the moment below — Wyoming, United States, 2419; all people are English-speaking Americans unless stated. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 어두운 숲 너머로 세 갈래 나무 문양의 신호등 불빛이 선명하게 켜진 찰나\n\nLOCATION (lock): Outside in the nighttime forest beyond the concealed mountainside exit. A distant signal light bearing a three-branched tree emblem flashes clearly among the dark woods. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: three-branched signal light in the middle-center of the frame, background.\n- KEY BACKGROUND ELEMENTS: Distant three-branched signal light (Fully illuminated for a single instant before switching off) — Its signal-bearing face is directed toward the group's position and visibly forms the three-branched tree emblem; used as Primary distant reveal and visual match to the transmitter engraving; Forest between group and signal (Dark, with the signal visible beyond it) — The view passes between the trees toward the distant illuminated point; used as Establishes concealment, distance, and the signal's hidden origin.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The dark night forest remains cool and muted while the briefly illuminated signal provides the shot's precise point of contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The rice-grain-sized transmitter retains the engraved three-branched tree emblem, and the jumper belt remains cut along its inner seam. The distant forest is dark beyond the secret exit. A matching three-branched-tree signal light shines clearly once in the distant forest before going dark.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
  "roll_prompts": {
   "A": "Create ONE FINAL photorealistic live-action film still of the moment below — Wyoming, United States, 2419; all people are English-speaking Americans unless stated. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 어두운 숲 너머로 세 갈래 나무 문양의 신호등 불빛이 선명하게 켜진 찰나\n\nLOCATION (lock): Outside in the nighttime forest beyond the concealed mountainside exit. A distant signal light bearing a three-branched tree emblem flashes clearly among the dark woods. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: three-branched signal light in the middle-center of the frame, background.\n- KEY BACKGROUND ELEMENTS: Distant three-branched signal light (Fully illuminated for a single instant before switching off) — Its signal-bearing face is directed toward the group's position and visibly forms the three-branched tree emblem; used as Primary distant reveal and visual match to the transmitter engraving; Forest between group and signal (Dark, with the signal visible beyond it) — The view passes between the trees toward the distant illuminated point; used as Establishes concealment, distance, and the signal's hidden origin.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The dark night forest remains cool and muted while the briefly illuminated signal provides the shot's precise point of contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The rice-grain-sized transmitter retains the engraved three-branched tree emblem, and the jumper belt remains cut along its inner seam. The distant forest is dark beyond the secret exit. A matching three-branched-tree signal light shines clearly once in the distant forest before going dark.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.",
   "B": "Create ONE FINAL photorealistic live-action film still of the moment below — Wyoming, United States, 2419; all people are English-speaking Americans unless stated. TIME OF DAY (lock): night.\n\nSHOT TEXT (authoritative, Korean): 어두운 숲 너머로 세 갈래 나무 문양의 신호등 불빛이 선명하게 켜진 찰나\n\nLOCATION (lock): Outside in the nighttime forest beyond the concealed mountainside exit. A distant signal light bearing a three-branched tree emblem flashes clearly among the dark woods. The shot takes place here — the attached LOCATION PHOTOGRAPH shows the exact spot.\n\nCAMERA & FRAME (follow exactly — this is the composition authority for this still; the reference images supply identity and place, never the framing):\n- FRAMING SCALE: wide shot\n- FRAME LAYOUT: three-branched signal light in the middle-center of the frame, background.\n- KEY BACKGROUND ELEMENTS: Distant three-branched signal light (Fully illuminated for a single instant before switching off) — Its signal-bearing face is directed toward the group's position and visibly forms the three-branched tree emblem; used as Primary distant reveal and visual match to the transmitter engraving; Forest between group and signal (Dark, with the signal visible beyond it) — The view passes between the trees toward the distant illuminated point; used as Establishes concealment, distance, and the signal's hidden origin.\nCompose the frame exactly as specified above — subject scale and screen placement. The viewing angle is deliberately left open here; choose the angle that serves the shot. Keep true physical scale between people and background elements: fixtures and distant objects occupy only the small screen area their real size and distance dictate; never enlarge a background object into a foreground presence.\n\nLIGHTING & MOOD (these govern light, color and surface rendering):\n- LIGHTING & MOOD: The dark night forest remains cool and muted while the briefly illuminated signal provides the shot's precise point of contrast.\n- MATERIAL REALISM: every object and surface must read as a real,\n  physical material — correct texture, weight, wear and light response\n  (metal reflects, fabric drapes and creases, liquid is glossy and\n  pools, painted or drawn marks sit ON a surface and follow its\n  curvature and lighting). Nothing may look like a flat sticker, a\n  doodle or a graphic overlay pasted onto the frame.\n\nREALIZE FIGURATIVE LANGUAGE AS A LIVE-ACTION SHOT: the Korean shot\ntext may describe characters metaphorically, figuratively or with\nexaggeration. Photograph what a real movie camera would actually\nrecord on a physical set — exaggerated or figurative impressions\nbecome realistic staging within whatever this brief already fixes,\nnot literal fantasy imagery.\n\nPROPS FACE THE RIGHT WAY: every handheld or used object must be\noriented exactly as its real-world use requires. A person reading,\nwatching or operating something (a phone, a photograph, a paper,\nany device) has its functional side — screen, front, page — facing\nTHEIR OWN eyes; the camera then sees whatever side the staging\ngeometry implies (often its back). Show the functional side to the\ncamera ONLY when the shot text itself stages it toward the viewer.\nNever flip, mirror or reverse an object's front and back.\n\nDRAWN MARKS KEEP THEIR SHAPE: any mark the text describes as drawn,\npainted or traced (a circle, a line, a symbol) is a STROKE sitting on\nthe surface — an outline whose interior still shows the underlying\nsurface (wall, skin, paper). Render its stated shape faithfully: a\ndrawn circle stays an open ring of brush-width, never filled into a\nsolid disc, unless the text explicitly says it is filled.\n\nCARRIED STATE (persist exactly — must match the neighbouring shots of this scene): The rice-grain-sized transmitter retains the engraved three-branched tree emblem, and the jumper belt remains cut along its inner seam. The distant forest is dark beyond the secret exit. A matching three-branched-tree signal light shines clearly once in the distant forest before going dark.\n\nNO PEOPLE IN THIS SHOT: the shot text shows only the place and its state — no living person appears in frame, and no face, and no figure standing, sitting or passing through. Do not invent someone to occupy the space.\n\nOne thing this rule does NOT forbid: the hand or arm of whoever is doing what the shot text describes. If the moment is an object being held, carried, operated, read, aimed, worn or handed over, the hand doing it belongs in frame, resting on or gripping that object exactly as its real use requires. An object photographed floating unsupported, with the hand that must be holding it erased, is a worse failure than the person you were avoiding — render the hand.\n\nNo readable writing anywhere in this image. Surfaces that would carry writing may be present, but nothing a viewer could read: stage every one of them out of legibility — a hand across, an oblique angle, shallow focus, or simply turned away. No caption, subtitle, watermark, logo or overlay.\n\nCOMPOSITION VARIATION (this candidate only): keep the place, the people and the moment exactly as specified, and keep from the CAMERA & FRAME contract its subject scale, screen placement and key background elements exactly. For this candidate only, its camera angle and height are deliberately left open: choose a camera position distinctly different from the obvious one, as a film director picking a second setup on the same blocking. Stay true to any explicit framing the SHOT TEXT itself dictates."
  },
  "shot_run_spend_attempt_count": 1,
  "readings": [
   {
    "label": "A",
    "direction": "카메라는 빽빽하게 늘어선 침엽수림 내부를 가로질러, 나무 기둥에 인접해 있는 빛나는 문양을 향하고 있습니다.",
    "built_space": "화면 전체가 나무 기둥들로 가득 차 있으며, 레퍼런스에서 확인되는 바위산 비탈이나 진흙 바닥 같은 지형적 특징을 전혀 찾아볼 수 없는 평탄한 숲 내부입니다.",
    "entities": "세 갈래 나무 문양을 띤 밝은 신호가 중앙에 보입니다. 인물은 등장하지 않습니다.",
    "hard_violations": [
     "[gemini-pro] 지정된 로케이션(바위가 있는 산비탈 입구 주변)의 지형적 특성과 재질(돌, 흙)이 완전히 누락된 채 다른 장소처럼 연출됨."
    ],
    "physics": "신호등 형체가 나무 기둥 겉면에 다소 모호하게 부착되거나 떠 있는 것처럼 보이며, 지지 기반이 명확하지 않습니다."
   },
   {
    "label": "B",
    "direction": "카메라는 숲의 나무들 사이 열린 공간을 지나, 원경의 산등성이에서 밝게 빛나는 세 갈래 문양의 신호를 향하고 있습니다.",
    "built_space": "전경 양측에 레퍼런스 사진과 질감이 일치하는 젖은 바위와 진흙 지면이 배치되어 있으며, 그 너머로 어두운 침엽수림이 자연스럽게 펼쳐져 지정된 위치의 공간감을 잘 살렸습니다.",
    "entities": "지시된 형태의 세 갈래 나무 문양 신호등이 프레임 중앙 원경에서 밝게 빛나고 있습니다. 화면 내에 사람은 존재하지 않습니다.",
    "hard_violations": [],
    "physics": "원경의 신호등은 언덕 지면에 단단히 고정되어 있으며, 비가 내리는 자연스러운 환경 속에 놓여 있습니다."
   }
  ],
  "cross_model_order": {
   "policy": "cross_model_order_v2_slot0_fwd_slot1_rev_parallel2",
   "slots": [
    {
     "model": "gemini-pro",
     "order": "forward",
     "display_to_canonical": {
      "A": "A",
      "B": "B"
     },
     "raw": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "샷 텍스트에 명시된 '어두운 숲 너머로' 빛나는 원경의 신호등을 정확한 거리감으로 구현했으며, 레퍼런스의 젖은 바위와 진흙 지형까지 충실히 반영하여 위치적 일관성을 유지했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "신호등이 '숲 너머' 원경이 아닌 숲 내부의 가까운 나무에 위치해 프롬프트가 요구하는 거리감과 스케일을 상실했으며, 레퍼런스의 지형적 특성(바위산 비탈)도 완전히 누락되었습니다."
       }
      ],
      "readings": [
       {
        "label": "B",
        "direction": "카메라는 숲의 나무들 사이 열린 공간을 지나, 원경의 산등성이에서 밝게 빛나는 세 갈래 문양의 신호를 향하고 있습니다.",
        "built_space": "전경 양측에 레퍼런스 사진과 질감이 일치하는 젖은 바위와 진흙 지면이 배치되어 있으며, 그 너머로 어두운 침엽수림이 자연스럽게 펼쳐져 지정된 위치의 공간감을 잘 살렸습니다.",
        "entities": "지시된 형태의 세 갈래 나무 문양 신호등이 프레임 중앙 원경에서 밝게 빛나고 있습니다. 화면 내에 사람은 존재하지 않습니다.",
        "hard_violations": [],
        "physics": "원경의 신호등은 언덕 지면에 단단히 고정되어 있으며, 비가 내리는 자연스러운 환경 속에 놓여 있습니다."
       },
       {
        "label": "A",
        "direction": "카메라는 빽빽하게 늘어선 침엽수림 내부를 가로질러, 나무 기둥에 인접해 있는 빛나는 문양을 향하고 있습니다.",
        "built_space": "화면 전체가 나무 기둥들로 가득 차 있으며, 레퍼런스에서 확인되는 바위산 비탈이나 진흙 바닥 같은 지형적 특징을 전혀 찾아볼 수 없는 평탄한 숲 내부입니다.",
        "entities": "세 갈래 나무 문양을 띤 밝은 신호가 중앙에 보입니다. 인물은 등장하지 않습니다.",
        "hard_violations": [
         "지정된 로케이션(바위가 있는 산비탈 입구 주변)의 지형적 특성과 재질(돌, 흙)이 완전히 누락된 채 다른 장소처럼 연출됨."
        ],
        "physics": "신호등 형체가 나무 기둥 겉면에 다소 모호하게 부착되거나 떠 있는 것처럼 보이며, 지지 기반이 명확하지 않습니다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "샷 텍스트에 명시된 '어두운 숲 너머로' 빛나는 원경의 신호등을 정확한 거리감으로 구현했으며, 레퍼런스의 젖은 바위와 진흙 지형까지 충실히 반영하여 위치적 일관성을 유지했습니다."
       },
       {
        "label": "A",
        "score": 3,
        "verdict_ko": "신호등이 '숲 너머' 원경이 아닌 숲 내부의 가까운 나무에 위치해 프롬프트가 요구하는 거리감과 스케일을 상실했으며, 레퍼런스의 지형적 특성(바위산 비탈)도 완전히 누락되었습니다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "카메라는 숲의 나무들 사이 열린 공간을 지나, 원경의 산등성이에서 밝게 빛나는 세 갈래 문양의 신호를 향하고 있습니다.",
        "built_space": "전경 양측에 레퍼런스 사진과 질감이 일치하는 젖은 바위와 진흙 지면이 배치되어 있으며, 그 너머로 어두운 침엽수림이 자연스럽게 펼쳐져 지정된 위치의 공간감을 잘 살렸습니다.",
        "entities": "지시된 형태의 세 갈래 나무 문양 신호등이 프레임 중앙 원경에서 밝게 빛나고 있습니다. 화면 내에 사람은 존재하지 않습니다.",
        "hard_violations": [],
        "physics": "원경의 신호등은 언덕 지면에 단단히 고정되어 있으며, 비가 내리는 자연스러운 환경 속에 놓여 있습니다."
       },
       {
        "label": "A",
        "direction": "카메라는 빽빽하게 늘어선 침엽수림 내부를 가로질러, 나무 기둥에 인접해 있는 빛나는 문양을 향하고 있습니다.",
        "built_space": "화면 전체가 나무 기둥들로 가득 차 있으며, 레퍼런스에서 확인되는 바위산 비탈이나 진흙 바닥 같은 지형적 특징을 전혀 찾아볼 수 없는 평탄한 숲 내부입니다.",
        "entities": "세 갈래 나무 문양을 띤 밝은 신호가 중앙에 보입니다. 인물은 등장하지 않습니다.",
        "hard_violations": [
         "지정된 로케이션(바위가 있는 산비탈 입구 주변)의 지형적 특성과 재질(돌, 흙)이 완전히 누락된 채 다른 장소처럼 연출됨."
        ],
        "physics": "신호등 형체가 나무 기둥 겉면에 다소 모호하게 부착되거나 떠 있는 것처럼 보이며, 지지 기반이 명확하지 않습니다."
       }
      ]
     },
     "ok": true
    },
    {
     "model": "gpt",
     "order": "reverse",
     "display_to_canonical": {
      "A": "B",
      "B": "A"
     },
     "raw": {
      "winner": "A",
      "ranking": [
       "A",
       "B"
      ],
      "verdicts": [
       {
        "label": "A",
        "score": 9,
        "verdict_ko": "신호가 정확히 화면 중앙의 먼 배경에 작게 놓이고, 젖은 암반과 침엽수림이 위치 사진의 산악 지형을 충실히 이어 받아 권위 있는 와이드숏 구성을 가장 잘 실현했다."
       },
       {
        "label": "B",
        "score": 7,
        "verdict_ko": "세 갈래 신호의 방향과 물리적 설치는 명확하지만 신호가 다소 크게 보이고 암반 산악 지형이 사라진 일반적인 밀림 구도로 바뀌어 장소 및 원거리 스케일 충실도가 A보다 낮다."
       }
      ],
      "readings": [
       {
        "label": "A",
        "direction": "사람·시선·무기는 없다. 세 갈래 신호의 발광면은 카메라와 출구 쪽 관찰 위치를 정면으로 향해 문양이 온전히 보이며, 숲 사이 시선축이 화면 정중앙의 신호에 닿는다.",
        "built_space": "인공 구조물은 먼 신호 하나뿐이며 출구 문이나 다른 설비는 프레임에 없다. 전경과 양옆에는 젖은 암반, 나무줄기와 수풀이 있고 그 사이 골짜기 너머 능선에 신호가 배치되어 위치 사진의 암석성 산악 침엽수림과 잘 연결된다. 반사 장면은 없다.",
        "entities": "사람·얼굴·신체는 전혀 없다. 화면 중앙 먼 곳에 밝은 세 갈래 나무 문양의 신호 하나가 있으며, 어두운 숲 너머에서 한순간 완전히 점등된 상태로 읽힌다. 송신기와 절단된 벨트는 이 장소 중심 와이드숏에 등장하지 않는다. 읽을 수 있는 글자나 로고도 없다.",
        "hard_violations": [],
        "physics": "나무와 수풀은 지면 및 암반에 뿌리내려 있고 움직이거나 공중에 뜬 물체는 없다. 신호의 빛나는 세로 몸체는 능선 표면과 맞닿아 있으며 하부 고정부는 거리와 능선에 가려진 것으로 읽혀, 명백한 무지지 부유로 보이지 않는다."
       },
       {
        "label": "B",
        "direction": "사람·시선·무기는 없다. 세 갈래 신호의 발광면은 카메라 쪽을 정면으로 향하며 숲 사이의 관찰축에 맞아 문양이 선명하게 보인다.",
        "built_space": "인공 설비는 중앙에서 약간 왼쪽에 있는 신호 한 개이며, 수직 기둥과 뒤쪽의 좁은 지지 프레임이 보인다. 주위는 빽빽한 침엽수림이지만 위치 사진의 특징적인 젖은 암반 사면과 산악 출구 주변 지형은 거의 보이지 않아 정확한 장소 연결성이 약하다. 반사 장면은 없다.",
        "entities": "사람·얼굴·신체는 없다. 세 갈래 나무 문양의 점등 신호 하나와 어두운 숲이 보이고 문자는 없다. 신호 문양은 맞지만 원거리 배경물치고 상대적으로 크고 설치물이 두드러진다. 송신기와 벨트는 보이지 않는다.",
        "hard_violations": [],
        "physics": "신호는 눈에 보이는 수직 기둥과 프레임에 고정되어 있어 물리적으로 지지된다. 나무들은 지면에 뿌리내린 상태이며 공중에 뜨거나 지원 없이 움직이는 물체는 없다."
       }
      ],
      "all_candidates_fail": false
     },
     "normalized": {
      "winner": "B",
      "ranking": [
       "B",
       "A"
      ],
      "verdicts": [
       {
        "label": "B",
        "score": 9,
        "verdict_ko": "신호가 정확히 화면 중앙의 먼 배경에 작게 놓이고, 젖은 암반과 침엽수림이 위치 사진의 산악 지형을 충실히 이어 받아 권위 있는 와이드숏 구성을 가장 잘 실현했다."
       },
       {
        "label": "A",
        "score": 7,
        "verdict_ko": "세 갈래 신호의 방향과 물리적 설치는 명확하지만 신호가 다소 크게 보이고 암반 산악 지형이 사라진 일반적인 밀림 구도로 바뀌어 장소 및 원거리 스케일 충실도가 A보다 낮다."
       }
      ],
      "all_candidates_fail": false,
      "readings": [
       {
        "label": "B",
        "direction": "사람·시선·무기는 없다. 세 갈래 신호의 발광면은 카메라와 출구 쪽 관찰 위치를 정면으로 향해 문양이 온전히 보이며, 숲 사이 시선축이 화면 정중앙의 신호에 닿는다.",
        "built_space": "인공 구조물은 먼 신호 하나뿐이며 출구 문이나 다른 설비는 프레임에 없다. 전경과 양옆에는 젖은 암반, 나무줄기와 수풀이 있고 그 사이 골짜기 너머 능선에 신호가 배치되어 위치 사진의 암석성 산악 침엽수림과 잘 연결된다. 반사 장면은 없다.",
        "entities": "사람·얼굴·신체는 전혀 없다. 화면 중앙 먼 곳에 밝은 세 갈래 나무 문양의 신호 하나가 있으며, 어두운 숲 너머에서 한순간 완전히 점등된 상태로 읽힌다. 송신기와 절단된 벨트는 이 장소 중심 와이드숏에 등장하지 않는다. 읽을 수 있는 글자나 로고도 없다.",
        "hard_violations": [],
        "physics": "나무와 수풀은 지면 및 암반에 뿌리내려 있고 움직이거나 공중에 뜬 물체는 없다. 신호의 빛나는 세로 몸체는 능선 표면과 맞닿아 있으며 하부 고정부는 거리와 능선에 가려진 것으로 읽혀, 명백한 무지지 부유로 보이지 않는다."
       },
       {
        "label": "A",
        "direction": "사람·시선·무기는 없다. 세 갈래 신호의 발광면은 카메라 쪽을 정면으로 향하며 숲 사이의 관찰축에 맞아 문양이 선명하게 보인다.",
        "built_space": "인공 설비는 중앙에서 약간 왼쪽에 있는 신호 한 개이며, 수직 기둥과 뒤쪽의 좁은 지지 프레임이 보인다. 주위는 빽빽한 침엽수림이지만 위치 사진의 특징적인 젖은 암반 사면과 산악 출구 주변 지형은 거의 보이지 않아 정확한 장소 연결성이 약하다. 반사 장면은 없다.",
        "entities": "사람·얼굴·신체는 없다. 세 갈래 나무 문양의 점등 신호 하나와 어두운 숲이 보이고 문자는 없다. 신호 문양은 맞지만 원거리 배경물치고 상대적으로 크고 설치물이 두드러진다. 송신기와 벨트는 보이지 않는다.",
        "hard_violations": [],
        "physics": "신호는 눈에 보이는 수직 기둥과 프레임에 고정되어 있어 물리적으로 지지된다. 나무들은 지면에 뿌리내린 상태이며 공중에 뜨거나 지원 없이 움직이는 물체는 없다."
       }
      ]
     },
     "ok": true
    }
   ],
   "models": [
    "gemini-pro",
    "gpt"
   ],
   "slot_winner_match": true,
   "slot_winner": {
    "gemini-pro": "B",
    "gpt": "B"
   },
   "route": "cross_slot_agree"
  },
  "dual": {
   "models": [
    "gemini-pro",
    "gpt"
   ],
   "normalized": {
    "A": 1.111,
    "B": 2.0
   },
   "adjusted": {
    "A": 0.861,
    "B": 2.0
   },
   "violations": {
    "A": [
     "[gemini-pro] 지정된 로케이션(바위가 있는 산비탈 입구 주변)의 지형적 특성과 재질(돌, 흙)이 완전히 누락된 채 다른 장소처럼 연출됨."
    ]
   },
   "per_model_winner": {
    "gemini-pro": "B",
    "gpt": "B"
   },
   "agreed": true
  },
  "totals": {
   "B": 2000,
   "A": 861
  },
  "selected": "B",
  "ranking": [
   "B",
   "A"
  ],
  "verdicts": [
   {
    "label": "B",
    "score": 2000,
    "verdict_ko": "샷 텍스트에 명시된 '어두운 숲 너머로' 빛나는 원경의 신호등을 정확한 거리감으로 구현했으며, 레퍼런스의 젖은 바위와 진흙 지형까지 충실히 반영하여 위치적 일관성을 유지했습니다."
   },
   {
    "label": "A",
    "score": 861,
    "verdict_ko": "신호등이 '숲 너머' 원경이 아닌 숲 내부의 가까운 나무에 위치해 프롬프트가 요구하는 거리감과 스케일을 상실했으며, 레퍼런스의 지형적 특성(바위산 비탈)도 완전히 누락되었습니다.  ★위반: [gemini-pro] 지정된 로케이션(바위가 있는 산비탈 입구 주변)의 지형적 특성과 재질(돌, 흙)이 완전히 누락된 채 다른 장소처럼 연출됨."
   }
  ],
  "refs": [
   {
    "label": "LOCATION PHOTOGRAPH — the exact place of this shot: its architecture, materials, fixed features and lighting mood are spatial truth; stage the moment inside this place. Never copy its camera framing.",
    "path": "/Users/manta/Documents/Projects/TheRoad-I1/projects/37154490-a4ed-4b37-9092-489c054d8a72/episodes/cc1d3d0b-626e-4f20-bc33-f5f9e2c6ee7e/images/background_chain/L12B01.png",
    "asset_id": "2546bba5-27f9-48bb-9915-75d5afdc77c5",
    "role": "location_plate"
   }
  ],
  "critique_skipped": true,
  "shot_run_produced": true,
  "shot_run_uid": "06a9bdfa-6d33-7ebb-afdd-78e25dddf765",
  "ref_mode": "플레이트만 (배경 전용)",
  "share_plan": {
   "ref_plan": "background"
  }
 },
 "S7sh13::cine": {
  "applied": true,
  "attempted_at": "2026-09-05T09:24:44.414159+00:00",
  "fingerprint": "8327be069565d3edf651ebfd177d01434c324adfe2224c3b2b309c910592d2a7",
  "fingerprint_version": 2,
  "provider": "grok",
  "endpoint": "openrouter/chat-completions",
  "model": "x-ai/grok-imagine-image-2.0",
  "pack": "24.202608252115",
  "source_file": "S7sh13_sel.png",
  "source_sha256": "697c35a9f8dec1fe6b0b19972471de5b765d4dc1eda62789e9d9e7fee4db9ee7",
  "file": "S7sh13_cine.png",
  "staged_sha256": "b601ad410e964ab454edc48073aa8aa0ed8fc7b55a54db8baf21d318929698f7",
  "latency_ms": 11959
 }
}