{"chunk_end": 104, "chunk_start": 1, "chunk_summary": "No actionable findings; this Alembic migration only defines closed-world schema/index identifiers and contains no semantic string judgment or scenario-dependent prompt pollution.", "duration_ms": 4850, "findings": [], "path": "backend/alembic/versions/001_add_indexes.py", "scan_kind": "python", "sha256": "443d404214d268646a852a16b9fb0dae2be7d48f7e5246ca1280fc7ed869aed2"}
{"chunk_end": 70, "chunk_start": 1, "chunk_summary": "No actionable findings; this Alembic environment file only configures migration metadata and database connectivity.", "duration_ms": 5152, "findings": [], "path": "backend/alembic/env.py", "scan_kind": "python", "sha256": "0ebc7e60b781371151de2a3e04377b883bf6dc5258eec68551ce6d915acb66c0"}
{"chunk_end": 65, "chunk_start": 1, "chunk_summary": "No actionable findings.", "duration_ms": 8981, "findings": [], "path": "backend/alembic/versions/003_resume_integrity.py", "scan_kind": "python", "sha256": "8ce8dd8d21121f9b607f2ad1c9dac6df425a9f83d9b80fb807be0988d41667f9"}
{"chunk_end": 48, "chunk_start": 1, "chunk_summary": "No actionable findings.", "duration_ms": 7513, "findings": [], "path": "backend/alembic/versions/006_llm_call_log_metadata_json.py", "scan_kind": "python", "sha256": "1b72ded9be069eae18d32f41b6f416b9804e4877927fce0731f13475004634bf"}
{"chunk_end": 46, "chunk_start": 1, "chunk_summary": "No actionable findings; this chunk only contains closed-world auth route/cookie handling and no open-world semantic string judgments or scenario-polluting prompt logic.", "duration_ms": 5800, "findings": [], "path": "backend/app/api/v1/auth.py", "scan_kind": "python", "sha256": "897f59b16fd7a5874d60124f0bfeb419ae516004bf924b003684733adf31e85e"}
{"chunk_end": 60, "chunk_start": 1, "chunk_summary": "No actionable findings; this migration only performs closed-world schema checks and does not use string patterns for open-world semantic judgment.", "duration_ms": 14632, "findings": [], "path": "backend/alembic/versions/007_d6_entity_canon_metadata_json.py", "scan_kind": "python", "sha256": "dd607b5a3858f9f017e6f5f9b32eb2c1e81a61faf1a53b7ba9dbcdf8ce30c307"}
{"chunk_end": 72, "chunk_start": 1, "chunk_summary": "No actionable findings; this chunk only adds columns and explanatory comments.", "duration_ms": 29785, "findings": [], "path": "backend/alembic/versions/002_phase5_image_asset_variants.py", "scan_kind": "python", "sha256": "e4e87dc3c8f6b93e5f1dd0ea46beeece8ebe6c77003fb5157d59d60b5ad47783"}
{"chunk_end": 60, "chunk_start": 1, "chunk_summary": "Docstring lines 3-8 contain scenario-specific incident labels and pipeline names that should be generalized; the migration logic itself is otherwise fine.", "duration_ms": 44411, "findings": [{"category": "scenario_dependent_code", "evidence": "\"night_body_blood_curtain_red_circle\" / \"location_axis1_axis2_axis3_axisN\" / \"scene_image_pipeline\"", "line_end": 8, "line_start": 3, "recommended_fix": "Rewrite the docstring to a generic explanation of why `variant_label` needs 255 characters, and move the PID/example strings to the issue tracker or commit message outside the migration file.", "severity": "P2", "why_problematic": "This rationale hard-codes one incident’s label pattern and pipeline names into source text. That is scenario-specific pollution rather than generic schema rationale, and it can leak into downstream LLM/documentation context as an unintended exemplar set."}], "path": "backend/alembic/versions/004_variant_label_extend.py", "scan_kind": "python", "sha256": "4d8a88e7e7074489bb79c48cbb9cb62b23bf132db35add90d286cb0f93954b52"}
{"chunk_end": 212, "chunk_start": 1, "chunk_summary": "No actionable findings; this router only uses closed-world path/extension handling and contains no prompt or open-world semantic string judgments.", "duration_ms": 22579, "findings": [], "path": "backend/app/api/v1/exports.py", "scan_kind": "python", "sha256": "2020ad132c8eb0656b13fabe61e1981975228e27cd7c09e7ec805cd852f16059"}
{"chunk_end": 70, "chunk_start": 1, "chunk_summary": "No actionable findings; this router chunk does not use string-pattern semantic judgment or scenario-dependent prompt pollution.", "duration_ms": 5541, "findings": [], "path": "backend/app/api/v1/operations.py", "scan_kind": "python", "sha256": "0f7f23fe55215e99777c21681da6ba640ac71643716099be716f6c272a926691"}
{"chunk_end": 154, "chunk_start": 1, "chunk_summary": "No actionable findings in this chunk.", "duration_ms": 60035, "findings": [], "path": "backend/app/api/deps.py", "scan_kind": "python", "sha256": "483314672232f13f52c439ddde4113e4cc80d4d9c9ecedee6b0a6f9313f87cfb"}
{"chunk_end": 99, "chunk_start": 1, "chunk_summary": "No actionable findings; this chunk only performs closed-world file-path validation and migration bookkeeping.", "duration_ms": 77383, "findings": [], "path": "backend/alembic/versions/005_file_path_relative_check.py", "scan_kind": "python", "sha256": "63389a9d05e169558dc925ede63fceb3a0d990d8cc66b4ca85c640f9a2fe056e"}
{"chunk_end": 422, "chunk_start": 1, "chunk_summary": "No actionable findings; this chunk only uses closed-world status/operation constants and non-routing log/docstrings, with no open-world semantic string matching or scenario-dependent prompt pollution.", "duration_ms": 62310, "findings": [], "path": "backend/app/api/v1/episodes.py", "scan_kind": "python", "sha256": "63af6d843b5f1747846e3e921cf2787ed67a29846924d55773b26004c2473c65"}
{"chunk_end": 236, "chunk_start": 1, "chunk_summary": "No actionable findings; this chunk is CRUD-only and does not contain open-world string-pattern semantic judgments.", "duration_ms": 15101, "findings": [], "path": "backend/app/api/v1/prompts.py", "scan_kind": "python", "sha256": "122494141f3e88c71eda968f511c7e8d06566fc93f08738a132bae1706040aa4"}
{"chunk_end": 84, "chunk_start": 1, "chunk_summary": "No actionable findings; this file only wires FastAPI user routes to service methods and does not perform semantic string judgment or scenario-dependent prompt logic.", "duration_ms": 5118, "findings": [], "path": "backend/app/api/v1/users.py", "scan_kind": "python", "sha256": "6976912b349a4e82df9ee64157335cca7cb55bc0178040ecc9669a78001d75d6"}
{"chunk_end": 570, "chunk_start": 1, "chunk_summary": "1 actionable finding: provider labels are guessed from model-name prefixes instead of structured model metadata.", "duration_ms": 108680, "findings": [{"category": "semantic_string_judgment", "evidence": "provider = \"openai\" if current_model.startswith(\"gpt\") else \"gemini\"", "line_end": 529, "line_start": 523, "recommended_fix": "Build an alias->provider lookup from AVAILABLE_MODELS (or have _resolve_model return provider together with the model alias) and use that structured value here instead of a startswith() heuristic.", "severity": "P2", "why_problematic": "This infers an LLM provider from a name prefix rather than from declared metadata, so the UI will silently mislabel any alias that does not follow the current naming convention and the response will drift from the actual model config."}], "path": "backend/app/api/v1/projects.py", "scan_kind": "python", "sha256": "3aa863e13caa986487d44188e1d22e782e543cb053da74423a706f1c39a18079"}
{"chunk_end": 298, "chunk_start": 1, "chunk_summary": "No actionable findings in this chunk.", "duration_ms": 96648, "findings": [], "path": "backend/app/core/applicability.py", "scan_kind": "python", "sha256": "8cc4c734e80e9b97e773f8c16db66ae397ee7d17a4477f8f86c75e34f41eb45e"}
{"chunk_end": 1016, "chunk_start": 1, "chunk_summary": "Two actionable issues: a prompt-text substring lookup contract and an unvalidated variant enum.", "duration_ms": 137605, "findings": [{"category": "semantic_string_judgment", "evidence": "\"prompt_used contains 'outlook_id:{outlook_id}'\"", "line_end": 109, "line_start": 105, "recommended_fix": "Persist the outlook linkage in structured metadata (e.g. a dedicated `outlook_id` column or relation table) and query that field directly. Keep any text-marker parsing only as a temporary migration/backfill path, not the primary lookup.", "severity": "P1", "why_problematic": "Composite-image membership is being inferred from a magic substring inside generated prompt text, so open-world meaning depends on scenario prose rather than structured metadata. That is brittle and can silently break when prompt templates or wording change."}, {"category": "schema_or_enum_drift", "evidence": "variant: str  # \"original\" | \"A\" | \"B\"", "line_end": 45, "line_start": 44, "recommended_fix": "Replace the field with `Literal[\"original\", \"A\", \"B\"]` or an `Enum`, and update the service signature/docs to match the enforced contract.", "severity": "P2", "why_problematic": "The request model documents a closed set of allowed values but does not enforce it, so invalid strings can reach the service layer and fail late or behave inconsistently."}], "path": "backend/app/api/v1/images.py", "scan_kind": "python", "sha256": "5b74e60458b29711ced0f37357bf4401cf87a5b19b1b52a81fe798e58a34f6b9"}
{"chunk_end": 268, "chunk_start": 1, "chunk_summary": "One docstring contains a dated, scenario-specific operational note; otherwise no actionable findings.", "duration_ms": 143201, "findings": [{"category": "scenario_dependent_code", "evidence": "2026-05-11 사용자 권장안 #2 축소형: scene_detail step 의 atomic-fail 대비.", "line_end": 195, "line_start": 186, "recommended_fix": "Replace the docstring with a neutral one-line purpose statement and move the historical rationale to a changelog, ticket, or commit note.", "severity": "P2", "why_problematic": "The endpoint docstring embeds a one-off incident/history note instead of a stable API description. This is scenario-dependent prose that can drift and leak into generated docs or downstream model context."}], "path": "backend/app/api/v1/steps.py", "scan_kind": "python", "sha256": "475e38d587f4bc934cf5ebd1edd7a6d457e47f1204a977998d8ccfb592964670"}
{"chunk_end": 82, "chunk_start": 1, "chunk_summary": "No actionable findings in this chunk.", "duration_ms": 70128, "findings": [], "path": "backend/app/core/checkpoint_io.py", "scan_kind": "python", "sha256": "9921f13f1dd3fafa94bdb8d7ed2f6a0642f772854ae7dcc168958e3df2a4ddea"}
{"chunk_end": 213, "chunk_start": 1, "chunk_summary": "No actionable findings; this chunk only defines closed-world enums and regex validators.", "duration_ms": 96199, "findings": [], "path": "backend/app/core/bg_state_vocab.py", "scan_kind": "python", "sha256": "d6c2d1fccafa20ba72e40e90ca8a76f14d4225eebe57c37a205c010e50e9aafd"}
{"chunk_end": 8, "chunk_start": 1, "chunk_summary": "No actionable findings; this chunk only contains a package docstring and a closed-world export list.", "duration_ms": 2668, "findings": [], "path": "backend/app/core/dto/__init__.py", "scan_kind": "python", "sha256": "08aff79b29f622f9a3b9d2722fd95e43d14ee3c512de50b4cd6fee5912c05d64"}
{"chunk_end": 264, "chunk_start": 1, "chunk_summary": "No actionable findings in this chunk.", "duration_ms": 53400, "findings": [], "path": "backend/app/core/database.py", "scan_kind": "python", "sha256": "3255c3bb7dbac684feae1f4fa2b966b762df9cb41a1ef2a41d02ae7e68c73abe"}
{"chunk_end": 330, "chunk_start": 1, "chunk_summary": "Two prompt-pollution issues: one validator message inlines a closed allowed-token list, and another serializes the full raw intent object into retry-visible text.", "duration_ms": 179052, "findings": [{"category": "llm_closed_list_instruction", "evidence": "space_key hint {hint!r} not in allowed_space_keys {sorted(allowed)} ... controlled vocab strict — fail-fast.", "line_end": 68, "line_start": 66, "recommended_fix": "Raise a stable error code and attach `allowed_space_keys` as structured metadata for the caller; keep the retry-facing message generic and free of allowed-token lists.", "severity": "P2", "why_problematic": "This turns a structured enum check into free-text guidance that can be surfaced to the LLM retry path, scattering the allowed vocabulary into prompt text instead of keeping it in structured context."}, {"category": "scenario_dependent_code", "evidence": "f\"intent missing required fields {missing}: {intent}\"", "line_end": 184, "line_start": 182, "recommended_fix": "Include only machine-readable field names and stable IDs in the exception text, and send the full intent to non-LLM structured logs or telemetry.", "severity": "P2", "why_problematic": "Serializing the raw intent dict into validator text can leak free-form scenario content (for example metadata fields such as `sub_location_label` / `state_label_raw`) into LLM-visible retry prompts."}], "path": "backend/app/core/bg_catalog.py", "scan_kind": "python", "sha256": "2225fe163ccb224e71c66211be106160a1aaef1c39fac8d4e2cf93eacbe4a480"}
{"chunk_end": 249, "chunk_start": 1, "chunk_summary": "1 actionable schema drift finding: one backend selector is documented as closed-world but remains a free string; the rest of the chunk is closed-world config validation.", "duration_ms": 135726, "findings": [{"category": "schema_or_enum_drift", "evidence": "scene_detail_llm: str = \"gpt\"  # 씬 상세 분석 LLM: \"gemini\" or \"gpt\"", "line_end": 90, "line_start": 90, "recommended_fix": "Make the field a Literal[\"gemini\", \"gpt\"] or an Enum (with an explicit alias map if you need one) so invalid settings fail fast during config load.", "severity": "P2", "why_problematic": "The comment documents a finite set of valid values, but the setting is an unconstrained string, so typos or unsupported values can slip through and later force ad hoc string branching."}], "path": "backend/app/core/config.py", "scan_kind": "python", "sha256": "08605a2d225b7e37cf36b1211e63abc722c2a5b55a6dd1f3afb3e3896fc7f034"}
{"chunk_end": 191, "chunk_start": 1, "chunk_summary": "No actionable findings; this chunk uses structured manifest fields, closed-world ID handling, and exact enum checks only.", "duration_ms": 89683, "findings": [], "path": "backend/app/core/entity_protection.py", "scan_kind": "python", "sha256": "e181a4a35b8014686e298d8c09967a00e3be7867fd2c0fd005f3836515faaaa6"}
{"chunk_end": 218, "chunk_start": 1, "chunk_summary": "No actionable findings; the file only contains closed-world reserved-key validation and structured error serialization.", "duration_ms": 85722, "findings": [], "path": "backend/app/core/errors.py", "scan_kind": "python", "sha256": "ba89fa3d25ef305a82f0bc1b664d5f5b01a78c04656a81067792cb165042f7b2"}
{"chunk_end": 137, "chunk_start": 1, "chunk_summary": "Two actionable semantic-string issues found: ownership validation still consumes canonical names, and multi-background resolution collapses to an alpha-sorted bg_id.", "duration_ms": 173576, "findings": [{"category": "semantic_string_judgment", "evidence": "chain_bg_owned_by_shot; \"owned object names (English canonical, sorted)\"; \"post-parse owned judge / sentinel 검증에 소비\"", "line_end": 110, "line_start": 105, "recommended_fix": "Store stable object refs/IDs plus provenance in a structured type, and keep the display-name list only for prompt/rendering; make the judge compare IDs or refs, not names.", "severity": "P2", "why_problematic": "This contract feeds open-world ownership into a downstream judge as human-readable name strings, so validation can change when aliases, localization, or naming drift changes the text instead of the underlying object identity."}, {"category": "semantic_string_judgment", "evidence": "chain_bg_id_by_shot; \"first ok bid (deterministic alpha-sorted) only returned\"", "line_end": 119, "line_start": 112, "recommended_fix": "Represent multi-match cases as a candidate list plus explicit priority/selection metadata, and resolve with a structured rule instead of alpha-sorted fallback.", "severity": "P1", "why_problematic": "When several backgrounds match a shot, selecting the winner by alphabetical bg_id order makes semantic attachment depend on lexicographic string order and silently drops valid candidates."}], "path": "backend/app/core/dto/scene_analysis.py", "scan_kind": "python", "sha256": "5d17c2fdffaa87103bf4484442ea4485b1da846eb1a343445430c5fd3cd8ce84"}
{"chunk_end": 63, "chunk_start": 1, "chunk_summary": "No actionable findings; this chunk is closed-world enum validation for framing_scale with only soft diagnostics.", "duration_ms": 31038, "findings": [], "path": "backend/app/core/framing_scale.py", "scan_kind": "python", "sha256": "a48840fd0a5c37397cb26a10a57d4bd4333071aa55ad4bae0f6db74f786ca4e3"}
{"chunk_end": 175, "chunk_start": 1, "chunk_summary": "No actionable findings; this chunk only contains path-normalization helpers and legacy-compatibility comments.", "duration_ms": 82407, "findings": [], "path": "backend/app/core/file_paths.py", "scan_kind": "python", "sha256": "39921edd78a9cb4a0ade74a2728088eff7b2d104f6ee7f95c19b8da0a8992679"}
{"chunk_end": 205, "chunk_start": 1, "chunk_summary": "No actionable findings; this chunk only enforces closed-world entity metadata shape and emits diagnostic error messages.", "duration_ms": 219504, "findings": [], "path": "backend/app/core/entity_metadata.py", "scan_kind": "python", "sha256": "893d93150185304ee8d6b294842e6b970f8a7590185eb9681df3095795d0be49"}
{"chunk_end": 42, "chunk_start": 1, "chunk_summary": "No actionable findings; this chunk only performs thread lifecycle management and logging without semantic string judgment.", "duration_ms": 11222, "findings": [], "path": "backend/app/core/job_manager.py", "scan_kind": "python", "sha256": "fee1d9730d64a20295f3dbc70474835423be6fa6632498e616d282a36a410696"}
{"chunk_end": 35, "chunk_start": 1, "chunk_summary": "No actionable findings; this chunk is plain structured logging config without semantic string judgment or scenario-dependent prompt pollution.", "duration_ms": 4997, "findings": [], "path": "backend/app/core/logging_config.py", "scan_kind": "python", "sha256": "b6e8ab379da48f01ee8f87b1878a8d25a6bb38859dce8238521a89ee7f1a4d4c"}
{"chunk_end": 62, "chunk_start": 1, "chunk_summary": "No actionable findings; this file only handles closed-world file I/O for skip ID JSON and does not perform semantic string judgment or scenario-dependent prompting.", "duration_ms": 8729, "findings": [], "path": "backend/app/core/low_freq_skip.py", "scan_kind": "python", "sha256": "5a119ae11d83690059124ebe4ab43860481a2bac301864835aaf9d9e972f9f9d"}
{"chunk_end": 85, "chunk_start": 1, "chunk_summary": "No actionable findings in this chunk.", "duration_ms": 16613, "findings": [], "path": "backend/app/core/pipeline_cache.py", "scan_kind": "python", "sha256": "0efc4fe218ff2295389ecfb153eede4bfa30670817c8b6059307c716350ee318"}
{"chunk_end": 363, "chunk_start": 1, "chunk_summary": "One low-risk finding: `phrase_diagnostic` still infers prompt coverage from lowercase substring checks against closed phrase lists.", "duration_ms": 139392, "findings": [{"category": "semantic_string_judgment", "evidence": "`phrase_diagnostic()`: `prompt_lower = (t2i_prompt or \"\").lower()`, `label in prompt_lower`, `any(v in prompt_lower for v in zone_variants)`, `any(v in prompt_lower for v in depth_variants)`", "line_end": 331, "line_start": 315, "recommended_fix": "Replace the text scan with structured machine-readable metadata from the renderer/prompt-card (or remove it entirely); if kept, fence it to debug-only telemetry and never use it for validation or routing.", "severity": "P2", "why_problematic": "This is an open-world semantic check over free-form prompt text implemented with lowercase substring/phrase-list matching, which is brittle and can false-match unrelated prose."}], "path": "backend/app/core/frame_spatial_contract.py", "scan_kind": "python", "sha256": "dc8cd3121d7ca06ba098541ab7a40a722f0c79326fa6c4f65ed6b0ded2483f80"}
{"chunk_end": 46, "chunk_start": 1, "chunk_summary": "Scenario-specific workflow details are hardcoded in the `CompletionReport` docstring.", "duration_ms": 117980, "findings": [{"category": "scenario_dependent_code", "evidence": "Block B B1, plan v2.1.3 / spec V5 §2.1; B7/B8; DB row / PNG / cp 파일 부재", "line_end": 25, "line_start": 17, "recommended_fix": "Rewrite the docstring to keep only generic field semantics, and move the origin-to-meaning mappings/examples into a structured enum or external SOT keyed by stable IDs.", "severity": "P2", "why_problematic": "This docstring bakes one pipeline's block IDs, plan/spec references, and artifact examples into a reusable integrity-report schema, so downstream tooling that reuses the text inherits scenario-specific jargon instead of a structured source of truth."}], "path": "backend/app/core/integrity_report.py", "scan_kind": "python", "sha256": "746d4ecf33b66cd7e4c2b3d94f850ba980b2ca941cf49393fbaeef54dc4bce79"}
{"chunk_end": 31, "chunk_start": 1, "chunk_summary": "No actionable findings; this chunk only contains password hashing, session token generation, and datetime-based expiry checks without open-world string judgments or scenario-dependent prompt pollution.", "duration_ms": 3527, "findings": [], "path": "backend/app/core/security.py", "scan_kind": "python", "sha256": "efce0d75e3a2e4ccdc8c675afcc00854a5af7fbe5799da350b7caf2bf30957ad"}
{"chunk_end": 130, "chunk_start": 1, "chunk_summary": "One shared validator error path hard-codes shot_dependency_t2i-specific rerun guidance; the rest is closed-world shape validation.", "duration_ms": 121141, "findings": [{"category": "scenario_dependent_code", "evidence": "`shot_dependency_t2i schema_version ... cp 를 force re-run 하세요.`", "line_end": 124, "line_start": 120, "recommended_fix": "Replace the message with a generic legacy-shape diagnostic for keep_elements, and let the caller attach any producer/version/rerun hint from its own metadata or error wrapper.", "severity": "P2", "why_problematic": "This shared validator emits pipeline-specific names and remediation text in an exception message, so the output is tied to one scenario instead of the generic keep_elements contract and can become misleading when reused by other importers."}], "path": "backend/app/core/keep_elements.py", "scan_kind": "python", "sha256": "1fa17ab361a1f74991d78c8965aa818897f82f6e9cbd652423308c3f853293fe"}
{"chunk_end": 83, "chunk_start": 1, "chunk_summary": "No actionable findings in this chunk.", "duration_ms": 34870, "findings": [], "path": "backend/app/core/settings_registry.py", "scan_kind": "python", "sha256": "a75c16fd433e6ca3dbba88f5a6ed532714c3670dc287f6992b8cef98fc7d62d7"}
{"chunk_end": 166, "chunk_start": 1, "chunk_summary": "One actionable issue: the legacy episode fallback can load another episode's planning-doc context into the current prompt.", "duration_ms": 173995, "findings": [{"category": "scenario_dependent_code", "evidence": "for ep_dir in sorted(base.iterdir()): ... legacy_cp = alt; break", "line_end": 93, "line_start": 88, "recommended_fix": "Do not scan arbitrary episode directories here. Either return empty context when the exact `episode_id` path is absent, or load only a deterministic legacy location and verify the manifest's embedded episode id/slug matches `episode_id` before accepting it.", "severity": "P1", "why_problematic": "If the exact `episode_id` manifest is missing, this fallback scans every episode directory and accepts the first manifest it finds, so the current episode can inherit another episode's planning-doc text. That pollutes downstream prompts with the wrong scenario and makes behavior depend on filesystem order rather than the requested episode."}], "path": "backend/app/core/planning_doc_context.py", "scan_kind": "python", "sha256": "96a5e66e6ef2bd8f932a78d98da7cf04cc0bb01723a449a16a4d79929afbe7ea"}
{"chunk_end": 225, "chunk_start": 1, "chunk_summary": "No actionable findings; this chunk uses closed-world step IDs/enums and has no open-world string-pattern routing or scenario-polluting prompt logic.", "duration_ms": 116914, "findings": [], "path": "backend/app/core/step_catalog.py", "scan_kind": "python", "sha256": "a9349acfefc4619ad2fbd259fab29f7803b2677237d2b146736d9687bef7ef3d"}
{"chunk_end": 125, "chunk_start": 1, "chunk_summary": "No actionable findings; this chunk is a step registry with closed-world step IDs and optional imports only.", "duration_ms": 11339, "findings": [], "path": "backend/app/core/steps/__init__.py", "scan_kind": "python", "sha256": "7dfe3f4a7e6f7f2fb6c660abc77c47b48790b723847c664652cecacc26c88b42"}
{"chunk_end": 129, "chunk_start": 1, "chunk_summary": "One actionable schema drift issue: the fresh-evidence validator does not enforce the declared confidence enum, so unexpected confidence values can slip through when evidence is present.", "duration_ms": 165754, "findings": [{"category": "schema_or_enum_drift", "evidence": "VALID_LLM_CONFIDENCE = (\"high\", \"medium\", \"low\") ... if not sf and confidence != \"low\" ... if not (sf or vi or cd) and confidence != \"low\"", "line_end": 128, "line_start": 108, "recommended_fix": "Add an upfront `confidence in VALID_LLM_CONFIDENCE` check (after the legacy rejection) and reject unknown values with `step.contract_violation`; then keep the low-confidence evidence rule as a separate branch using the validated enum/constant.", "severity": "P1", "why_problematic": "`assert_fresh_llm_evidence()` only special-cases `legacy` and the literal `\"low\"`; any other unexpected `confidence` token (including `None`) can still pass whenever at least one evidence list is non-empty, so the declared confidence enum is not actually enforced."}], "path": "backend/app/core/steps/_evidence_helpers.py", "scan_kind": "python", "sha256": "1b11a4a10b12e90373ed92fd688d5eb5db3a49220fc3c31d94987b7fdf77b9aa"}
{"chunk_end": 161, "chunk_start": 1, "chunk_summary": "Regex-based bracket stripping and bare-key fallback are using surface-form heuristics to resolve entity names, which is the actionable semantic debt in this chunk.", "duration_ms": 535475, "findings": [{"category": "semantic_string_judgment", "evidence": "`_BRACKET_PATTERN.sub`, `normalize_name`, `if not _BRACKET_PATTERN.search(raw)`, `bare = normalize_name(query)`, `lookup_name`", "line_end": 160, "line_start": 51, "recommended_fix": "Move alias/canonical relationships into structured data (for example `aliases[]`, `variant_of_id`, or a dedicated alias table in the SOT), build exact alias indices from that source, and keep runtime normalization limited to harmless whitespace cleanup only.", "severity": "P1", "why_problematic": "The module treats any bracketed suffix as disposable and then resolves queries through the stripped \"bare\" form. That is a string-pattern heuristic over open-world entity names, so distinct canonical records can collapse to the same key and be silently misassociated."}], "path": "backend/app/core/name_matcher.py", "scan_kind": "python", "sha256": "6a06d3cde26de3c8fec7c158b189b543c3b86df0d99b9e9cdda903b03f62f797"}
{"chunk_end": 116, "chunk_start": 1, "chunk_summary": "No actionable findings in this chunk.", "duration_ms": 26400, "findings": [], "path": "backend/app/core/steps/_owned_judge.py", "scan_kind": "python", "sha256": "6280d7a602b422b4d3a0815310a128f83cc6eacf7b0ab3ba4944eb703268fdcb"}
{"chunk_end": 1597, "chunk_start": 1, "chunk_summary": "Found one actionable string-based routing dependency in legacy contract-drift handling.", "duration_ms": 351746, "findings": [{"category": "semantic_string_judgment", "evidence": "if \"schema_version mismatch\" in mismatch_reason: ... self.step_id in _LEGACY_SCHEMA_BUMP_ALLOWLIST", "line_end": 769, "line_start": 752, "recommended_fix": "Return a structured mismatch result from `_check_cp_mismatch()` (for example `kind='schema_version'|'config_hash'` plus details/message) and switch on `kind` here; keep the free-text reason only for logging and user-facing errors.", "severity": "P1", "why_problematic": "This branches on human-readable mismatch text from `_check_cp_mismatch()` instead of a structured mismatch kind, so a wording/copy/localization change can silently flip a valid legacy schema-bump from rerun to block and couple resume routing to message phrasing."}], "path": "backend/app/core/step_runner.py", "scan_kind": "python", "sha256": "1bc98850586701b74aec05907a3329fb336d1b8e4412efc5724e836aeafc121a"}
{"chunk_end": 457, "chunk_start": 1, "chunk_summary": "Two actionable findings: prompt semantics are still inferred from hard-coded phrase lists/regexes, and the resulting text-derived labels are used to hard-fail validation.", "duration_ms": 544136, "findings": [{"category": "semantic_string_judgment", "evidence": "_CHARACTER_TOKENS, _BACKGROUND_TOKENS, _OBJECT_TOKENS, _FROM_THE_REFERENCE_RE, _is_plural_reference_images, _has_generic_instruction_signal, classify_from_the_reference", "line_end": 190, "line_start": 47, "recommended_fix": "Move reference kind/scope into structured metadata from the prompt-card SOT or upstream parser and have validation read that enum directly. If a heuristic must remain, keep it diagnostic-only and never use it to continue/skip or to derive production labels.", "severity": "P1", "why_problematic": "This block infers whether 'from the reference' means character/background/object by scanning prompt text with fixed token lists, regexes, and line-local substring checks, and it can silently skip the guard on heuristic 'generic' lines. That is open-world semantic meaning decided by surface strings, so wording drift can change classification or bypass validation."}, {"category": "semantic_string_judgment", "evidence": "classify_from_the_reference(prompt), ref_type == 'character', ref_type == 'background', ref_type == 'object', raise RefContractError", "line_end": 457, "line_start": 425, "recommended_fix": "Require an explicit structured from_reference/reference_kind field from the prompt source or SOT and make this branch consume that field. Treat classify_from_the_reference as warning/telemetry only if it is kept at all.", "severity": "P1", "why_problematic": "The phantom guard turns the text-derived ref_type into a hard validation decision. A misread prompt phrase can raise or suppress the error path even when attached_meta is correct, because the branch is keyed off string-classified labels rather than structured reference metadata."}], "path": "backend/app/core/ref_contract_validator.py", "scan_kind": "python", "sha256": "93a34075f568486f74bd384f96f18208bab54f8e5777726c6d7beef955a7b504"}
{"chunk_end": 214, "chunk_start": 1, "chunk_summary": "One scenario-dependent prompt-leakage path remains: floor-plan `prompt_text` is carried forward into background planning instead of a purely structured representation.", "duration_ms": 158999, "findings": [{"category": "scenario_dependent_prompt", "evidence": "`prompt_text`, `floor_plan_prompts`, `run_background_chain_planning(...)`", "line_end": 185, "line_start": 84, "recommended_fix": "Keep floor-plan facts in structured fields (for example location_ids, room/group ids, indoor/outdoor flags, adjacency) and generate any helper narrative from one canonical template at prompt-composition time; do not forward raw upstream prompt_text verbatim.", "severity": "P2", "why_problematic": "This loads free-form floor-plan prose from a prior checkpoint and forwards it as prompt context for a later LLM step, so the planner can become sensitive to scenario-specific wording rather than only stable structured floor-plan facts."}], "path": "backend/app/core/steps/background_chain_planning_step.py", "scan_kind": "python", "sha256": "0fed58a03fae448319826f84ba117fe61266792824f6f7500ea105d649ad5d39"}
{"chunk_end": 995, "chunk_start": 1, "chunk_summary": "One actionable finding: SceneDirectorStep still falls back from schema-validated short IDs to normalized name/prefix string matching, which can silently misbind entities.", "duration_ms": 174046, "findings": [{"category": "blind_string_mutation", "evidence": "`name_to_uuid[name_upper]` and `raw_id.replace(\"CHAR_\", \"\").replace(\"BG_\", \"\").replace(\"PROP_\", \"\")`", "line_end": 863, "line_start": 842, "recommended_fix": "Remove the runtime name/prefix fallback and accept only schema-validated short IDs (`short_to_uuid`). If alias recovery is genuinely needed, move it into an explicit offline alias table keyed by canonical IDs and fail closed on any unresolved `raw_id` instead of normalizing it in-place.", "severity": "P1", "why_problematic": "This block turns raw LLM output into canonical entity UUIDs by uppercasing, space-to-underscore normalization, and prefix stripping. That bypasses the short-ID enum contract and can silently resolve a hallucinated or formatted alias to the wrong entity, which is exactly the kind of open-world semantic mapping that should not be handled by blind string mutation."}], "path": "backend/app/core/steps/analysis_steps_legacy.py", "scan_kind": "python", "sha256": "5309486c802bcbd8c46d563320b1b652467837e59f9414993412e25513d9e3a3"}
{"chunk_end": 366, "chunk_start": 1, "chunk_summary": "1 actionable issue: source_language is silently blanked and left for the LLM to infer from scene text, which is scenario-dependent prompt behavior.", "duration_ms": 125529, "findings": [{"category": "scenario_dependent_prompt", "evidence": "\"LLM이 scene_segments에서 자체 추론하도록 빈 문자열로 inject\"; source_language = extract_source_language(rules_cp) or \"\"", "line_end": 250, "line_start": 247, "recommended_fix": "Require source_language in the visual_world_rules checkpoint (or supply a closed-world default from settings), validate it before building the prompt, and fail/skip the step when it is absent instead of injecting an empty string.", "severity": "P1", "why_problematic": "When the structured source-language field is missing, this code suppresses it instead of validating/failing closed, so prompt generation can fall back to open-ended inference from scene text. That creates scenario-dependent behavior and hides an upstream schema gap."}], "path": "backend/app/core/steps/background_prompt_step.py", "scan_kind": "python", "sha256": "9318a00456c3f8bb0eaf22772efac3dc477b8daba55892adea74588c6be52359"}
{"chunk_end": 629, "chunk_start": 1, "chunk_summary": "No actionable findings; the chunk uses closed-world status/ID checks and structural DB/file validation only.", "duration_ms": 315957, "findings": [], "path": "backend/app/core/steps/background_chain_render_step.py", "scan_kind": "python", "sha256": "db955002c4caf6db1a7f2927358a8e72e4d761f4dfa4bec668d8a4c36376f52f"}
{"chunk_end": 766, "chunk_start": 1, "chunk_summary": "No actionable findings; this chunk uses closed-world ID/status/path checks and data pass-through without open-world semantic string judgment or scenario-specific prompt pollution.", "duration_ms": 61556, "findings": [], "path": "backend/app/core/steps/background_render_step.py", "scan_kind": "python", "sha256": "61f88d19935e0b1efe05d6cdf48d004f8901bd7705a8577ad4c8f9516bf5d7b9"}
{"chunk_end": 264, "chunk_start": 1, "chunk_summary": "No actionable findings; this chunk mainly loads structured checkpoints, counts shots by ID, and forwards raw location context to the LLM without string-pattern semantic routing.", "duration_ms": 322916, "findings": [], "path": "backend/app/core/steps/background_classify_step.py", "scan_kind": "python", "sha256": "1ffb71738766db06c062a8702c057a161718c17d4b721d15138ca50a742f33eb"}
{"chunk_end": 107, "chunk_start": 1, "chunk_summary": "One actionable finding: raw director notes from a previous checkpoint are appended to the character-extraction prompt, creating scenario-dependent prompt pollution.", "duration_ms": 110117, "findings": [{"category": "scenario_dependent_prompt", "evidence": "\"director_notes\", \"[시각적 규칙]\", `user_prompt = f\"[시각적 규칙]\\n{visual_rules}\\n\\n{user_prompt}\"`", "line_end": 71, "line_start": 62, "recommended_fix": "Do not feed raw `director_notes` into this extraction step. Pass only schema-backed world-rule fields (for example normalized aliases/IDs or a compact structured rule summary), or move the notes to a separate, explicitly structured channel and keep the character-list prompt focused on scene text only.", "severity": "P1", "why_problematic": "This step is meant to extract characters from the scene block, but it also injects free-form notes from `visual_world_rules` into the same prompt. That lets the LLM base character extraction on unrelated scenario-specific prose, names, or props instead of a structured SOT, and the resulting list is later reused to constrain names downstream."}], "path": "backend/app/core/steps/character_list_step.py", "scan_kind": "python", "sha256": "ee4727eab61e855b22a5898064a08f163d1b34b4b45a0d79472a6489ae8b5254"}
{"chunk_end": 95, "chunk_start": 1, "chunk_summary": "One actionable finding: `_build_beat_shot_context` flattens raw beat/shot descriptions and character names into an unstructured string that is later used as relation-extraction context.", "duration_ms": 111114, "findings": [{"category": "scenario_dependent_prompt", "evidence": "`[씬 {si}]`, `Beat {b.get('beat_index', '?')}: [{b.get('change_type', '')}] {b.get('description', '')}`, `chars = ', '.join(sh.get('characters', []))`, `Shot {sh.get('shot_index', '?')}: [{chars}] {sh.get('description', '')}`", "line_end": 92, "line_start": 85, "recommended_fix": "Pass structured scene/beat/shot records to `extract_entity_relations` (IDs plus relation-relevant fields) and render any prompt text from stable canonical fields only; avoid concatenating raw descriptions or names here.", "severity": "P2", "why_problematic": "This serializes scene/shot checkpoint text and character names into a freeform context blob, so the downstream extractor can make relation judgments from scenario-specific prose instead of a structured source-of-truth."}], "path": "backend/app/core/steps/entity_relation_step.py", "scan_kind": "python", "sha256": "b1a46b315c5130b07a2b6a7569645a2432bfda4bcba2a1d7a27a110c321ebe71"}
{"chunk_end": 369, "chunk_start": 1, "chunk_summary": "No actionable findings; this chunk mostly handles structured checkpoint loading, exact ID/schema handling, and prompt assembly.", "duration_ms": 481473, "findings": [], "path": "backend/app/core/steps/background_planner_step.py", "scan_kind": "python", "sha256": "53d33dcd02ffd8685de2d501c41a75e5556ef64a21aa374b9bfa3ed61ecdbfb8"}
{"chunk_end": 519, "chunk_start": 1, "chunk_summary": "No actionable findings; this chunk only uses closed-world enums/IDs and structured post-processing.", "duration_ms": 532986, "findings": [], "path": "backend/app/core/steps/background_master_plan_step.py", "scan_kind": "python", "sha256": "f316adda85aac755b904d1cc15c8d3ff7ed749c8c9584a6289f4d9cd457d646c"}
{"chunk_end": 304, "chunk_start": 1, "chunk_summary": "No actionable findings; this chunk only uses closed-world status/ID validation and structural prompt orchestration, with no open-world string-pattern judgment in this file.", "duration_ms": 77409, "findings": [], "path": "backend/app/core/steps/floor_plan_prompt_step.py", "scan_kind": "python", "sha256": "3d0f71b1f1d384eeb6b46210236a9910106809dec6aed137cf941ee4f45ab8be"}
{"chunk_end": 537, "chunk_start": 1, "chunk_summary": "No actionable findings in this chunk; the file uses closed-world ID/enum checks and does not judge open-world semantics by string patterns.", "duration_ms": 90707, "findings": [], "path": "backend/app/core/steps/floor_plan_render_step.py", "scan_kind": "python", "sha256": "95bc28000e5f70e5ac3f3830e7679e68d6c937b7ce401cc96a21bafcc17a6b06"}
{"chunk_end": 781, "chunk_start": 1, "chunk_summary": "A heuristic first-sentence/first-line summary of prior floor-plan prompt text is injected into later prompts, which can misrepresent the context and is the only actionable semantic-string debt here.", "duration_ms": 150727, "findings": [{"category": "blind_string_mutation", "evidence": "text.find(sep); if 0 < idx < 200: return text[:idx + 1].strip(); first_line = text.split(\"\\n\", 1)[0].strip()", "line_end": 734, "line_start": 724, "recommended_fix": "Persist an explicit structured summary/layout_digest field from the upstream generation step (or run a dedicated summarizer with structured output) and feed that field forward; do not derive downstream context by slicing the first sentence/line of prompt_text.", "severity": "P1", "why_problematic": "This compresses arbitrary LLM output into a \"summary\" using punctuation boundaries and a fixed length cutoff, so the downstream prompt context can silently drop or distort the actual spatial/layout meaning of the earlier floor plan prompt."}], "path": "backend/app/core/steps/location_floor_plan_step.py", "scan_kind": "python", "sha256": "86bc164f3624c24cfe7edc1704db823d4f3d9122c7f53fd39106b85d11dcb146"}
{"chunk_end": 553, "chunk_start": 1, "chunk_summary": "No actionable findings in this chunk.", "duration_ms": 150577, "findings": [], "path": "backend/app/core/steps/outlook_steps.py", "scan_kind": "python", "sha256": "7cff875a4cc7dd2b8d2a3bc0961d812f257d97eba72e149a570f06b4e35548d2"}
{"chunk_end": 242, "chunk_start": 1, "chunk_summary": "No actionable findings; this chunk only loads structured prompts/schemas and uses closed-world checkpoint/schema validation.", "duration_ms": 58577, "findings": [], "path": "backend/app/core/steps/scene_camera_flow_step.py", "scan_kind": "python", "sha256": "657e39ab6fa5b6c7e0a3f31c562ee529d72342fdf1be7e23ac81d058dbcef351"}
{"chunk_end": 217, "chunk_start": 1, "chunk_summary": "This chunk is clean; no actionable findings.", "duration_ms": 269988, "findings": [], "path": "backend/app/core/steps/planning_doc_step.py", "scan_kind": "python", "sha256": "30736eedb414dddb5c432e9cde26f18ac1a90b52404ba4abdd013f8df3485131"}
{"chunk_end": 889, "chunk_start": 1, "chunk_summary": "One actionable semantic-string judgment remains: location visuals are gated by the free-text '실패' prefix instead of a structured status field.", "duration_ms": 160472, "findings": [{"category": "semantic_string_judgment", "evidence": "if summary.startswith(\"실패\"):", "line_end": 268, "line_start": 266, "recommended_fix": "Add a structured `analysis_status`/`is_failed` field to the `location_consistency` checkpoint, backfill legacy rows once, and branch on that enum or boolean; keep `analysis_summary` descriptive only.", "severity": "P1", "why_problematic": "This branches on localized free text in `analysis_summary`, so wording drift or unrelated summaries beginning with that token will silently change whether `fixed_visual_description` is injected downstream."}], "path": "backend/app/core/steps/scene_context_loader.py", "scan_kind": "python", "sha256": "1cd6a7d7b27f853057101287ef0c718d2337d8e7e89a3f65096e6395b01ed410"}
{"chunk_end": 162, "chunk_start": 1, "chunk_summary": "No actionable findings; this chunk uses checkpoint IDs, DB-backed shot-type data, and generic diversity guidance without string-based semantic routing or scenario-specific prompt pollution.", "duration_ms": 70295, "findings": [], "path": "backend/app/core/steps/shot_cinematography_step.py", "scan_kind": "python", "sha256": "553ac84f1a94ced477d79e1c148187e782717b8332a723c27413affb05b1d128"}
{"chunk_end": 126, "chunk_start": 1, "chunk_summary": "No actionable findings; this chunk only performs checkpoint loading, closed-world presence checks, and version stamping.", "duration_ms": 34460, "findings": [], "path": "backend/app/core/steps/shot_director_step.py", "scan_kind": "python", "sha256": "9a337c33f08a7b50b1b2344fde61ef56d39e90922639effa8e06e34012114abe"}
{"chunk_end": 302, "chunk_start": 1, "chunk_summary": "One actionable finding: the LLM prompt injects scenario-specific location and entity names into an open-world dependency decision, which can bias results away from the shot/T2I evidence.", "duration_ms": 87034, "findings": [{"category": "scenario_dependent_prompt", "evidence": "`loc_name = location_names.get(loc_id, loc_id)` and `\"Entities: {', '.join(ent_names_list)}\"` inside `user_prompt`", "line_end": 265, "line_start": 246, "recommended_fix": "Remove human-readable location/entity names from the prompt, or pass them only as separate structured metadata with stable IDs. Keep the natural-language prompt limited to the shot description and T2I text unless those extra fields are explicitly required by the model contract.", "severity": "P1", "why_problematic": "This step is supposed to judge shot dependency from shot descriptions/T2I, but the prompt adds free-form location and entity names from upstream story data. Those scenario-specific labels can become a lexical shortcut for the LLM and bias continuity judgments toward name overlap instead of the actual visual/story evidence."}], "path": "backend/app/core/steps/shot_dependency_t2i_step.py", "scan_kind": "python", "sha256": "4e5bf81c1d9040583870f1e8f14636e453dad73093f557bbbaf9b1511ec9cc8a"}
{"chunk_end": 309, "chunk_start": 1, "chunk_summary": "No actionable findings; this chunk uses closed-world IDs and status enums only, with no pattern-based semantic judgments or scenario-dependent prompt pollution.", "duration_ms": 50634, "findings": [], "path": "backend/app/core/steps/shot_essence_extraction_step.py", "scan_kind": "python", "sha256": "bd3cf91c9b396438817a9061f82d5b2f1d5414d825f17ae6641c171361bf1820"}
{"chunk_end": 46, "chunk_start": 1, "chunk_summary": "No actionable findings; this chunk only loads checkpoints, validates required inputs, and forwards structured data to the staging pipeline.", "duration_ms": 7421, "findings": [], "path": "backend/app/core/steps/shot_staging_step.py", "scan_kind": "python", "sha256": "e86b746584f044e566326671dc8be66843ca343ea2d8b7e09356a77dcf07d935"}
{"chunk_end": 287, "chunk_start": 1, "chunk_summary": "No actionable findings; this chunk uses structured shot indices and generic status/fallback strings, with no open-world semantic string gating or scenario-dependent prompt pollution visible.", "duration_ms": 82650, "findings": [], "path": "backend/app/core/steps/shot_selection_step.py", "scan_kind": "python", "sha256": "0b3d7f87247c5760f000b0ed7d6031a2ed5722df3d4e2c087fc2ef055d16a254"}
{"chunk_end": 524, "chunk_start": 1, "chunk_summary": "No actionable findings; this chunk confines string handling to closed-world ID/status validation and normal LLM prompt assembly.", "duration_ms": 132279, "findings": [], "path": "backend/app/core/steps/shot_validator_step.py", "scan_kind": "python", "sha256": "4c28fc69a4a9e706c558d73f8bae3e411531373f1529fa9d2b553333ec3a1b86"}
{"chunk_end": 54, "chunk_start": 1, "chunk_summary": "No actionable findings; this chunk only performs file existence fallback and data loading without string-pattern semantic judgment or scenario-dependent prompt pollution.", "duration_ms": 5070, "findings": [], "path": "backend/app/core/steps/text_steps.py", "scan_kind": "python", "sha256": "dee0d749567714398fc791817e5732c769049f44dd42787e16593f185023fd08"}
{"chunk_end": 40, "chunk_start": 1, "chunk_summary": "No actionable findings; this chunk uses closed-world task/status bookkeeping only and does not perform open-world semantic string judgment or scenario-dependent prompt logic.", "duration_ms": 6858, "findings": [], "path": "backend/app/core/task_registry.py", "scan_kind": "python", "sha256": "ff6afe0ab1644dfaa038b42db7fadb121b972c02662ff83b685bfd4aa28ff348"}
{"chunk_end": 197, "chunk_start": 1, "chunk_summary": "SceneSummaryStep has a documented/output key mismatch (`scene_summaries` vs `summaries`); no other actionable findings in this chunk.", "duration_ms": 138099, "findings": [{"category": "schema_or_enum_drift", "evidence": "Docstring says `scene_summaries: [{scene_index, scene_summary}, ...]`, but the return uses `{\"data\": {\"summaries\": scene_summaries}}`.", "line_end": 196, "line_start": 140, "recommended_fix": "Use one canonical field name across the docstring, return payload, and any readers. If compatibility matters, emit both keys briefly and migrate consumers to the chosen name.", "severity": "P2", "why_problematic": "The step contract is internally inconsistent: readers following the documented field name will not find the emitted list under that key, which can silently break downstream checkpoint consumers or tests."}], "path": "backend/app/core/steps/summary_steps.py", "scan_kind": "python", "sha256": "68cadacda2b891546478fce5e24ebdc5117e38140cbe27bd27070780fcb4a949"}
{"chunk_end": 258, "chunk_start": 1, "chunk_summary": "No actionable findings; this chunk is limited to checkpoint I/O and hash-based integrity checks.", "duration_ms": 38404, "findings": [], "path": "backend/app/core/steps/t2i_review_step.py", "scan_kind": "python", "sha256": "7f44d1520245d3690182a66bd7e2d6892be7b59cc8939e0b4c85abeb75ae6baf"}
{"chunk_end": 20, "chunk_start": 1, "chunk_summary": "No actionable findings in this chunk.", "duration_ms": 4776, "findings": [], "path": "backend/app/i18n/loader.py", "scan_kind": "python", "sha256": "1df58ae74fe67baf31a49079fd1cbddf2dabf03081e7ba321cd802e024319af9"}
{"chunk_end": 156, "chunk_start": 1, "chunk_summary": "No actionable findings; this file is a version registry with closed-world metadata only.", "duration_ms": 14291, "findings": [], "path": "backend/app/core/version_registry.py", "scan_kind": "python", "sha256": "b4454265a347ae52a317b1277af5b517365c83f12e58fc952e6630c916bda0c1"}
{"chunk_end": 42, "chunk_start": 1, "chunk_summary": "No actionable findings; this file only persists opaque activity-log fields and does not perform open-world semantic string judgment.", "duration_ms": 3348, "findings": [], "path": "backend/app/logging/activity_logger.py", "scan_kind": "python", "sha256": "997d9bc02147260f7d5770a934833ee80b3a205db9e22867aa632dd587764d10"}
{"chunk_end": 16, "chunk_start": 1, "chunk_summary": "No actionable findings; this chunk only defines a logging table schema and does not perform semantic string judgments.", "duration_ms": 4717, "findings": [], "path": "backend/app/logging/models.py", "scan_kind": "python", "sha256": "5e479a106bc17dc09334a3b3871991ec6787e18bdf9d372fe805dc4c9a19d6c8"}
{"chunk_end": 51, "chunk_start": 1, "chunk_summary": "No actionable findings; this chunk only defines SQLAlchemy models and closed-world identifiers, with no open-world semantic string judgment or scenario-dependent prompt content.", "duration_ms": 7916, "findings": [], "path": "backend/app/models/catalog.py", "scan_kind": "python", "sha256": "9d4a647c9247795de8bc255e2c91bbe888ea3fb924797258c1032c869d7f5ac7"}
{"chunk_end": 130, "chunk_start": 1, "chunk_summary": "No actionable findings; this chunk uses closed-world environment checks, startup side effects, and fixed status constants without open-world string judgment.", "duration_ms": 16316, "findings": [], "path": "backend/app/main.py", "scan_kind": "python", "sha256": "49472efd68728a06703b1264499b0c8c0dc8dc57e7b020a1fec14e155b4587bc"}
{"chunk_end": 341, "chunk_start": 1, "chunk_summary": "No actionable findings; this chunk is schema-only and does not include open-world semantic judgment or scenario-polluting prompt text.", "duration_ms": 61555, "findings": [], "path": "backend/app/models/project.py", "scan_kind": "python", "sha256": "c31a4345928225a45b84f3d64938a3481dd19f3858014ef6700b16e2db71a1dd"}
{"chunk_end": 351, "chunk_start": 1, "chunk_summary": "No actionable findings in this chunk.", "duration_ms": 75878, "findings": [], "path": "backend/app/modules/gemini_i2i_editor.py", "scan_kind": "python", "sha256": "3ecd9be0d08929461578a029b148fa228612db99edd966f28d8d46d44b5d3a8a"}
{"chunk_end": 139, "chunk_start": 1, "chunk_summary": "No actionable findings; this module only records and aggregates closed-world generation trace metadata without open-world string judgment.", "duration_ms": 15081, "findings": [], "path": "backend/app/modules/generation_tracker.py", "scan_kind": "python", "sha256": "635516c04a47c0e75bbe854a51872482c73a7b9b55195c7d16f4df442b04c6fa"}
{"chunk_end": 214, "chunk_start": 1, "chunk_summary": "Hard-coded semantic relation labels drive dependency routing in one allowlist; one open-world string-debt hotspot was found.", "duration_ms": 141050, "findings": [{"category": "semantic_string_judgment", "evidence": "VISUAL_REFERENCE_RULES", "line_end": 26, "line_start": 18, "recommended_fix": "Move the relation-family dependency policy into the structured SOT or a generated enum-backed config, and have the graph builder consume that shared contract instead of hardcoding semantic label lists here.", "severity": "P1", "why_problematic": "This embeds a semantic allowlist (`identity`, `transformation`, `possession`) in code to decide when a visual dependency exists. That is open-world ontology routing via string labels, so any upstream relation-family rename, localization, or vocabulary expansion will silently change generation order unless this policy is sourced from the canonical schema/ontology."}], "path": "backend/app/modules/entity_dependency.py", "scan_kind": "python", "sha256": "d7d1dc91f1acdd43734d0097877c1e5646006e307759d8e080654f80375d5ab3"}
{"chunk_end": 1, "chunk_start": 1, "chunk_summary": "No actionable findings; this file contains only a neutral module docstring.", "duration_ms": 1935, "findings": [], "path": "backend/app/modules/llm/__init__.py", "scan_kind": "python", "sha256": "154c2f746a44ce90bdb76dee3da6fbcef2f58bd62941312a6fdb533cc3f4882d"}
{"chunk_end": 953, "chunk_start": 1, "chunk_summary": "Two actionable semantic-string heuristics remain: free-text prompt exemptions and free-text entity-name matching.", "duration_ms": 172650, "findings": [{"category": "semantic_string_judgment", "evidence": "_FACE_CLOSE_UP_PATTERNS, _is_face_close_up, body_part_focus_rule.trigger_phrases, t.lower() in prompt_lower", "line_end": 230, "line_start": 139, "recommended_fix": "Replace the regex/phrase-list gate with a structured upstream field (for example a `body_part_focus_rule.applies` boolean or `focus_scope` enum) emitted by the prompt-card builder. Keep the text examples only for docs/tests, not for validator routing.", "severity": "P1", "why_problematic": "This path decides whether to skip visible-ID enforcement by matching free-text prompt wording and phrase lists. That is open-world semantic classification by lexical patterns, so wording changes or incidental matches can falsely exempt or reject a prompt."}, {"category": "semantic_string_judgment", "evidence": "name_lower = name.lower(), prompt_lower.find(name_lower, offset), _entity_specific_id_in_window", "line_end": 572, "line_start": 544, "recommended_fix": "Have the producer emit structured mention metadata keyed by entity ID or alias (for example an `entity_mentions` list) and validate that metadata instead of scanning prompt prose. If the metadata cannot be added yet, keep name matching as diagnostics only and remove it from contract enforcement.", "severity": "P1", "why_problematic": "The validator infers entity identity from canonical-name substring matches and a ±60-char / same-sentence heuristic. That is brittle open-world text matching, not structured identity data, so paraphrases, aliases, or incidental name matches will change pass/fail behavior."}], "path": "backend/app/core/visible_entities_validator.py", "scan_kind": "python", "sha256": "366e2185c5df853112075146a74dc8d81d4c36e5fe4441689b7cc81833b80485"}
{"chunk_end": 122, "chunk_start": 1, "chunk_summary": "No actionable findings; this chunk only handles closed-world checkpoint metadata and atomic persistence.", "duration_ms": 12075, "findings": [], "path": "backend/app/modules/image_checkpoint.py", "scan_kind": "python", "sha256": "b253475dfa2921c3c105af37d0654414f3438c8f17adef04d4f4e31fcc8081a4"}
{"chunk_end": 18, "chunk_start": 1, "chunk_summary": "No actionable findings; this abstract base class only defines a structured LLM interface.", "duration_ms": 4969, "findings": [], "path": "backend/app/modules/llm/base.py", "scan_kind": "python", "sha256": "e80908ebc287a538add83673361d6a5652e7508a279e0c95d3940885d1c8b3a9"}
{"chunk_end": 90, "chunk_start": 1, "chunk_summary": "No actionable findings; this chunk only handles closed-world Gemini API key loading and round-robin selection.", "duration_ms": 15559, "findings": [], "path": "backend/app/modules/llm/gemini_key_pool.py", "scan_kind": "python", "sha256": "b0eb3589a24814832578ac71eada7146f061dfa877735336bc58949c7976e48b"}
{"chunk_end": 282, "chunk_start": 1, "chunk_summary": "One actionable issue: free-form prior memory is threaded into the entity-extraction prompt, while the remaining lines are structured schema/enum validation.", "duration_ms": 123867, "findings": [{"category": "scenario_dependent_prompt", "evidence": "prior_memory: Optional[str] = None; memory_block = prior_memory or \"(없음 / None)\"; series_memory_block=memory_block", "line_end": 268, "line_start": 246, "recommended_fix": "Replace prior_memory with a structured memory object (canonical IDs/facts) or remove it from the extractor prompt; if continuity context is needed, pass only validated, machine-readable fields instead of raw text.", "severity": "P1", "why_problematic": "An unconstrained prior-memory string is injected into the extraction prompt, so previous scenario prose can bias the model toward names, places, or props that are not grounded in the current screenplay or a structured SOT."}], "path": "backend/app/modules/entity_extractor_legacy.py", "scan_kind": "python", "sha256": "9edba4cc7829246a732320b6db3828ad8c1692f47fd691cf57b98eb2994998ed"}
{"chunk_end": 180, "chunk_start": 1, "chunk_summary": "This chunk has two actionable string-semantics issues: normalized name lookup on checkpoint text and raw location-string equality both influence dependency selection.", "duration_ms": 517225, "findings": [{"category": "blind_string_mutation", "evidence": "“name_matcher: 원본 + 괄호·공백 정규화 둘 다 키로” and `lookup_name(name_to_sid, name)`", "line_end": 95, "line_start": 71, "recommended_fix": "Carry `short_id`/alias data through upstream checkpoints and resolve shot characters via structured IDs; if aliases are needed, use a canonical alias table from the SOT rather than name normalization in this step.", "severity": "P1", "why_problematic": "This block resolves open-world entity names from checkpoint text by locally normalizing/matching strings, so identity depends on ad hoc text forms instead of a canonical structured ID or alias source. That can silently merge distinct entities or miss valid ones, changing the overlap score."}, {"category": "semantic_string_judgment", "evidence": "`if prev[\"location\"] != cur_loc:` with `cur_loc` coming from `scene_director.primary_location`", "line_end": 142, "line_start": 137, "recommended_fix": "Compare a canonical `location_id` or normalized location key from the scene SOT, or resolve `primary_location` to a stable ID upstream before dependency scoring.", "severity": "P1", "why_problematic": "Same-background routing is decided by exact equality on scenario text from a checkpoint. Any wording variation or paraphrase in `primary_location` changes which prior shot is treated as a dependency."}], "path": "backend/app/core/steps/shot_dependency_step.py", "scan_kind": "python", "sha256": "6db76375617c27cd5402dc6c8b9a6177fc84261bf73fe26f54e2293104acf226"}
{"chunk_end": 82, "chunk_start": 1, "chunk_summary": "No actionable findings in this chunk; this file only logs LLM call payloads and truncates long fields for storage.", "duration_ms": 18007, "findings": [], "path": "backend/app/modules/llm/llm_logger.py", "scan_kind": "python", "sha256": "bc280499bbdb46679b12250b3904c4c770a5cf5128a42d9172beff71e0d09c30"}
{"chunk_end": 221, "chunk_start": 1, "chunk_summary": "No actionable findings in lines 1-221; this client only performs transport, retry, JSON parsing, and logging without open-world string judgment or scenario-specific prompt pollution.", "duration_ms": 53761, "findings": [], "path": "backend/app/modules/llm/gemini_text_client.py", "scan_kind": "python", "sha256": "ab01305191af5c3df4f4b3f5085b038ead3cc5c3a2e918c810577467921fc89b"}
{"chunk_end": 169, "chunk_start": 1, "chunk_summary": "No actionable findings; this chunk only performs closed-world OpenAI response parsing, retries, and generic logging without open-world semantic string judgment or scenario-dependent prompt pollution.", "duration_ms": 15960, "findings": [], "path": "backend/app/modules/llm/openai_client.py", "scan_kind": "python", "sha256": "b09b7d78ae5185fd4ffc7c1a080f9c8f4bb45efeca66da594fef3bdfe7212d2f"}
{"chunk_end": 289, "chunk_start": 1, "chunk_summary": "One actionable finding: the traits formatter falls back to every string field in `stable_traits`, which can leak scenario text into the validation prompt.", "duration_ms": 101576, "findings": [{"category": "scenario_dependent_code", "evidence": "`visual_anchor_traits` and `if isinstance(v, str): visual_traits.append(v)`", "line_end": 164, "line_start": 156, "recommended_fix": "Require a dedicated `visual_anchor_traits` list (or a small whitelist of visual fields) and remove the fallback over `traits_data.items()`. If no visual anchors are present, return an empty traits block or fail fast, and normalize legacy records upstream instead of inferring from arbitrary strings.", "severity": "P2", "why_problematic": "When the explicit visual-anchor list is missing, this code scavenges every string-valued field from the traits dict and injects it into the LLM prompt. That can pull in names, backstory, or other scenario-specific prose instead of a bounded visual SOT, polluting the validator with open-world text that should not steer image judgment."}], "path": "backend/app/modules/image_validator.py", "scan_kind": "python", "sha256": "e6401c59a6af4492ce33f2619f42f239905bc5551008b0b4e4591a93cfe7948b"}
{"chunk_end": 35, "chunk_start": 1, "chunk_summary": "The only actionable issue is a heuristic language-detection function that classifies open-world text by character counts and hard thresholds.", "duration_ms": 14448, "findings": [{"category": "semantic_string_judgment", "evidence": "\"Simple language detection based on character ranges.\"; \"ko_count > 50\"; \"ja_count > 50\"; `return \"en\"`", "line_end": 35, "line_start": 26, "recommended_fix": "Replace this heuristic with upstream language metadata or a dedicated language-ID service/library that returns confidence and an \"unknown\" state; if you only need a hint, expose script counts as diagnostics and avoid branching on them.", "severity": "P1", "why_problematic": "This function infers a semantic property (language) from raw Unicode ranges and fixed thresholds, which is open-world text judgment by string pattern. It is brittle for mixed/short PDFs and bakes in a hard English-default routing assumption."}], "path": "backend/app/modules/pdf_parser.py", "scan_kind": "python", "sha256": "6dae8e6cd873a2b8f8b6bcf8de055862e1c3c0105bbeefe6eff5d75ffc091a79"}
{"chunk_end": 369, "chunk_start": 1, "chunk_summary": "No actionable findings; the file only uses Gemini API response enums/shape and diagnostic logging.", "duration_ms": 116933, "findings": [], "path": "backend/app/modules/llm/gemini_image_client.py", "scan_kind": "python", "sha256": "da5180d7d8c0c04163c45e1c6f95e1d82d0243af4597aefb939e2dfd6baf3162"}
{"chunk_end": 361, "chunk_start": 1, "chunk_summary": "No actionable findings in this chunk.", "duration_ms": 32530, "findings": [], "path": "backend/app/modules/pdf_renderer.py", "scan_kind": "python", "sha256": "24d6b3c7ccf5da2095f93686ae6c526580f18ac2e3bb0a756c255f7a6b00f5ce"}
{"chunk_end": 216, "chunk_start": 1, "chunk_summary": "One low-risk semantic heuristic is used to infer provider labels from model names; otherwise no other actionable findings.", "duration_ms": 125732, "findings": [{"category": "semantic_string_judgment", "evidence": "provider=\"google_ai\" if \"gemini\" in model or \"imagen\" in model else \"fal.ai\"", "line_end": 138, "line_start": 138, "recommended_fix": "Pass provider/model_family explicitly from the caller or resolve it with a centralized exact model-to-provider map; avoid substring-based inference and use a clear fallback such as \"unknown\".", "severity": "P2", "why_problematic": "This infers an open-world semantic label (provider) from free-form model text via substring matching, so trace metadata depends on naming conventions and can silently misclassify new or aliased model IDs."}], "path": "backend/app/modules/llm/image_tracer.py", "scan_kind": "python", "sha256": "0ebe23c2d9b90987dfc52fac1097d9edc98d866707618f5e881a6d5fece24a91"}
{"chunk_end": 36, "chunk_start": 1, "chunk_summary": "No actionable findings; this chunk only resolves worker counts from explicit env names and clamps numeric values without open-world semantic string judging.", "duration_ms": 3816, "findings": [], "path": "backend/app/modules/pipeline/_workers.py", "scan_kind": "python", "sha256": "022738fed30d0efce3ee58a0a4ef4bac167dcd1e61c1257148162d97ea593559"}
{"chunk_end": 58, "chunk_start": 1, "chunk_summary": "Generic DAG batching logic is fine, but the module docstring still embeds scenario-specific phase/caller names.", "duration_ms": 86495, "findings": [{"category": "scenario_dependent_code", "evidence": "Phase 5.3 background_chain_render; Phase 7 floor_plan_render / background_render", "line_end": 4, "line_start": 3, "recommended_fix": "Rewrite the module docstring to describe the batching behavior generically, and move any historical provenance to changelog or caller-local documentation.", "severity": "P2", "why_problematic": "These lines hard-code pipeline-phase and caller names inside a reusable helper docstring, which makes the utility scenario-bound and carries stale domain context instead of a generic DAG-level description."}], "path": "backend/app/modules/pipeline/_dag_levels.py", "scan_kind": "python", "sha256": "50b9da1dd402a115f0a6951f0b978e3ed43213794210fbeec653b06882cc504a"}
{"chunk_end": 281, "chunk_start": 1, "chunk_summary": "One scenario-specific noun is hard-coded into the vision prompt; otherwise no actionable findings.", "duration_ms": 127846, "findings": [{"category": "scenario_dependent_prompt", "evidence": "\"Validate this webbook PDF page.\"", "line_end": 147, "line_start": 147, "recommended_fix": "Change the text to a generic instruction like \"Validate this PDF page.\" and, if document-type context is needed, inject it as a structured prompt variable or source it from the loaded prompt/SOT rather than inline prose.", "severity": "P2", "why_problematic": "This LLM instruction hard-codes the project-specific noun \"webbook\" instead of keeping the validation prompt generic or sourcing document context from structured metadata/prompt templates, which can bias the model and make the validator brittle across scenarios."}], "path": "backend/app/modules/pdf_validator.py", "scan_kind": "python", "sha256": "5cf00222ec7f7687470cb6a65aa529f7ddb6a29be7cc8e0c2464fdd5518ffaaf"}
{"chunk_end": 51, "chunk_start": 1, "chunk_summary": "Name matching is implemented as regex/prefix heuristics over free-form character names, so VE filtering can be wrong; use structured aliases/IDs instead.", "duration_ms": 311743, "findings": [{"category": "semantic_string_judgment", "evidence": "`_PAREN_RE.sub('', name).strip()`, `shot_name == base_name(entity_name)`, `entity_name.startswith(shot_name)`, `if any(match_shot_char(cn, ename) for cn in shot_char_names):`", "line_end": 50, "line_start": 11, "recommended_fix": "Store canonical entity IDs plus approved aliases in the SOT, resolve shot names to IDs before this module, and make `filter_ve_by_shot_chars` compare IDs/set membership only. Keep any parenthetical normalization in data ingestion, not in the matcher, and remove the prefix fallback.", "severity": "P1", "why_problematic": "Free-form character names are normalized by stripping parenthetical text and then compared with exact/prefix string checks to decide whether a `C*` VE survives filtering. That is surface-text heuristic matching on open-world names, so unrelated names can collide and legitimate aliases/variants can be dropped or retained incorrectly."}], "path": "backend/app/modules/name_matcher.py", "scan_kind": "python", "sha256": "431470212b80e7a3bd4b33fcb142e5499adc75515995ec4e849041ce8a75e34e"}
{"chunk_end": 224, "chunk_start": 1, "chunk_summary": "No actionable findings; the visible logic only performs allowed closed-world ID/enum validation and structural checks.", "duration_ms": 164349, "findings": [], "path": "backend/app/modules/pipeline/background_classify.py", "scan_kind": "python", "sha256": "49b9985cf8fb46be2e33e57cd8ed86814808df13e625a334b417f431d1bfeea4"}
{"chunk_end": 499, "chunk_start": 1, "chunk_summary": "No actionable findings; the chunk only uses closed-world ID/schema checks and prompt serialization without open-world string-pattern judgments.", "duration_ms": 89312, "findings": [], "path": "backend/app/modules/pipeline/background_planner.py", "scan_kind": "python", "sha256": "647640786fa18df4962693aba3275f4d1a2c69b11d1cb8821117c14f74bb1d32"}
{"chunk_end": 778, "chunk_start": 1, "chunk_summary": "One actionable finding: the shared client appends a movie/fiction framing suffix to fallback prompts across all call paths.", "duration_ms": 523117, "findings": [{"category": "scenario_dependent_prompt", "evidence": "\"영화 프레이밍 system suffix\"; `safe_system = system_prompt + SAFETY_SYSTEM_SUFFIX` (repeated in `call_structured`, `call_text`, and `call_multiturn`).", "line_end": 766, "line_start": 551, "recommended_fix": "Move the framing suffix out of the shared client and gate it behind explicit per-step/per-project config (for example, a STEP_MANIFEST or project_config flag). Keep fallback retries to generic sanitization unless the caller explicitly opts into that scenario framing.", "severity": "P2", "why_problematic": "A fixed fiction/movie frame is injected into sanitized retry prompts inside a shared LLM client. That is scenario-specific prompt pollution: unrelated tasks inherit extra framing text instead of keeping the fallback path generic and opt-in per workflow."}], "path": "backend/app/modules/llm/llm_client.py", "scan_kind": "python", "sha256": "16dda2eae6b7463c05729ddb95db0d5fbc4a5bacf78e9769c2f11af368df8340"}
{"chunk_end": 280, "chunk_start": 1, "chunk_summary": "One blocking semantic validator still enforces ASCII/English noun form on free-text owned objects; the rest of the chunk is structural.", "duration_ms": 137760, "findings": [{"category": "semantic_string_judgment", "evidence": "item.encode(\"ascii\") ... \"owned MUST be English canonical common nouns\"", "line_end": 210, "line_start": 196, "recommended_fix": "Remove the ASCII/English-only raise from this validator. Keep only list/type/duplication/size checks here, and if canonical object names are required, normalize them in a separate ontology/translation step or explicit enum-backed field while preserving the raw multilingual value.", "severity": "P0", "why_problematic": "This decides validity from character set and an English-noun requirement instead of a structured vocabulary or schema, so legitimate multilingual or domain-specific `objects_owned_by_background` labels are rejected and retried as failures."}], "path": "backend/app/modules/pipeline/background_prompt.py", "scan_kind": "python", "sha256": "de5f48a96c5a19b9b07b12506ddaa30556c2053f398f2284c175be142c76a3b2"}
{"chunk_end": 166, "chunk_start": 1, "chunk_summary": "One actionable finding: moderation detection is driven by keyword matching on exception text, which can misroute the sanitizer retry path.", "duration_ms": 115814, "findings": [{"category": "semantic_string_judgment", "evidence": "msg = str(exc).lower(); is_moderation = any(k in msg for k in (\"moderation\", \"safety\", \"content_policy\", \"prohibited\", \"policy\", \"blocked\", \"violates\", \"violation\"))", "line_end": 138, "line_start": 124, "recommended_fix": "Detect moderation from structured SDK error data instead of message text: use the typed exception, documented error code, HTTP status, or response/body enum fields. Keep `str(exc)` for logging only, and do not use keyword scanning to decide retry routing.", "severity": "P1", "why_problematic": "This treats free-text provider exception prose as a moderation classifier. The branch into PromptSanitizer/retry depends on substring hits in an unstructured message, so wording changes can cause false moderation retries or miss real moderation blocks."}], "path": "backend/app/modules/pipeline/background_render.py", "scan_kind": "python", "sha256": "497a49b7050919807a148e8fa32be91353e756dea2991ba82b21f4691213d15d"}
{"chunk_end": 101, "chunk_start": 1, "chunk_summary": "Entity removal is keyed off LLM-returned names with prefix stripping and exact string matching instead of stable entity IDs, creating brittle semantic routing.", "duration_ms": 72769, "findings": [{"category": "semantic_string_judgment", "evidence": "`raw_name = d[\"name\"]`, `re.sub(r'^[CLP]\\d{2,3}\\s*', '', raw_name).strip()`, `if e[\"name\"] not in rnames`", "line_end": 90, "line_start": 71, "recommended_fix": "Require the schema to return `short_id` (or another canonical entity key) in each decision and filter by `e['short_id']` instead of `e['name']`. Keep any ID-prefix stripping only for logging/debugging, not for routing.", "severity": "P1", "why_problematic": "The filter decision depends on the text form of the LLM output and the entity name spelling, so a prefix, spacing, duplicate-name, or wording variation can remove the wrong entity or miss the intended one. This is open-world semantic routing being decided by string matching rather than a stable key."}], "path": "backend/app/modules/pipeline/entity_filter.py", "scan_kind": "python", "sha256": "6942f826c21f51b159a159a2e14e22d2dd3826fee9b826bf1a8623cadca757d3"}
{"chunk_end": 116, "chunk_start": 1, "chunk_summary": "Both entity extractors rely on dynamic response-key lookups with silent empty-list fallback, which can mask schema drift as a normal empty extraction.", "duration_ms": 145936, "findings": [{"category": "schema_or_enum_drift", "evidence": "`key = f\"{entity_type}s\" if entity_type != \"prop\" else \"props\"` and `result.get(key, [])`", "line_end": 46, "line_start": 45, "recommended_fix": "Use an explicit enum-to-root-key mapping and validate presence with a required lookup (or fail fast in the structured-response validator) instead of defaulting to an empty list.", "severity": "P1", "why_problematic": "The response root key is inferred by ad hoc pluralization, and a missing/renamed key is downgraded to `[]`. That makes a malformed structured response indistinguishable from a legitimate 'no entities' case."}, {"category": "schema_or_enum_drift", "evidence": "`key = f\"{entity_type}s\" if entity_type != \"prop\" else \"props\"` and `result.get(key, [])`", "line_end": 93, "line_start": 92, "recommended_fix": "Apply the same explicit key validation here: map allowed entity types to fixed response keys, assert the key exists, and surface a validation error instead of returning `[]`.", "severity": "P1", "why_problematic": "This repeats the same silent fallback pattern in the list-based extractor, so a schema mismatch or payload drift is hidden and the pipeline returns an empty result as if extraction succeeded."}], "path": "backend/app/modules/pipeline/entity_extractor_v4.py", "scan_kind": "python", "sha256": "e13e19f1e6b9079f28d905440e6d106f69367d417d47349add85469ee51e11bb"}
{"chunk_end": 150, "chunk_start": 1, "chunk_summary": "Heuristic name normalization and short_id ordering are used to seed base/variant entity relations, which is the actionable semantic-string debt in this chunk.", "duration_ms": 129865, "findings": [{"category": "semantic_string_judgment", "evidence": "`base_name(name)` / `members.sort(key=lambda x: (len(x[0]), x[0]))` / `base_short_id`, `variant_short_id`", "line_end": 48, "line_start": 30, "recommended_fix": "Replace this with explicit structured relation metadata from the SOT or entity schema (for example, canonical/variant flags or a relation table). If you need heuristic retrieval, keep it neutral and do not emit base/variant labels or short_id-based direction; let the model or validator determine the actual relation.", "severity": "P1", "why_problematic": "This block infers an open-world relation from string heuristics: it collapses entities by normalized name, then assigns base/variant direction by short_id order. That is not structured relation data, so unrelated names can be grouped together and the downstream LLM is pre-biased by a possibly wrong candidate relation."}], "path": "backend/app/modules/pipeline/entity_relation.py", "scan_kind": "python", "sha256": "702fe956160167b22f8c13170e231f0b3e61c7c143a8d691b9f574bd4f20209b"}
{"chunk_end": 37, "chunk_start": 1, "chunk_summary": "No actionable findings; this chunk only constructs a summary prompt and forwards the episode text to structured LLM summarization.", "duration_ms": 8308, "findings": [], "path": "backend/app/modules/pipeline/episode_summarizer.py", "scan_kind": "python", "sha256": "ab7011132784352e2ada55950c411304521d1c9cde4bb6212c2f263baedef8b5"}
{"chunk_end": 43, "chunk_start": 1, "chunk_summary": "The entity review prompt flattens structured entity records into human-readable strings before the LLM checks omissions/extras, which is a brittle semantic-string mutation.", "duration_ms": 104717, "findings": [{"category": "blind_string_mutation", "evidence": "`_fmt(items)` renders `e['name']` plus `len(e.get('scene_appearances', []))` into a comma-separated string, then injects it into the review prompt.", "line_end": 31, "line_start": 21, "recommended_fix": "Pass the raw entity objects or JSON payload (preserving name, type, scene IDs, evidence, confidence) into the prompt, and keep any human-readable summary separate from the machine-readable review context.", "severity": "P2", "why_problematic": "This collapses structured entity data into a flat surface-form summary, so the reviewer can only reason over names and counts instead of structured fields/evidence. That makes the omission/extraneous-item judgment more dependent on string overlap and other brittle heuristics."}], "path": "backend/app/modules/pipeline/entity_reviewer.py", "scan_kind": "python", "sha256": "fcef947cadd9d7c4eb00bbed3df22fa77435056e08c2d59951dbad87a87a800e"}
{"chunk_end": 115, "chunk_start": 1, "chunk_summary": "No actionable findings in this chunk.", "duration_ms": 22812, "findings": [], "path": "backend/app/modules/pipeline/floor_plan_render.py", "scan_kind": "python", "sha256": "ec4b73480bd0a2ebd76fd42e029f2144bc5201052901c84bd988f77964836551"}
{"chunk_end": 526, "chunk_start": 1, "chunk_summary": "Two actionable findings: a GPT review label is used as a hard entity-pruning gate, and Turn 2 re-injects prior model-generated prose into the next prompt.", "duration_ms": 349517, "findings": [{"category": "semantic_string_judgment", "evidence": "if e.get(\"importance\") == \"none\" and int(e.get(\"appearances\", 0)) < 2", "line_end": 196, "line_start": 191, "recommended_fix": "Keep the review as advisory metadata, or gate pruning on a numeric confidence/evidence signal tied back to the screenplay. Do not delete entities solely because the model emitted the literal label 'none'.", "severity": "P1", "why_problematic": "This turns an LLM's semantic label into a hard delete/routing decision for open-world entities. A borderline or drifted label can remove a valid character/location/prop from all downstream turns, so semantic judgment is hidden behind a string compare instead of source-grounded validation."}, {"category": "scenario_dependent_prompt", "evidence": "\"[시나리오 기반 상세 정보]\" ... desc = gpt_detail.get(\"description\", \"\") ... traits = \", \".join(gpt_detail.get(\"visual_traits\", [])) ... \"+ extra_context\"", "line_end": 452, "line_start": 440, "recommended_fix": "Pass the Turn 1.7 result as structured metadata or evidence refs instead of freeform prose, or omit it from the natural-language prompt and use it only as a non-routing hint.", "severity": "P2", "why_problematic": "The next entity-detail prompt is conditioned on prose synthesized by a previous model pass, not on source evidence. That can reintroduce hallucinated or over-specific details as if they were factual context and bias downstream `description` / `t2i_prompt` generation."}], "path": "backend/app/modules/pipeline/entity_extractor_v2_legacy.py", "scan_kind": "python", "sha256": "491a07d32324e43bc3d0823cf4b20d94b9635a232efd731ff7f180a324f8c123"}
{"chunk_end": 205, "chunk_start": 1, "chunk_summary": "No actionable findings.", "duration_ms": 55321, "findings": [], "path": "backend/app/modules/pipeline/location_floor_plan.py", "scan_kind": "python", "sha256": "ed1c1f785630980449d2de23bce336c12221b32c106d6ff80b5881dab1eb6551"}
{"chunk_end": 284, "chunk_start": 1, "chunk_summary": "One actionable finding: the output validator gates free-form `t2i_prompt` text with a script-range regex, which is a brittle proxy for the actual encoding/language contract.", "duration_ms": 112434, "findings": [{"category": "semantic_string_judgment", "evidence": "_NON_ASCII_TEXT_RE.search(t2i) / \"contains non-ASCII text\"", "line_end": 216, "line_start": 213, "recommended_fix": "If the pipeline truly requires ASCII-only, replace this gate with a deterministic `t2i.isascii()` check (or an explicit schema field such as `prompt_language`/`output_charset`). If the intent is only to block specific scripts, rename the rule to that narrower policy and avoid implying a broader semantic validation.", "severity": "P1", "why_problematic": "This decides whether a free-form generated prompt is valid by matching a handful of Unicode script ranges instead of a structured contract. The rule is mislabeled as \"non-ASCII\": it will reject Korean/Hanja/CJK/Kana text, but still lets other non-ASCII characters through, so semantically correct outputs can fail while genuinely invalid ones slip past."}], "path": "backend/app/modules/pipeline/floor_plan_prompt.py", "scan_kind": "python", "sha256": "2b5573a261a1a27f638ee9a2d6e993d0b343f3e50ac66ccf20fdf19f37421029"}
{"chunk_end": 134, "chunk_start": 1, "chunk_summary": "The only actionable issue is scenario-specific world-rule text being folded into the system prompt; otherwise no actionable findings.", "duration_ms": 124024, "findings": [{"category": "scenario_dependent_prompt", "evidence": "visual_world_rules and [시각적 세계관 규칙 — 인물 물리적 존재 판단 시 반드시 참고]", "line_end": 96, "line_start": 94, "recommended_fix": "Keep the system prompt invariant and pass world rules as structured data (for example, rule IDs plus typed fields in a JSON/context object) or enforce them upstream; do not concatenate raw rule text into the system message.", "severity": "P1", "why_problematic": "This turns a runtime list of world rules into free-form system-prompt prose, so the model’s physical-presence/outlook reasoning can shift with scenario-specific strings instead of a stable structured source of truth."}], "path": "backend/app/modules/pipeline/outlook_extractor.py", "scan_kind": "python", "sha256": "06b9ac1042d32428b3d8c26caa1da8e0be00760eded0d9c3c60abfe756ee8387"}
{"chunk_end": 84, "chunk_start": 1, "chunk_summary": "No actionable findings in this chunk.", "duration_ms": 21524, "findings": [], "path": "backend/app/modules/pipeline/scene_dependency_extractor.py", "scan_kind": "python", "sha256": "6383ce4de75a8cc7f349689891e01fd0dd52d72da8aafe6f8f966a9f413795b4"}
{"chunk_end": 63, "chunk_start": 1, "chunk_summary": "No actionable findings in this chunk.", "duration_ms": 48334, "findings": [], "path": "backend/app/modules/pipeline/scene_dependency_v2.py", "scan_kind": "python", "sha256": "621ea80cec9ecf872356bf7abbca36077cc221d644290becd497ae6faa203cfe"}
{"chunk_end": 255, "chunk_start": 1, "chunk_summary": "No actionable findings; this chunk uses structured IDs/enums and generic prompt context without forbidden string-pattern semantic gating.", "duration_ms": 288124, "findings": [], "path": "backend/app/modules/pipeline/outlook_extractor_v2.py", "scan_kind": "python", "sha256": "5171f914c84b3f5138ec04a16502840b8783021d06d71fee74857d344eabf905"}
{"chunk_end": 125, "chunk_start": 1, "chunk_summary": "No actionable findings in this chunk.", "duration_ms": 11346, "findings": [], "path": "backend/app/modules/pipeline/scene_summarizer.py", "scan_kind": "python", "sha256": "98d3e551a2a11f10bc389764409c065358a34de89f1d5b5bee1128421524c7eb"}
{"chunk_end": 111, "chunk_start": 1, "chunk_summary": "No actionable findings in this chunk.", "duration_ms": 534893, "findings": [], "path": "backend/app/modules/pipeline/outlook_merger.py", "scan_kind": "python", "sha256": "5efbf5429fc2cd12e1d25c9407ccad251840f279c826c7935df59bce83238a76"}
{"chunk_end": 85, "chunk_start": 1, "chunk_summary": "One actionable finding: the scene-director schema silently widens from canonical short_ids to raw entity names when short_id is missing.", "duration_ms": 336027, "findings": [{"category": "schema_or_enum_drift", "evidence": "sid = e.get(\"short_id\", e.get(\"name\", \"?\")); items_def[\"present_entity_ids\"][\"items\"] = {\"type\": \"string\", \"enum\": all_ids}", "line_end": 59, "line_start": 34, "recommended_fix": "Require canonical short_id values for every entity (or normalize through short_id_map before building the schema). If a short_id is missing, fail closed or log an error instead of adding the name to the enum; keep response_schema values restricted to short_ids only.", "severity": "P1", "why_problematic": "The function is documented to use a short_id enum, but this fallback lets scenario-specific names become valid outputs whenever a short_id is absent. That mixes free-text labels into a structured ID contract, so downstream code can no longer rely on canonical IDs and the LLM's answer space drifts with the scene data."}], "path": "backend/app/modules/pipeline/scene_director_v2.py", "scan_kind": "python", "sha256": "a6511cb91674a8bce17425e8c5e5ae9c33f4e3faf7e4df02c2dcef641a50c5a6"}
{"chunk_end": 525, "chunk_start": 1, "chunk_summary": "One actionable schema/enum drift issue: the scene validation payload overloads `severity` and adds skip/unavailable marker fields outside the declared validation schema.", "duration_ms": 361941, "findings": [{"category": "schema_or_enum_drift", "evidence": "`_SCENE_VALIDATION_SCHEMA`; `severity=\"not_run_cost_policy\"`; `_scene_lvm_skipped`; `severity=\"unavailable\"`; `_validation_unavailable`", "line_end": 352, "line_start": 181, "recommended_fix": "Split evaluated validation from execution-state metadata: either return a wrapper like `{\"evaluation_state\": \"evaluated|skipped|unavailable\", \"validation\": {...}, \"skip_reason\": ...}` or define a `oneOf`/union schema with separate evaluated vs skipped variants. Keep `severity` reserved for actual image-quality results only.", "severity": "P1", "why_problematic": "The declared validation schema only covers `matches_prompt`, `severity`, `issues`, and `score`, but the skip and moderation-override branches return extra marker keys and non-schema severity values. That turns one validation object into two different contracts, so strict consumers/schema validators can reject skipped cases or misread them as ordinary failures."}], "path": "backend/app/modules/pipeline/scene_image_pipeline.py", "scan_kind": "python", "sha256": "5d6383246ebb757e1d612256a3abfbc0f964cf92ecc1fd17d0ae09e72357b70d"}
{"chunk_end": 337, "chunk_start": 1, "chunk_summary": "One actionable finding: the post-LLM gaze-exclusion block uses free-text description/name matching and rewrites `visible_entity_ids`, so open-world visibility is no longer LLM-only.", "duration_ms": 161021, "findings": [{"category": "semantic_string_judgment", "evidence": "`base_name(nm)`; `detect_gaze_pattern_exclusions(desc, name_to_char_id)`; `ls['visible_entity_ids'] = [sid for sid in ve_before if sid not in excluded_ids]`", "line_end": 199, "line_start": 171, "recommended_fix": "Remove the deterministic post-process from `visible_entity_ids`. If exclusion metadata is needed, keep it audit-only (or have the LLM/upstream SOT emit a dedicated `excluded_offscreen_entity_ids` field) and never back-propagate it into the visible list; any alias resolution should come from structured entity metadata, not text matching.", "severity": "P1", "why_problematic": "This block turns a free-text shot description plus name aliases into an exclusion decision and then mutates the core `visible_entity_ids` field. That makes open-world visibility depend on a heuristic string match instead of the structured LLM SOT, so paraphrases/aliases can silently drop correct entities and contaminate downstream context."}], "path": "backend/app/modules/pipeline/shot_director.py", "scan_kind": "python", "sha256": "becac71f7b5224ed6be94088c15448f827b6229f53f731ee955c604ab8db9d6d"}
{"chunk_end": 78, "chunk_start": 1, "chunk_summary": "No actionable findings; the prompt is limited to structural PDF text cleanup and does not introduce scenario-dependent string-pattern semantic rules.", "duration_ms": 83257, "findings": [], "path": "backend/app/modules/pipeline/text_cleaner.py", "scan_kind": "python", "sha256": "c9e826bc21f7993a9b265dd8e989e9f58000b6e44a85c1b4bfa2aec7186db266"}
{"chunk_end": 48, "chunk_start": 1, "chunk_summary": "No actionable findings; this chunk only loads a prompt/schema and forwards scenario text to the LLM.", "duration_ms": 7844, "findings": [], "path": "backend/app/modules/pipeline/visual_world_rules.py", "scan_kind": "python", "sha256": "d175b0ca2f77e250eb3e5f62c48a7af863d5ea6426c80c7c7a03b630a4b1c411"}
{"chunk_end": 47, "chunk_start": 1, "chunk_summary": "No actionable findings; this chunk only performs closed-world PNG extension validation and metadata writing.", "duration_ms": 2483, "findings": [], "path": "backend/app/modules/png_metadata.py", "scan_kind": "python", "sha256": "b11942c89529fb36b00d1f2b737d9343b53835d5fbe6dfd85e72b69ee578f1b9"}
{"chunk_end": 86, "chunk_start": 1, "chunk_summary": "No actionable findings; this chunk only persists progress strings and statuses without semantic string judgments or scenario-polluting prompts.", "duration_ms": 4953, "findings": [], "path": "backend/app/modules/progress_tracker.py", "scan_kind": "python", "sha256": "668fae07268d57f39931347c5ce69298e415a505ee85005e39345d4387725908"}
{"chunk_end": 403, "chunk_start": 1, "chunk_summary": "No actionable findings; this chunk only uses closed-world version/config routing and prompt/schema loading logic.", "duration_ms": 93306, "findings": [], "path": "backend/app/modules/prompt_loader.py", "scan_kind": "python", "sha256": "4f096399c0a91a5d5bc3adbc065988b68a1ea77178a4d4a8da063de484233a1b"}
{"chunk_end": 86, "chunk_start": 1, "chunk_summary": "No actionable findings; this file only serializes LLM trace records and does not perform semantic string-based routing or scenario-polluting prompt logic.", "duration_ms": 6260, "findings": [], "path": "backend/app/modules/prompt_tracer_legacy.py", "scan_kind": "python", "sha256": "a3de142f72aa3d16ff0c2108ae34640343b3b6be72783f726101a0cbd41f6538"}
{"chunk_end": 297, "chunk_start": 1, "chunk_summary": "No actionable findings in this chunk.", "duration_ms": 327067, "findings": [], "path": "backend/app/modules/pipeline/shot_staging.py", "scan_kind": "python", "sha256": "2e103d7cf3b31422d96abcbbe9730debfb96d88b64252c5544feab66e91171fb"}
{"chunk_end": 247, "chunk_start": 1, "chunk_summary": "One actionable finding: the final sanitized prompt is mutated based on a brittle prefix-string check against LLM-generated text.", "duration_ms": 134458, "findings": [{"category": "blind_string_mutation", "evidence": "if not sanitized.startswith(strategy[\"prefix\"].strip()[:40]):\n    sanitized = strategy[\"prefix\"] + sanitized", "line_end": 230, "line_start": 229, "recommended_fix": "Do not infer prefix presence from the generated text. Keep the strategy prefix as structured metadata (for example, a `prefix_applied` or `strategy_id` field) and compose the final prompt deterministically outside the model response, or store the body separately and concatenate it without a text-based startswith check.", "severity": "P1", "why_problematic": "This inspects open-ended LLM output with a partial literal prefix match to infer whether the strategy prefix was already applied, then mutates the final prompt based on that surface form. Semantically equivalent outputs or minor wording changes will be treated as missing the prefix, causing duplicate or inconsistent prompt bodies."}], "path": "backend/app/modules/prompt_sanitizer.py", "scan_kind": "python", "sha256": "1683afd2c14de05d98cb740416b11672fa98f0049d44330251143670f2538d90"}
{"chunk_end": 454, "chunk_start": 1, "chunk_summary": "Two P0 heuristics use hardcoded phrase lists and substring/proximity matching to infer visibility and off-camera drift from free-text shot descriptions, which is open-world semantic routing/fail-fast risk.", "duration_ms": 329360, "findings": [{"category": "semantic_string_judgment", "evidence": "`_KOREAN_GAZE_STEMS`, `KOREAN_GAZE_VERB_RE`, `KOREAN_FRAMING_RE`, `any(m in description for m in _CLOSE_UP_MARKERS)`, `name in target_span`, `description.find(d_phrase)`, `window.rfind(name)`", "line_end": 302, "line_start": 36, "recommended_fix": "Move the decision to structured upstream fields (for example `gaze_target_id`, `gaze_subject_id`, or an explicit `exclude_visible_ids` list from shot_staging/scene SOT) and keep any text matcher as non-routing diagnostics only.", "severity": "P0", "why_problematic": "This producer-side block converts arbitrary description prose into visible-entity exclusions using closed verb/noun lists, regexes, and substring/window heuristics, so routing depends on wording instead of structured SOT or explicit entity IDs."}, {"category": "semantic_string_judgment", "evidence": "`OFFSCREEN_RE.finditer(camera_direction)`, `_PROXIMITY_PRE/_PROXIMITY_POST`, `pre_text.rfind(name)`, `if name in post_text`", "line_end": 445, "line_start": 310, "recommended_fix": "Require shot_staging to emit structured drift/offscreen metadata per entity (for example `offscreen_entity_ids` or per-character in-frame flags) and reconcile by exact IDs; keep phrase/proximity scans only as logging/telemetry, not as validation input.", "severity": "P0", "why_problematic": "This consumer-side validator treats natural-language off-camera phrases plus nearby name substrings as authoritative drift evidence, so fail-fast behavior can flip on phrasing rather than on structured in-frame/offscreen state."}], "path": "backend/app/modules/pipeline/shot_visibility.py", "scan_kind": "python", "sha256": "1b551eb799e33ae1cd96ed94b561fda2b0921cd0a2f8d77e4f2ba515d2af15db"}
{"chunk_end": 287, "chunk_start": 1, "chunk_summary": "No actionable findings; the string checks here are limited to closed-world path/version/status metadata.", "duration_ms": 50541, "findings": [], "path": "backend/app/modules/provenance.py", "scan_kind": "python", "sha256": "449131b46eb16420f309910a745d6435af6627edf44c22e71a3472aba467f1fb"}
{"chunk_end": 242, "chunk_start": 1, "chunk_summary": "One production prompt line hardcodes a story-specific example, creating scenario-dependent pollution in the visible-entity verifier.", "duration_ms": 526467, "findings": [{"category": "scenario_dependent_prompt", "evidence": "\"동녘(강의원)\" 같은 표현 → 동녘의 몸만 있고, 강의원은 원격에서 조종 중이므로 강의원은 false", "line_end": 144, "line_start": 144, "recommended_fix": "Remove the named example from the live prompt. If an example is needed, replace it with anonymized placeholders or source it from structured SOT/test fixtures outside production.", "severity": "P1", "why_problematic": "This injects a concrete character name and plot relationship from one scenario into a production validation prompt, which can bias the LLM toward that specific story instead of the structured scene data."}], "path": "backend/app/modules/pipeline/scene_validator.py", "scan_kind": "python", "sha256": "8e27a9499877233d3a1e9ece29fbd4dd3aa5656c9d29dfbaa4b82f4e4671d42d"}
{"chunk_end": 70, "chunk_start": 1, "chunk_summary": "No actionable findings; this chunk only contains closed-world short-ID prefixing and DB mapping logic, with no open-world semantic string judgment or scenario pollution.", "duration_ms": 4919, "findings": [], "path": "backend/app/modules/short_id.py", "scan_kind": "python", "sha256": "a12da612ad174a03abd3fc704048950edc892d574fbad599ae73f53b5b9a325e"}
{"chunk_end": 579, "chunk_start": 1, "chunk_summary": "No actionable findings; the chunk uses structured scene/entity data without open-world string-pattern judgment.", "duration_ms": 79252, "findings": [], "path": "backend/app/modules/scene_image_generator.py", "scan_kind": "python", "sha256": "ecdff80635095f8ac3d4b2171cc50a33304e8e36ba0c278b02703a0e849e06ad"}
{"chunk_end": 63, "chunk_start": 1, "chunk_summary": "No actionable findings; this chunk only defines a generic structured style-rules prompt/schema and shows no string-pattern semantic routing or scenario-specific prompt pollution.", "duration_ms": 52362, "findings": [], "path": "backend/app/modules/style_rules_generator.py", "scan_kind": "python", "sha256": "d371c3250f2dbae4b3655e1b522c6a618ed7cc7b96cbd915a616e1f0c2bb856d"}
{"chunk_end": 396, "chunk_start": 1, "chunk_summary": "Hardcoded Korean baseline-clothing instructions are baked into the character prompt rules, which is the only actionable issue in this chunk.", "duration_ms": 174323, "findings": [{"category": "scenario_dependent_prompt", "evidence": "현대 한국/근미래 한국 기준의 현실적 기본 복장을 사용하라. / Use realistic present-day / near-future Korean baseline clothing.", "line_end": 86, "line_start": 64, "recommended_fix": "Remove the locale-specific clothing rule from the static entity-type list and source any default costume baseline from structured world data (for example `world_guide.costume_guardrails` or a dedicated `clothing_baseline` field) that is appended only when present.", "severity": "P1", "why_problematic": "This injects a fixed Korean present-day/near-future wardrobe assumption into every character reference prompt instead of deriving the costume baseline from `world_guide` or entity metadata, so it can skew outputs for worlds that are not explicitly Korean modern/future."}], "path": "backend/app/modules/reference_image_generator.py", "scan_kind": "python", "sha256": "968475c02cf20d169fef0a67556ee73f3be39cb8ff3f3279f61c1094931e370f"}
{"chunk_end": 152, "chunk_start": 1, "chunk_summary": "One actionable issue: Rule 1 still infers immobilized routing from a hardcoded `gaze_target` phrase list.", "duration_ms": 132772, "findings": [{"category": "semantic_string_judgment", "evidence": "`if gaze not in IMMOBILIZED_GAZE:` and `IMMOBILIZED_GAZE = frozenset({\"dead\", \"unconscious\", \"severely_injured\"})`", "line_end": 102, "line_start": 94, "recommended_fix": "Promote immobility to an upstream structured field/enum (for example `subject_state.immobility_state`) and have `build_semantic_contract()` consume that field; keep any string list only for validation or backward-compatibility tests, not for semantic routing.", "severity": "P1", "why_problematic": "This routes sanitizer behavior from a fixed token list over LLM-produced SOT, so the open-world concept of immobilization is inferred by literal word membership instead of an explicit structured state."}], "path": "backend/app/modules/semantic_contract_router.py", "scan_kind": "python", "sha256": "3eabedd46fc41df8eb486c7c37b479c82096db368111f44188c3cc7912667d56"}
{"chunk_end": 257, "chunk_start": 1, "chunk_summary": "Found a brittle prefix-based scene-heading heuristic that can bias prompt context, plus a language fallback that can mislabel unsupported inputs in the prompt.", "duration_ms": 160718, "findings": [{"category": "semantic_string_judgment", "evidence": "`detect_scene_headings`, `HEADING_PREFIXES`, `HEADING_MAX_LEN`, `_format_heading_catalog`", "line_end": 199, "line_start": 179, "recommended_fix": "Replace the heuristic with upstream screenplay-parser metadata or a dedicated AST-based heading extractor; if you keep a fallback, keep it diagnostics-only and do not use it to build prompt input.", "severity": "P1", "why_problematic": "This decides which raw screenplay lines count as headings using a tiny fixed prefix list and a length cutoff, so nonstandard sluglines are silently omitted before the resulting catalog is fed into the LLM prompt."}, {"category": "scenario_dependent_prompt", "evidence": "`SUPPORTED_LANGUAGE_MAP.get(language, 'Korean')`, `source_language_name=lang_name`", "line_end": 246, "line_start": 229, "recommended_fix": "Validate `language` against `SUPPORTED_LANGUAGE_MAP` and raise on unknown codes, or use an explicit unknown/echoed label instead of hard-coding Korean.", "severity": "P2", "why_problematic": "Unsupported language codes are silently relabeled as Korean in the prompts, which injects incorrect scenario metadata and hides caller bugs instead of surfacing them."}], "path": "backend/app/modules/scene_still_extractor_legacy.py", "scan_kind": "python", "sha256": "baa63faac463130f5931fae18b7f24ec4d990a1643049af5e823c7ca0267931a"}
{"chunk_end": 182, "chunk_start": 1, "chunk_summary": "One low-risk prompt-pollution fallback: the world guide is serialized wholesale when the summary field is missing.", "duration_ms": 130248, "findings": [{"category": "scenario_dependent_prompt", "evidence": "world_guide.get(\"world_setting_summary\", \"\") ... _format_json_block(world_guide)", "line_end": 156, "line_start": 153, "recommended_fix": "Keep a dedicated visual-summary field and use only that in the prompt; if it is missing, omit world_context or build it from a whitelist of prompt-safe keys instead of serializing the full guide.", "severity": "P2", "why_problematic": "If the dedicated summary is absent, the code injects the entire world_guide object into the prompt. That can carry plot, metadata, or other scenario-specific non-visual fields into T2I generation instead of keeping the context visual-only."}], "path": "backend/app/modules/t2i_prompt_composer.py", "scan_kind": "python", "sha256": "eb3ec3ea7ab066c510d18ef2ade7cc2449656a9317a31df1201834b0ca025fc7"}
{"chunk_end": 316, "chunk_start": 1, "chunk_summary": "No actionable findings; this chunk stays within schema/ID validation and generic retry text, with no open-world string-pattern judgments or scenario-specific prompt pollution visible.", "duration_ms": 77153, "findings": [], "path": "backend/app/modules/webbook_generator.py", "scan_kind": "python", "sha256": "9ec93175aa0c2d3159f2896f374778595166dd28f420ea5bc780870523a45941"}
{"chunk_end": 16, "chunk_start": 1, "chunk_summary": "No actionable findings; this file only defines plain Pydantic request/response schemas and does not perform string-based semantic judgment or scenario-dependent prompting.", "duration_ms": 2780, "findings": [], "path": "backend/app/schemas/auth.py", "scan_kind": "python", "sha256": "c2a78063e1a4ff91ccb3722ce4b76ae354c44ad3396920e7ecb3f919beb06ab2"}
{"chunk_end": 20, "chunk_start": 1, "chunk_summary": "No actionable findings; this chunk only defines closed-world response schemas and contains no open-world semantic string judgment or scenario-dependent prompt content.", "duration_ms": 2381, "findings": [], "path": "backend/app/schemas/common.py", "scan_kind": "python", "sha256": "efa97cf6a53a58a07adb5498f1f44e7bdb1ea42e1accf6593cc8128cea8e2b56"}
{"chunk_end": 86, "chunk_start": 1, "chunk_summary": "No actionable findings; this chunk only defines response/update schemas and contains no semantic string judgment or scenario-specific prompt logic.", "duration_ms": 8093, "findings": [], "path": "backend/app/schemas/entity.py", "scan_kind": "python", "sha256": "9ab2573cd9eff34c336c3a09d57d0ff5b1d84d53b976a0c01afd506865a20149"}
{"chunk_end": 178, "chunk_start": 1, "chunk_summary": "No actionable findings in this chunk.", "duration_ms": 114045, "findings": [], "path": "backend/app/modules/variation_recommender.py", "scan_kind": "python", "sha256": "69145ad1e970b20fa74035168c091d03c232cc7c5b192b9f3ca0f6c21c4ba7a0"}
{"chunk_end": 29, "chunk_start": 1, "chunk_summary": "No actionable findings; this schema file only declares Pydantic fields and does not perform open-world semantic string judgments or scenario-dependent prompt pollution.", "duration_ms": 3105, "findings": [], "path": "backend/app/schemas/episode.py", "scan_kind": "python", "sha256": "a72dc87a847d7c558a06f1abcfbcd98dfd41f3e1e40ec2a273018d70ff7c8996"}
{"chunk_end": 63, "chunk_start": 1, "chunk_summary": "No actionable findings; this chunk only defines export/import schemas and response models.", "duration_ms": 6371, "findings": [], "path": "backend/app/schemas/export.py", "scan_kind": "python", "sha256": "30349132c6762b409151f67c5a537f72d87049357a821e4ada65badaa81d0244"}
{"chunk_end": 50, "chunk_start": 1, "chunk_summary": "No actionable findings; this chunk is a pure Pydantic response schema with no semantic string judgment or scenario-dependent prompt text.", "duration_ms": 3416, "findings": [], "path": "backend/app/schemas/operation.py", "scan_kind": "python", "sha256": "c4dee1d33fc72dfaa6a198d9aa276c3e6689d996005cd164e98dd91b8d4cb7b3"}
{"chunk_end": 30, "chunk_start": 1, "chunk_summary": "No actionable findings; this chunk only declares response schemas and contains no semantic string judgment or scenario-specific prompt content.", "duration_ms": 3745, "findings": [], "path": "backend/app/schemas/trace.py", "scan_kind": "python", "sha256": "3bab02539298dd74bc98eddae9ed88e8aefb8aed869b53630b49524baed63926"}
{"chunk_end": 138, "chunk_start": 1, "chunk_summary": "No actionable findings in this chunk.", "duration_ms": 73130, "findings": [], "path": "backend/app/modules/world_guide_generator.py", "scan_kind": "python", "sha256": "8a0040e8cdf5ce23c74f9a2250f7e6cbc45ad1adc140785b5d420a0bb23fdd27"}
{"chunk_end": 16, "chunk_start": 1, "chunk_summary": "User role is modeled as a free-form string despite an explicit two-value set, so the schema and intended enum diverge.", "duration_ms": 49224, "findings": [{"category": "schema_or_enum_drift", "evidence": "role: str  # 'admin' | 'creator' ... role: Optional[str] = None", "line_end": 15, "line_start": 9, "recommended_fix": "Define a shared `UserRole` `Enum` or `Literal['admin', 'creator']` and use it in both `UserCreate` and `UserUpdate`; if roles are meant to be data-driven, remove the closed-list comment and validate from a structured source of truth.", "severity": "P2", "why_problematic": "The schema documents a closed role set but accepts any string, which pushes role validation and branching into ad hoc string comparisons elsewhere instead of enforcing the domain at the boundary."}], "path": "backend/app/schemas/user.py", "scan_kind": "python", "sha256": "454ee8dcddb1a547cf1b2d44876d36861449a1fb0892266ba0570686ce43e627"}
{"chunk_end": 68, "chunk_start": 1, "chunk_summary": "No actionable findings in this chunk.", "duration_ms": 3971, "findings": [], "path": "backend/app/services/auth_service.py", "scan_kind": "python", "sha256": "baa838999fa13b569c7daa02f52487f6d17f6303caea0fd3b9eb12c9ec22584f"}
{"chunk_end": 26, "chunk_start": 1, "chunk_summary": "No actionable findings; this package init only re-exports sync services and documents the canonical projection flow without semantic string judgment or scenario-dependent prompt content.", "duration_ms": 3279, "findings": [], "path": "backend/app/services/checkpoint_sync/__init__.py", "scan_kind": "python", "sha256": "6950afde87f19ddbf7a7f2ef7501c20cab26b30cf2e3b0d9707d643560377a44"}
{"chunk_end": 85, "chunk_start": 1, "chunk_summary": "No actionable findings; this chunk contains only schema declarations, comments, and a closed-world file-path normalizer.", "duration_ms": 73089, "findings": [], "path": "backend/app/schemas/image.py", "scan_kind": "python", "sha256": "36fc5caeec3ffedca36d21dbabe52fd31526363862bd65fd022d47d0cd1caca2"}
{"chunk_end": 78, "chunk_start": 1, "chunk_summary": "No actionable findings; this chunk is limited to structural dataclass contract definitions and comments, with no open-world string judgments or scenario-dependent prompt pollution.", "duration_ms": 7415, "findings": [], "path": "backend/app/services/checkpoint_sync/_scene_still_contracts.py", "scan_kind": "python", "sha256": "97663cc81d9af81732e0760ef170dcc3f5ac46cbc24d4b9eefa3e1fbb447714b"}
{"chunk_end": 774, "chunk_start": 1, "chunk_summary": "No actionable findings; this chunk uses closed-world enums/IDs and structured manifest lookups only.", "duration_ms": 64736, "findings": [], "path": "backend/app/services/analysis_dispatch_service.py", "scan_kind": "python", "sha256": "9c128dce46be495f153fd8f15a3349a8ca94ad23146bbd68ac57d13747444d1e"}
{"chunk_end": 293, "chunk_start": 1, "chunk_summary": "No actionable findings; this chunk uses only closed-world status and ID checks.", "duration_ms": 63675, "findings": [], "path": "backend/app/services/checkpoint_sync/episode_projection_service.py", "scan_kind": "python", "sha256": "09b092c7f0cf9cb4d25f17fe560deed5bdc44e387b9d67359d9482a6071b5cbd"}
{"chunk_end": 47, "chunk_start": 1, "chunk_summary": "The chunk has stringly typed status and role fields in the project/member schemas even though the allowed values are documented as closed sets.", "duration_ms": 171376, "findings": [{"category": "schema_or_enum_drift", "evidence": "status: Optional[str] = None  # 'active' | 'archived' | 'deleted' / status: str", "line_end": 20, "line_start": 13, "recommended_fix": "Replace this with a shared `ProjectStatus` `Literal['active', 'archived', 'deleted']` or `Enum`, and use that type anywhere project status is accepted or returned.", "severity": "P1", "why_problematic": "The project-status contract is documented as a finite set but implemented as free-form strings across the request/response schema, so invalid statuses can slip through and later code must re-implement the contract with string checks."}, {"category": "schema_or_enum_drift", "evidence": "role: str  # 'owner' | 'member' / role: str", "line_end": 43, "line_start": 30, "recommended_fix": "Replace these with a shared `MemberRole` `Literal['owner', 'member']` or `Enum`, and use it for `MemberAdd.role`, `MemberUpdate.role`, and `MemberResponse.role`.", "severity": "P1", "why_problematic": "Member roles are documented as a closed set but modeled as free-form strings in the add/update/response schemas, which allows unsupported roles and pushes validation into ad hoc string comparisons."}], "path": "backend/app/schemas/project.py", "scan_kind": "python", "sha256": "db325ae49a0a3e2eb855ff3392b7ffa2aa1f5d4963bd2f4a21842a04217c4a02"}
{"chunk_end": 105, "chunk_start": 1, "chunk_summary": "Runtime logic is closed-world on status enums; the only actionable debt is scenario-specific prose in the shared is_cp_syncable docstring.", "duration_ms": 125688, "findings": [{"category": "scenario_dependent_code", "evidence": "scene_detail 1 shot fail → status=\"partial\"; Track B P1-2 partial 정책; [\"scenes\"], [\"characters\", \"locations\", \"props\"], [\"outlooks\"]", "line_end": 42, "line_start": 19, "recommended_fix": "Rewrite the docstring as a neutral contract summary only (completed/partial/other behavior), and move incident notes plus example key lists to a changelog, test fixture, or structured policy/SOT document.", "severity": "P2", "why_problematic": "This shared base docstring mixes a generic sync contract with incident history, phase tags, and domain-specific example keys. That is scenario pollution in common code and can leak into downstream docs or prompt text instead of coming from a structured source of truth."}], "path": "backend/app/services/checkpoint_sync/_base.py", "scan_kind": "python", "sha256": "63961d4d3b025abc4d5fbcaed4c929ef1250aa9f4c05f8f9f8f6c9f4a8d5c491"}
{"chunk_end": 260, "chunk_start": 1, "chunk_summary": "One actionable issue: entity canon upserts are keyed by raw `name`, so open-world text can silently collide or drift across syncs.", "duration_ms": 159852, "findings": [{"category": "semantic_string_judgment", "evidence": "existing canon 매핑 (name → id) / existing_canons = {e.name: e ...} / existing = existing_canons.get(name)", "line_end": 139, "line_start": 47, "recommended_fix": "Key the lookup on a stable structured identifier from the checkpoint (persistent source UUID, source_entity_key, or equivalent). If only names are available, use at least a composite key like (entity_type, normalized_name) plus explicit duplicate/ambiguity checks, and fail closed instead of overwriting on collision.", "severity": "P1", "why_problematic": "This treats a human-readable entity name as the canonical join key for update-vs-insert. Names are open-world strings: they can rename, vary by checkpoint, or collide across entity types, and the dict comprehension will silently drop duplicates, so the sync can merge the wrong rows or create duplicates."}], "path": "backend/app/services/checkpoint_sync/entity_sync_service.py", "scan_kind": "python", "sha256": "d93b0ef98a45c52fe596c45bdb0c0f69aa161af00e7f7c4ec167ba1570200581"}
{"chunk_end": 153, "chunk_start": 1, "chunk_summary": "One actionable debt: repair-mode status transitions are keyed off free-form error text, making failed→synced repair depend on message wording instead of a stable token.", "duration_ms": 124896, "findings": [{"category": "semantic_string_judgment", "evidence": "\"sync_error = :prev_err\"", "line_end": 49, "line_start": 40, "recommended_fix": "Persist and compare a stable repair token (for example a failure/attempt UUID, exception class/code, or row-version column) and keep sync_error as diagnostic text only; if exact matching is required, compare a machine-generated fingerprint instead of the raw message.", "severity": "P1", "why_problematic": "This uses a human-readable exception message as the predicate for a state transition, so wording drift, truncation, or localization can silently block a valid repair or match the wrong failure."}], "path": "backend/app/services/checkpoint_sync/orchestrator.py", "scan_kind": "python", "sha256": "2ddbb1e40fcb91e679e13b21f95d97b571430d3503dcf8accc184e668024ffd0"}
{"chunk_end": 201, "chunk_start": 1, "chunk_summary": "One actionable finding: a hardcoded prose fallback is written into persisted relation reasons when the checkpoint omits them.", "duration_ms": 84440, "findings": [{"category": "scenario_dependent_code", "evidence": "rel.get(\"reason\", \"시각적 변형 — 기본 요소에 의존\")", "line_end": 64, "line_start": 62, "recommended_fix": "Preserve missing reasons as `None`/empty (or fail validation upstream) and only render a fallback label in presentation code, or source the fallback from a structured SOT/config rather than embedding prose in the sync path.", "severity": "P2", "why_problematic": "When `reason` is missing, the sync layer invents a human-language explanation and stores it as authoritative `continuity_reason`, so absent source data becomes hardcoded prose instead of a structured missing-value state."}], "path": "backend/app/services/checkpoint_sync/relation_sync_service.py", "scan_kind": "python", "sha256": "6e92eee3c6bdec60861c5972e32d868f03819c1b7c7efa46cdf47f1c3147694e"}
{"chunk_end": 97, "chunk_start": 1, "chunk_summary": "No actionable findings in this chunk.", "duration_ms": 35765, "findings": [], "path": "backend/app/services/checkpoint_sync/scene_still_sync_service.py", "scan_kind": "python", "sha256": "713434e0a55bff3a92581f12cc3044077f046ac52a3c0bd5c798c95749abd83f"}
{"chunk_end": 108, "chunk_start": 1, "chunk_summary": "One scenario-specific comment should be generalized; otherwise no actionable findings.", "duration_ms": 88874, "findings": [{"category": "scenario_dependent_code", "evidence": "\"1~2 씬 fail이 64 shot 차단하지 않음 (2026-05-01 사고)\"", "line_end": 29, "line_start": 28, "recommended_fix": "Rewrite the comment to state the general rule only (for example, that partial checkpoints should not block remaining scene sync and empty cascade output is unsyncable), or move the historical incident reference to an external ticket/ADR.", "severity": "P2", "why_problematic": "This permanent code comment bakes a specific incident date and shot-count example into the implementation notes, which is scenario-specific baggage rather than a reusable invariant."}], "path": "backend/app/services/checkpoint_sync/scene_still_checkpoint_loader.py", "scan_kind": "python", "sha256": "7f7f750650e328e72bbb53317eedb31699f08e50e97fb52d864fc211747470a6"}
{"chunk_end": 198, "chunk_start": 1, "chunk_summary": "One actionable issue: `_shot_ve` parses generated prompt text to infer visible entities, so semantic meaning is decided by string matching.", "duration_ms": 64771, "findings": [{"category": "semantic_string_judgment", "evidence": "`t2i_prompt` join, `_BARE_ID_RE.findall(text)`, `return [sid for sid in director_ve if sid in used]`", "line_end": 76, "line_start": 69, "recommended_fix": "Add an explicit structured field such as `shot_vars[*].visible_entity_ids` or `mentioned_entity_ids` and derive `visible_entities_json` from that field only. If a regex heuristic is needed during migration, keep it diagnostic-only and out of production normalization/routing.", "severity": "P1", "why_problematic": "This treats generated `t2i_prompt` prose as an authoritative source for which entities are visible by regex-matching bare IDs and then filtering `director_ve` against that set. That is open-world semantic inference from text, not closed-world validation, so wording changes can silently alter planned still metadata."}], "path": "backend/app/services/checkpoint_sync/scene_still_normalizer.py", "scan_kind": "python", "sha256": "c0c29435dd7a4a4aafa0fa5d348a19901db4489d5eef09f5904ab3bf52d3e12f"}
{"chunk_end": 213, "chunk_start": 1, "chunk_summary": "No actionable findings; this chunk only propagates structured checkpoint fields and closed-world IDs without string-based semantic judgment.", "duration_ms": 49367, "findings": [], "path": "backend/app/services/checkpoint_sync/scene_still_writer.py", "scan_kind": "python", "sha256": "1f52aa60a54f0f34744d909cec355179d8d2a900cb87266cc37ff1ae2cd39704"}
{"chunk_end": 248, "chunk_start": 1, "chunk_summary": "No actionable findings; this chunk uses closed-world ID/path handling and standard database/file operations, with no clear open-world string-pattern semantic judgment or scenario-dependent prompt pollution.", "duration_ms": 58243, "findings": [], "path": "backend/app/services/episode_service.py", "scan_kind": "python", "sha256": "f1eaaedaff9fdc20ef72f00e40125caeb7adef762f13ef365f1eb67d9c1edcdd"}
{"chunk_end": 241, "chunk_start": 1, "chunk_summary": "No actionable findings; the module uses generic image-selection prompts and numeric angle clamping without closed-list semantic or scenario-specific string judgments.", "duration_ms": 41698, "findings": [], "path": "backend/app/services/fal_angle_helpers.py", "scan_kind": "python", "sha256": "f77edf941603c41d5c372a5b0f61292ed7ebc77b0fa39e076087e278ebbf247e"}
{"chunk_end": 248, "chunk_start": 1, "chunk_summary": "No actionable findings; this chunk only performs closed-world JSON/ID handling and passes prompt overrides through without hardcoded semantic string judgments.", "duration_ms": 38091, "findings": [], "path": "backend/app/services/image_composer_service.py", "scan_kind": "python", "sha256": "3542b129ba7f9899f2936d64d2b65b09fac2c32783ffaf153e3e3cc35d1c3fc7"}
