{"candidate_reason": "python scope discovery", "chunk_end": 104, "chunk_start": 1, "chunk_summary": "No actionable findings; the file contains standard database migration logic for index creation using closed-world schema identifiers.", "duration_ms": 3994, "findings": [], "path": "backend/alembic/versions/001_add_indexes.py", "scan_kind": "python", "sha256": "443d404214d268646a852a16b9fb0dae2be7d48f7e5246ca1280fc7ed869aed2"}
{"candidate_reason": "python scope discovery", "chunk_end": 48, "chunk_start": 1, "chunk_summary": "No actionable findings; the file contains standard Alembic migration logic for database schema management using closed-world column name checks.", "duration_ms": 4076, "findings": [], "path": "backend/alembic/versions/006_llm_call_log_metadata_json.py", "scan_kind": "python", "sha256": "1b72ded9be069eae18d32f41b6f416b9804e4877927fce0731f13475004634bf"}
{"candidate_reason": "python scope discovery", "chunk_end": 70, "chunk_start": 1, "chunk_summary": "No actionable findings; this is a standard Alembic environment configuration file using closed-world schema definitions.", "duration_ms": 4448, "findings": [], "path": "backend/alembic/env.py", "scan_kind": "python", "sha256": "0ebc7e60b781371151de2a3e04377b883bf6dc5258eec68551ce6d915acb66c0"}
{"candidate_reason": "python scope discovery", "chunk_end": 60, "chunk_start": 1, "chunk_summary": "No actionable findings; the file is a standard Alembic migration script managing database schema columns.", "duration_ms": 4447, "findings": [], "path": "backend/alembic/versions/007_d6_entity_canon_metadata_json.py", "scan_kind": "python", "sha256": "dd607b5a3858f9f017e6f5f9b32eb2c1e81a61faf1a53b7ba9dbcdf8ce30c307"}
{"candidate_reason": "python scope discovery", "chunk_end": 46, "chunk_start": 1, "chunk_summary": "No actionable findings; the file contains standard FastAPI authentication endpoints using closed-world protocol and session identifiers.", "duration_ms": 2935, "findings": [], "path": "backend/app/api/v1/auth.py", "scan_kind": "python", "sha256": "897f59b16fd7a5874d60124f0bfeb419ae516004bf924b003684733adf31e85e"}
{"candidate_reason": "python scope discovery", "chunk_end": 65, "chunk_start": 1, "chunk_summary": "No actionable findings; this is a standard database migration script defining schema changes for recovery tracking and asset labeling.", "duration_ms": 7199, "findings": [], "path": "backend/alembic/versions/003_resume_integrity.py", "scan_kind": "python", "sha256": "8ce8dd8d21121f9b607f2ad1c9dac6df425a9f83d9b80fb807be0988d41667f9"}
{"candidate_reason": "python scope discovery", "chunk_end": 154, "chunk_start": 1, "chunk_summary": "The file provides standard FastAPI dependencies for database sessions, authentication, and project access control, with no actionable findings regarding semantic string judgment or scenario pollution.", "duration_ms": 11990, "findings": [], "path": "backend/app/api/deps.py", "scan_kind": "python", "sha256": "483314672232f13f52c439ddde4113e4cc80d4d9c9ecedee6b0a6f9313f87cfb"}
{"candidate_reason": "python scope discovery", "chunk_end": 72, "chunk_start": 1, "chunk_summary": "The migration script correctly adds structural columns (variant_index, variant_label, t2i_guide) to the image_asset table without implementing any pattern-based semantic logic.", "duration_ms": 16556, "findings": [], "path": "backend/alembic/versions/002_phase5_image_asset_variants.py", "scan_kind": "python", "sha256": "e4e87dc3c8f6b93e5f1dd0ea46beeece8ebe6c77003fb5157d59d60b5ad47783"}
{"candidate_reason": "python scope discovery", "chunk_end": 212, "chunk_start": 1, "chunk_summary": "no actionable findings", "duration_ms": 10103, "findings": [], "path": "backend/app/api/v1/exports.py", "scan_kind": "python", "sha256": "2020ad132c8eb0656b13fabe61e1981975228e27cd7c09e7ec805cd852f16059"}
{"candidate_reason": "python scope discovery", "chunk_end": 70, "chunk_start": 1, "chunk_summary": "no actionable findings", "duration_ms": 4483, "findings": [], "path": "backend/app/api/v1/operations.py", "scan_kind": "python", "sha256": "0f7f23fe55215e99777c21681da6ba640ac71643716099be716f6c272a926691"}
{"candidate_reason": "python scope discovery", "chunk_end": 422, "chunk_start": 1, "chunk_summary": "No actionable findings; the file serves as an API router using closed-world constants for status and operation tracking.", "duration_ms": 18271, "findings": [], "path": "backend/app/api/v1/episodes.py", "scan_kind": "python", "sha256": "63af6d843b5f1747846e3e921cf2787ed67a29846924d55773b26004c2473c65"}
{"candidate_reason": "python scope discovery", "chunk_end": 60, "chunk_start": 1, "chunk_summary": "The migration documentation reveals a design where semantic axes are concatenated into a single string label, leading to schema fragility and pattern-based semantic judgment.", "duration_ms": 24146, "findings": [{"category": "semantic_string_judgment", "evidence": "location_axis1_axis2_axis3_axisN", "line_end": 8, "line_start": 7, "recommended_fix": "Refactor the variant_label to use a structured format (e.g., JSONB) or a separate metadata table to store individual semantic axes.", "severity": "P1", "why_problematic": "Encoding multiple semantic dimensions into a single string forces downstream components to use string parsing to interpret scene structure, which is brittle and bypasses structured data validation."}, {"category": "scenario_dependent_code", "evidence": "night_body_blood_curtain_red_circle", "line_end": 5, "line_start": 3, "recommended_fix": "Generalize the documentation to describe the technical failure (length overflow) without embedding specific scenario content.", "severity": "P2", "why_problematic": "The migration docstring contains scenario-specific content strings, which is a form of scenario-dependent pollution in the codebase."}], "path": "backend/alembic/versions/004_variant_label_extend.py", "scan_kind": "python", "sha256": "4d8a88e7e7074489bb79c48cbb9cb62b23bf132db35add90d286cb0f93954b52"}
{"candidate_reason": "python scope discovery", "chunk_end": 921, "chunk_start": 1, "chunk_summary": "The file contains several instances of semantic judgment based on string patterns, including image type filtering by prompt substrings and entity extraction via regex from open-world text.", "duration_ms": 21612, "findings": [{"category": "semantic_string_judgment", "evidence": "if not img.prompt_used or (\"outlook_id:\" not in img.prompt_used and \"outfit:\" not in img.prompt_used and \"composite:\" not in img.prompt_used):", "line_end": 366, "line_start": 366, "recommended_fix": "Add a structured 'is_face_reference' boolean or 'reference_type' enum to the ImageAsset model instead of parsing the prompt string.", "severity": "P1", "why_problematic": "This logic determines if an image is a 'face reference' by checking for the absence of specific substrings within a prompt string. This is a fragile semantic judgment based on open-world text patterns."}, {"category": "semantic_string_judgment", "evidence": "_re.finditer(r'\\[\\[([^\\]]+)\\]\\+\\[([^\\]]+)\\]\\]', t2i_text)", "line_end": 433, "line_start": 433, "recommended_fix": "Use a formal parser for the T2I DSL or store these entity relationships in a structured 'visible_entities' field rather than extracting them from the prompt text via regex.", "severity": "P1", "why_problematic": "It uses regex to extract character and outlook names (open-world semantic entities) from a prompt string. This relies on a specific bracketed syntax being present in the middle of open-world text."}, {"category": "semantic_string_judgment", "evidence": "ImageAsset.prompt_used.like(f\"%composite:{c_canon.id}:{o_canon.id}%\")", "line_end": 461, "line_start": 461, "recommended_fix": "Create a join table or a JSONB metadata field to store composite entity IDs on the ImageAsset model.", "severity": "P1", "why_problematic": "It identifies a 'composite' image relationship by searching for a specific string pattern inside a prompt field. This is a semantic link stored as a string pattern rather than a structured database relation."}, {"category": "llm_closed_list_instruction", "evidence": "\"'Photorealistic cinematic still.' 같은 스타일 접두어도 그대로 유지하세요.\"", "line_end": 792, "line_start": 792, "recommended_fix": "Move specific style strings to a project-level configuration or a 'protected_phrases' list passed to the translation function.", "severity": "P2", "why_problematic": "The prompt instructs the LLM to treat a specific open-world phrase as a fixed token. This is scenario-dependent pollution in a translation utility that should ideally be agnostic to specific style strings."}], "path": "backend/app/api/v1/entities.py", "scan_kind": "python", "sha256": "e5f92eef6cfdc114ebcc82cef44fd9975cb444244f497a6c11c87c707a1717fe"}
{"candidate_reason": "python scope discovery", "chunk_end": 99, "chunk_start": 1, "chunk_summary": "The migration script enforces relative file paths but contains hardcoded logic duplicated across multiple system components, risking schema drift.", "duration_ms": 26281, "findings": [{"category": "schema_or_enum_drift", "evidence": "CHECK_EXPR = \"file_path = '' OR file_path NOT LIKE '/%'\"", "line_end": 34, "line_start": 30, "recommended_fix": "Centralize the CHECK_EXPR string in a shared constants file or a database schema utility that can be imported by the migration, the model definition, and the startup scripts.", "severity": "P1", "why_problematic": "The logic for validating relative paths is hardcoded in three separate locations (migration, database startup block, and ORM model). This duplication increases the risk of inconsistent validation if the definition of a valid path changes or needs to account for different environments, violating the principle of a single source of truth for domain nomenclature."}], "path": "backend/alembic/versions/005_file_path_relative_check.py", "scan_kind": "python", "sha256": "63389a9d05e169558dc925ede63fceb3a0d990d8cc66b4ca85c640f9a2fe056e"}
{"candidate_reason": "python scope discovery", "chunk_end": 84, "chunk_start": 1, "chunk_summary": "No actionable findings; the file contains standard FastAPI CRUD endpoints for user management without semantic string judgment or scenario pollution.", "duration_ms": 3699, "findings": [], "path": "backend/app/api/v1/users.py", "scan_kind": "python", "sha256": "6976912b349a4e82df9ee64157335cca7cb55bc0178040ecc9669a78001d75d6"}
{"candidate_reason": "python scope discovery", "chunk_end": 236, "chunk_start": 1, "chunk_summary": "The file provides a standard CRUD API for prompt template management with versioning support and contains no actionable findings regarding semantic string judgment or scenario-dependent pollution.", "duration_ms": 10612, "findings": [], "path": "backend/app/api/v1/prompts.py", "scan_kind": "python", "sha256": "122494141f3e88c71eda968f511c7e8d06566fc93f08738a132bae1706040aa4"}
{"candidate_reason": "python scope discovery", "chunk_end": 268, "chunk_start": 1, "chunk_summary": "no actionable findings", "duration_ms": 8009, "findings": [], "path": "backend/app/api/v1/steps.py", "scan_kind": "python", "sha256": "475e38d587f4bc934cf5ebd1edd7a6d457e47f1204a977998d8ccfb592964670"}
{"candidate_reason": "python scope discovery", "chunk_end": 82, "chunk_start": 1, "chunk_summary": "No actionable findings; the file provides standard atomic JSON I/O utilities using closed-world path and UUID logic.", "duration_ms": 3075, "findings": [], "path": "backend/app/core/checkpoint_io.py", "scan_kind": "python", "sha256": "9921f13f1dd3fafa94bdb8d7ed2f6a0642f772854ae7dcc168958e3df2a4ddea"}
{"candidate_reason": "python scope discovery", "chunk_end": 249, "chunk_start": 1, "chunk_summary": "The configuration file defines application settings and performs standard closed-world syntax and security validation for environment variables and IDs.", "duration_ms": 10756, "findings": [], "path": "backend/app/core/config.py", "scan_kind": "python", "sha256": "08605a2d225b7e37cf36b1211e63abc722c2a5b55a6dd1f3afb3e3896fc7f034"}
{"candidate_reason": "python scope discovery", "chunk_end": 298, "chunk_start": 1, "chunk_summary": "The file is a registry for step applicability validators and contains no actionable findings regarding semantic string judgment or scenario-dependent pollution.", "duration_ms": 15276, "findings": [], "path": "backend/app/core/applicability.py", "scan_kind": "python", "sha256": "8cc4c734e80e9b97e773f8c16db66ae397ee7d17a4477f8f86c75e34f41eb45e"}
{"candidate_reason": "python scope discovery", "chunk_end": 570, "chunk_start": 1, "chunk_summary": "The file manages project and member CRUD operations, but contains a hardcoded string heuristic for classifying LLM providers based on model name prefixes.", "duration_ms": 23151, "findings": [{"category": "semantic_string_judgment", "evidence": "provider = \"openai\" if current_model.startswith(\"gpt\") else \"gemini\"", "line_end": 523, "line_start": 523, "recommended_fix": "Retrieve the provider by looking up the model's metadata in the AVAILABLE_MODELS registry using the current_model alias, rather than performing a prefix check.", "severity": "P1", "why_problematic": "The code uses a hardcoded substring check ('gpt') on a model identifier to infer its provider. This is a fragile heuristic for open-world model names (which are added frequently) and bypasses the structured metadata available in the model registry (AVAILABLE_MODELS)."}], "path": "backend/app/api/v1/projects.py", "scan_kind": "python", "sha256": "3aa863e13caa986487d44188e1d22e782e543cb053da74423a706f1c39a18079"}
{"candidate_reason": "python scope discovery", "chunk_end": 330, "chunk_start": 1, "chunk_summary": "The file implements a deterministic background ID catalog system using structured profiles and semantic keys, avoiding pattern-based semantic judgment on open-world text.", "duration_ms": 14653, "findings": [], "path": "backend/app/core/bg_catalog.py", "scan_kind": "python", "sha256": "2225fe163ccb224e71c66211be106160a1aaef1c39fac8d4e2cf93eacbe4a480"}
{"candidate_reason": "python scope discovery", "chunk_end": 8, "chunk_start": 1, "chunk_summary": "No actionable findings; the file contains standard package initialization and exports for DTOs.", "duration_ms": 2625, "findings": [], "path": "backend/app/core/dto/__init__.py", "scan_kind": "python", "sha256": "08aff79b29f622f9a3b9d2722fd95e43d14ee3c512de50b4cd6fee5912c05d64"}
{"candidate_reason": "python scope discovery", "chunk_end": 264, "chunk_start": 1, "chunk_summary": "The file contains standard SQLAlchemy database configuration and idempotent raw SQL migrations for schema evolution, with no actionable findings regarding open-world semantic judgment or scenario pollution.", "duration_ms": 13051, "findings": [], "path": "backend/app/core/database.py", "scan_kind": "python", "sha256": "3255c3bb7dbac684feae1f4fa2b966b762df9cb41a1ef2a41d02ae7e68c73abe"}
{"candidate_reason": "python scope discovery", "chunk_end": 484, "chunk_start": 1, "chunk_summary": "The module implements asset readiness validation but relies on substring patterns in prompt fields and hardcoded semantic state strings to determine asset requirements.", "duration_ms": 20654, "findings": [{"category": "semantic_string_judgment", "evidence": "ImageAsset.prompt_used.like(\"%composite:%\")", "line_end": 252, "line_start": 252, "recommended_fix": "Migrate legacy data to ensure the asset_type column is correctly populated and query by that column instead of using LIKE on prompt_used.", "severity": "P1", "why_problematic": "Determines if an asset is a 'composite' type by performing a substring search on a text prompt field. This is a pattern-based judgment of asset nature instead of relying on structured metadata or the asset_type column."}, {"category": "llm_closed_list_instruction", "evidence": "_STATE_LABELS = (\"dead\", \"severely_injured\", \"unconscious\")", "line_end": 305, "line_start": 305, "recommended_fix": "Use a centralized SOT or Enum for character states that is shared between the manifest generator and the validator.", "severity": "P1", "why_problematic": "Defines a closed list of semantic character states. If the upstream LLM or manifest generator uses synonyms (e.g., 'deceased', 'fainted'), the readiness check will silently skip these requirements, leading to the fallback behavior this module is intended to prevent."}, {"category": "semantic_string_judgment", "evidence": "ImageAsset.prompt_used.like(\"%state_variant:%\")", "line_end": 325, "line_start": 325, "recommended_fix": "Query by the asset_type column directly; ensure the producer step populates this column correctly.", "severity": "P1", "why_problematic": "Similar to the composite check, it identifies 'state_variant' assets by parsing a string pattern within a prompt field rather than using structured type identifiers."}, {"category": "semantic_string_judgment", "evidence": "gaze = ca.get(\"gaze_target\", \"\") ... if gaze in (\"dead\", \"severely_injured\", \"unconscious\")", "line_end": 436, "line_start": 435, "recommended_fix": "Update the manifest schema to include a dedicated 'physical_state' or 'status' field and use a structured Enum for validation.", "severity": "P0", "why_problematic": "Overloads the 'gaze_target' field (semantically a spatial/directional attribute) to detect character physical states. This is a high-risk semantic drift where open-world story states are judged by string matching against a mislabeled field."}], "path": "backend/app/core/asset_readiness.py", "scan_kind": "python", "sha256": "254da5a0e1790186c83b8f42077b113e9f92c2d3fa3fcd1d4fd01d9551f8a5af"}
{"candidate_reason": "python scope discovery", "chunk_end": 137, "chunk_start": 1, "chunk_summary": "The file defines a SceneAnalysisContext DTO which serves as a structured container for pipeline data, and no actionable findings were identified.", "duration_ms": 11919, "findings": [], "path": "backend/app/core/dto/scene_analysis.py", "scan_kind": "python", "sha256": "5d17c2fdffaa87103bf4484442ea4485b1da846eb1a343445430c5fd3cd8ce84"}
{"candidate_reason": "python scope discovery", "chunk_end": 205, "chunk_start": 1, "chunk_summary": "no actionable findings", "duration_ms": 11923, "findings": [], "path": "backend/app/core/entity_metadata.py", "scan_kind": "python", "sha256": "893d93150185304ee8d6b294842e6b970f8a7590185eb9681df3095795d0be49"}
{"candidate_reason": "python scope discovery", "chunk_end": 175, "chunk_start": 1, "chunk_summary": "no actionable findings", "duration_ms": 10048, "findings": [], "path": "backend/app/core/file_paths.py", "scan_kind": "python", "sha256": "39921edd78a9cb4a0ade74a2728088eff7b2d104f6ee7f95c19b8da0a8992679"}
{"candidate_reason": "python scope discovery", "chunk_end": 63, "chunk_start": 1, "chunk_summary": "No actionable findings; the file implements strict closed-world enum validation for framing scales without semantic string guessing.", "duration_ms": 4063, "findings": [], "path": "backend/app/core/framing_scale.py", "scan_kind": "python", "sha256": "a48840fd0a5c37397cb26a10a57d4bd4333071aa55ad4bae0f6db74f786ca4e3"}
{"candidate_reason": "python scope discovery", "chunk_end": 42, "chunk_start": 1, "chunk_summary": "no actionable findings", "duration_ms": 2474, "findings": [], "path": "backend/app/core/job_manager.py", "scan_kind": "python", "sha256": "fee1d9730d64a20295f3dbc70474835423be6fa6632498e616d282a36a410696"}
{"candidate_reason": "python scope discovery", "chunk_end": 191, "chunk_start": 1, "chunk_summary": "The file implements entity protection logic using structured manifest data and graph relations, avoiding keyword-based semantic judgment.", "duration_ms": 15928, "findings": [], "path": "backend/app/core/entity_protection.py", "scan_kind": "python", "sha256": "e181a4a35b8014686e298d8c09967a00e3be7867fd2c0fd005f3836515faaaa6"}
{"candidate_reason": "python scope discovery", "chunk_end": 46, "chunk_start": 1, "chunk_summary": "The file defines dataclasses for integrity reporting with allowed status enums and contains no actionable semantic string judgments or scenario pollution.", "duration_ms": 7408, "findings": [], "path": "backend/app/core/integrity_report.py", "scan_kind": "python", "sha256": "746d4ecf33b66cd7e4c2b3d94f850ba980b2ca941cf49393fbaeef54dc4bce79"}
{"candidate_reason": "python scope discovery", "chunk_end": 35, "chunk_start": 1, "chunk_summary": "no actionable findings", "duration_ms": 3317, "findings": [], "path": "backend/app/core/logging_config.py", "scan_kind": "python", "sha256": "b6e8ab379da48f01ee8f87b1878a8d25a6bb38859dce8238521a89ee7f1a4d4c"}
{"candidate_reason": "python scope discovery", "chunk_end": 62, "chunk_start": 1, "chunk_summary": "no actionable findings", "duration_ms": 5102, "findings": [], "path": "backend/app/core/low_freq_skip.py", "scan_kind": "python", "sha256": "5a119ae11d83690059124ebe4ab43860481a2bac301864835aaf9d9e972f9f9d"}
{"candidate_reason": "python scope discovery", "chunk_end": 130, "chunk_start": 1, "chunk_summary": "No actionable findings; the module performs strict schema and enum validation for keep_elements without inspecting or judging open-world semantic content.", "duration_ms": 8442, "findings": [], "path": "backend/app/core/keep_elements.py", "scan_kind": "python", "sha256": "1fa17ab361a1f74991d78c8965aa818897f82f6e9cbd652423308c3f853293fe"}
{"candidate_reason": "python scope discovery", "chunk_end": 85, "chunk_start": 1, "chunk_summary": "No actionable findings; the file provides a technical utility for pipeline caching based on content hashes without semantic judgment.", "duration_ms": 6328, "findings": [], "path": "backend/app/core/pipeline_cache.py", "scan_kind": "python", "sha256": "0efc4fe218ff2295389ecfb153eede4bfa30670817c8b6059307c716350ee318"}
{"candidate_reason": "python scope discovery", "chunk_end": 363, "chunk_start": 1, "chunk_summary": "The file contains semantic validation logic using substring matching against hardcoded phrase lists and embedded Korean prompt instructions for LLM retries.", "duration_ms": 22271, "findings": [{"category": "semantic_string_judgment", "evidence": "ZONE_PHRASES: Dict[str, List[str]] = { ... }", "line_end": 41, "line_start": 25, "recommended_fix": "Define these concepts in a structured SOT and use LLM-based semantic evaluation or VLM verification instead of keyword matching.", "severity": "P1", "why_problematic": "Hardcoded lists of natural language synonyms (e.g., 'top-left', 'center-right') are used to judge the semantic presence of spatial concepts in open-world T2I prompts via substring matching."}, {"category": "semantic_string_judgment", "evidence": "any(v in prompt_lower for v in zone_variants)", "line_end": 324, "line_start": 322, "recommended_fix": "Replace substring matching with an LLM-based check that evaluates if the prompt's intent matches the spatial constraints.", "severity": "P1", "why_problematic": "The phrase_diagnostic function uses substring checks (lower().in(...)) to determine if an open-world prompt satisfies semantic spatial constraints. This is brittle and fails to account for linguistic variety."}, {"category": "scenario_dependent_prompt", "evidence": "\"[재시도 — 직전 응답의 frame_spatial_contract 가 다음 위반을 포함했습니다. 수정해서 재출력하세요:]\"", "line_end": 342, "line_start": 341, "recommended_fix": "Move the retry hint formatting and instructional text to a centralized prompt template system or localization resource.", "severity": "P2", "why_problematic": "Hardcoded Korean instructional text for LLM retry hints is embedded directly in the core logic file, creating scenario-dependent pollution."}], "path": "backend/app/core/frame_spatial_contract.py", "scan_kind": "python", "sha256": "dc8cd3121d7ca06ba098541ab7a40a722f0c79326fa6c4f65ed6b0ded2483f80"}
{"candidate_reason": "python scope discovery", "chunk_end": 31, "chunk_start": 1, "chunk_summary": "No actionable findings; the file contains standard security utility functions for password hashing and session management using closed-world primitives.", "duration_ms": 2785, "findings": [], "path": "backend/app/core/security.py", "scan_kind": "python", "sha256": "efce0d75e3a2e4ccdc8c675afcc00854a5af7fbe5799da350b7caf2bf30957ad"}
{"candidate_reason": "python scope discovery", "chunk_end": 276, "chunk_start": 1, "chunk_summary": "The pipeline gate logic uses regex and substring matching on the 'prompt_used' text field to identify composite assets and extract entity IDs, which is a brittle semantic judgment.", "duration_ms": 10177, "findings": [{"category": "semantic_string_judgment", "evidence": "ImageAsset.prompt_used.like(\"%composite:%\"), re.search(r'composite:([a-f0-9-]+):([a-f0-9-]+)', p)", "line_end": 156, "line_start": 144, "recommended_fix": "Store composite metadata (character_id, outlook_id) in structured columns or a JSONB metadata field in the ImageAsset model, and query those directly instead of performing string searches on the prompt.", "severity": "P1", "why_problematic": "The code determines if an image is a 'composite' and extracts associated entity IDs by parsing the 'prompt_used' string field. This field typically contains the natural language prompt or a decorated version of it. Using regex to drive pipeline validation logic from a text field is brittle and conflates metadata with the generative prompt."}, {"category": "semantic_string_judgment", "evidence": "ImageAsset.prompt_used.like(\"%outlook_id:%\")", "line_end": 238, "line_start": 238, "recommended_fix": "Use a structured 'asset_subtype' or 'metadata' column to identify composite/outlook-specific assets.", "severity": "P1", "why_problematic": "Pipeline status calculation relies on substring matching within the prompt text to count completed composite images. This is an open-world semantic judgment that fails if the prompt format changes or if the string appears naturally in a non-composite prompt."}], "path": "backend/app/core/pipeline_gate.py", "scan_kind": "python", "sha256": "7c845f2b6194ec9c311aa1825f965c6553892db648cd29cb090e356e7921c46f"}
{"candidate_reason": "python scope discovery", "chunk_end": 218, "chunk_start": 1, "chunk_summary": "The file defines application-specific error classes, one of which (VisibleStagingDriftError) documents a reliance on string-pattern matching within natural language descriptions to determine semantic visibility drift.", "duration_ms": 32704, "findings": [{"category": "semantic_string_judgment", "evidence": "camera_direction NL 에 명시적 off-camera/off-screen/화면 밖 phrase 가 있는데", "line_end": 104, "line_start": 103, "recommended_fix": "Refactor the upstream logic to use structured boolean flags or enums for visibility/off-screen status in the SOT, rather than parsing natural language strings.", "severity": "P1", "why_problematic": "The system determines semantic state (visibility) by searching for specific natural language substrings ('off-camera', etc.) in a description field. This is an open-world semantic judgment performed via string patterns rather than structured SOT data."}], "path": "backend/app/core/errors.py", "scan_kind": "python", "sha256": "ba89fa3d25ef305a82f0bc1b664d5f5b01a78c04656a81067792cb165042f7b2"}
{"candidate_reason": "python scope discovery", "chunk_end": 83, "chunk_start": 1, "chunk_summary": "The SettingsRegistry provides a centralized interface for feature flags and model selection, currently relying on environment fallbacks and step catalog defaults.", "duration_ms": 2838, "findings": [{"category": "scenario_dependent_code", "evidence": "return \"gpt\"", "line_end": 80, "line_start": 80, "recommended_fix": "Move the global default model identifier to app.core.config.settings or a dedicated constants file to avoid hardcoding specific model names in the registry logic.", "severity": "P1", "why_problematic": "Hardcoded fallback to 'gpt' as a default model name is a scenario-dependent string judgment. If the step catalog or project config fails to provide a model, the system defaults to a specific provider/model string rather than a configuration-driven default or raising a configuration error."}], "path": "backend/app/core/settings_registry.py", "scan_kind": "python", "sha256": "a75c16fd433e6ca3dbba88f5a6ed532714c3670dc287f6992b8cef98fc7d62d7"}
{"candidate_reason": "python scope discovery", "chunk_end": 213, "chunk_start": 1, "chunk_summary": "The file contains hardcoded semantic enums for narrative states and location types, which are used to enforce closed-list constraints on LLM outputs, leading to scenario-dependent pollution in core logic.", "duration_ms": 48546, "findings": [{"category": "scenario_dependent_code", "evidence": "STATE_CLASS_ENUM: FrozenSet[str] = frozenset({\"normal\", \"quiet\", \"busy\", \"busy_exit\", \"ransacked\", \"clean_after\", \"blood_scene\", \"intrusion\", \"arrival\", \"evidence_display\", \"dream_or_vision_state\"})", "line_end": 38, "line_start": 26, "recommended_fix": "Move the state vocabulary to a dynamic SOT or scenario configuration file and inject it into the validator.", "severity": "P1", "why_problematic": "Hardcoding narrative states like 'blood_scene' and 'ransacked' in core logic is scenario-dependent pollution. These are open-world semantic concepts that should be defined in a scenario-specific SOT. The code (lines 48-55) enforces this closed list on LLM outputs, causing retries on valid but non-matching semantic descriptions."}, {"category": "scenario_dependent_code", "evidence": "LOCATION_SPACE_KEY_VOCAB: FrozenSet[str] = frozenset({\"main\", \"kitchen\", \"rooftop\", \"stairs\", \"yard\", \"exterior\", \"office\"})", "line_end": 136, "line_start": 128, "recommended_fix": "Inject allowed space keys from the scenario definition or a structured SOT into the validation function instead of hardcoding them.", "severity": "P1", "why_problematic": "Hardcoding architectural labels like 'kitchen' and 'rooftop' restricts the system to specific environments. This list is tightly coupled with the entity_extractor system prompt (line 126), creating scenario-dependent pollution in a core utility file."}, {"category": "scenario_dependent_code", "evidence": "if kind == \"single_space\" and allowed != [\"main\"]:", "line_end": 205, "line_start": 202, "recommended_fix": "Allow the default space key to be defined in the SOT or scenario configuration.", "severity": "P2", "why_problematic": "Hardcodes 'main' as the only valid space key for single-space locations. This is a scenario-dependent semantic assumption that may not hold for all environments."}], "path": "backend/app/core/bg_state_vocab.py", "scan_kind": "python", "sha256": "d6c2d1fccafa20ba72e40e90ca8a76f14d4225eebe57c37a205c010e50e9aafd"}
{"candidate_reason": "python scope discovery", "chunk_end": 161, "chunk_start": 1, "chunk_summary": "The name matching logic uses regex patterns to strip bracketed suffixes and classify entities, which constitutes open-world semantic judgment via string patterns.", "duration_ms": 18443, "findings": [{"category": "blind_string_mutation", "evidence": "_BRACKET_PATTERN.sub(\"\", s)", "line_end": 74, "line_start": 51, "recommended_fix": "Replace string-based normalization with an explicit alias system or canonical ID mapping in the database.", "severity": "P1", "why_problematic": "Stripping bracketed content from character names to determine a 'bare' identity is a heuristic for semantic matching that can fail in open-world scenarios where brackets are part of the name or where multiple variants exist."}, {"category": "semantic_string_judgment", "evidence": "if not _BRACKET_PATTERN.search(raw)", "line_end": 141, "line_start": 138, "recommended_fix": "Use structured metadata (e.g., a 'is_variant' flag or 'parent_id') in the EntityCanon model to drive indexing logic.", "severity": "P1", "why_problematic": "The code classifies entities as 'base' or 'variant' based on the presence of brackets, which is a pattern-based judgment of open-world semantic relationships."}], "path": "backend/app/core/name_matcher.py", "scan_kind": "python", "sha256": "6a06d3cde26de3c8fec7c158b189b543c3b86df0d99b9e9cdda903b03f62f797"}
{"candidate_reason": "python scope discovery", "chunk_end": 125, "chunk_start": 1, "chunk_summary": "No actionable findings; the file contains a standard registry mapping internal step identifiers to their respective classes.", "duration_ms": 2944, "findings": [], "path": "backend/app/core/steps/__init__.py", "scan_kind": "python", "sha256": "7dfe3f4a7e6f7f2fb6c660abc77c47b48790b723847c664652cecacc26c88b42"}
{"candidate_reason": "python scope discovery", "chunk_end": 225, "chunk_start": 1, "chunk_summary": "The file serves as a centralized registry for pipeline steps, mapping manifest data to runner classes using internal identifiers and enums; no actionable semantic string judgments or scenario pollutions were found.", "duration_ms": 13829, "findings": [], "path": "backend/app/core/step_catalog.py", "scan_kind": "python", "sha256": "a9349acfefc4619ad2fbd259fab29f7803b2677237d2b146736d9687bef7ef3d"}
{"candidate_reason": "python scope discovery", "chunk_end": 116, "chunk_start": 1, "chunk_summary": "The file implements a judge wrapper for validating owned objects in T2I prompts using an LLM, with appropriate schema validation and no actionable semantic string judgment findings.", "duration_ms": 2409, "findings": [], "path": "backend/app/core/steps/_owned_judge.py", "scan_kind": "python", "sha256": "6280d7a602b422b4d3a0815310a128f83cc6eacf7b0ab3ba4944eb703268fdcb"}
{"candidate_reason": "python scope discovery", "chunk_end": 1154, "chunk_start": 1, "chunk_summary": "The file is a structural manifest defining pipeline steps, dependencies, and metadata using closed-world IDs and enums; no actionable findings were identified.", "duration_ms": 20094, "findings": [], "path": "backend/app/core/step_manifest.py", "scan_kind": "python", "sha256": "ca41b2e1efde1513640c9b57ce37a6fc42665ec7a0b111f849ba847a76fba548"}
{"candidate_reason": "python scope discovery", "chunk_end": 400, "chunk_start": 1, "chunk_summary": "The file provides technical helpers for normalizing and validating 'owned objects' data structures, focusing on schema integrity, ASCII constraints, and deterministic hashing without semantic string judgment.", "duration_ms": 18557, "findings": [], "path": "backend/app/core/steps/_owned_helpers.py", "scan_kind": "python", "sha256": "b9b4d7f230a479d642279017c5296d7a471f7e09c4e32cf0a9a6dcb1fc915dec"}
{"candidate_reason": "python scope discovery", "chunk_end": 129, "chunk_start": 1, "chunk_summary": "The file provides structural validation and backfilling for evidence-related fields in LLM outputs, using allowed enum checks and schema-based logic.", "duration_ms": 21372, "findings": [], "path": "backend/app/core/steps/_evidence_helpers.py", "scan_kind": "python", "sha256": "1b11a4a10b12e90373ed92fd688d5eb5db3a49220fc3c31d94987b7fdf77b9aa"}
{"candidate_reason": "python scope discovery", "chunk_end": 166, "chunk_start": 1, "chunk_summary": "The file contains hardcoded workflow logic and scenario-dependent prompt fragments (Korean labels) within the planning context helper.", "duration_ms": 36941, "findings": [{"category": "scenario_dependent_code", "evidence": "if section == \"characters_text\" and not self.is_first_episode:", "line_end": 47, "line_start": 44, "recommended_fix": "Remove the hardcoded episode check from the helper and allow the caller or the prompt template to determine which sections to include.", "severity": "P1", "why_problematic": "This hardcodes a semantic workflow rule that character and relationship information is only relevant for the first episode. This logic should be part of the prompt template or pipeline configuration, not buried in a core context helper, as it prevents these sections from being used in later episodes if needed."}, {"category": "scenario_dependent_prompt", "evidence": "f\"## 기획서 참고: {section}\", f\"\\n  외형: {c['visual_traits']}\", and manual string assembly in lines 138-146", "line_end": 164, "line_start": 52, "recommended_fix": "Use a template engine or a structured prompt builder to separate data formatting and localization from the core logic.", "severity": "P2", "why_problematic": "Hardcoded Korean labels and manual string formatting for character/relationship data are embedded in the Python logic. This creates scenario-dependent pollution and makes the prompt structure fragile and difficult to localize."}], "path": "backend/app/core/planning_doc_context.py", "scan_kind": "python", "sha256": "96a5e66e6ef2bd8f932a78d98da7cf04cc0bb01723a449a16a4d79929afbe7ea"}
{"candidate_reason": "python scope discovery", "chunk_end": 1597, "chunk_start": 1, "chunk_summary": "The StepRunner class uses substring matching on human-readable error messages to drive conditional routing logic for checkpoint mismatches.", "duration_ms": 27075, "findings": [{"category": "semantic_string_judgment", "evidence": "if \"schema_version mismatch\" in mismatch_reason:", "line_end": 752, "line_start": 752, "recommended_fix": "Refactor _check_cp_mismatch to return a structured result (e.g., a dataclass or Enum) that explicitly identifies the mismatch type (SCHEMA vs CONFIG_HASH) separately from the descriptive reason string.", "severity": "P1", "why_problematic": "Routing logic for allowing reruns on specific steps depends on substring matching against a human-readable string returned by _check_cp_mismatch. This creates a tight coupling between logging/error messages and execution flow, making the system brittle to message changes."}], "path": "backend/app/core/step_runner.py", "scan_kind": "python", "sha256": "1bc98850586701b74aec05907a3329fb336d1b8e4412efc5724e836aeafc121a"}
{"candidate_reason": "python scope discovery", "chunk_end": 629, "chunk_start": 1, "chunk_summary": "The file manages the background chain rendering step, including data loading from previous checkpoints, calling the rendering module, and synchronizing results with the database, with no actionable findings regarding semantic string judgment or scenario pollution.", "duration_ms": 14777, "findings": [], "path": "backend/app/core/steps/background_chain_render_step.py", "scan_kind": "python", "sha256": "db955002c4caf6db1a7f2927358a8e72e4d761f4dfa4bec668d8a4c36376f52f"}
{"candidate_reason": "python scope discovery", "chunk_end": 457, "chunk_start": 1, "chunk_summary": "The validator uses hardcoded noun and verb lists to perform open-world semantic classification of prompt text and instruction intent to enforce reference attachment contracts.", "duration_ms": 46659, "findings": [{"category": "semantic_string_judgment", "evidence": "_CHARACTER_TOKENS, _BACKGROUND_TOKENS, _OBJECT_TOKENS", "line_end": 62, "line_start": 49, "recommended_fix": "Replace post-hoc string classification with structured metadata from the prompt generation stage or use an LLM-based classifier.", "severity": "P1", "why_problematic": "Hardcoded lists of open-world nouns (body parts, furniture, electronics) are used to classify the semantic target of phrases in the prompt."}, {"category": "semantic_string_judgment", "evidence": "_GENERIC_VERB_TOKENS, _GENERIC_NOUN_TOKENS", "line_end": 88, "line_start": 87, "recommended_fix": "Identify instruction types during prompt construction and pass them as structured metadata.", "severity": "P1", "why_problematic": "Keyword lists are used to judge the semantic intent (generic vs specific) of prompt instructions to decide whether to skip validation."}, {"category": "semantic_string_judgment", "evidence": "def _classify_window(window: str) -> str:", "line_end": 118, "line_start": 91, "recommended_fix": "Use a semantic embedding similarity check or delegate classification to the component that generated the prompt.", "severity": "P1", "why_problematic": "Implements semantic classification of open-world text by searching for the 'nearest' noun from a closed list, which is prone to false positives/negatives in natural language."}, {"category": "semantic_string_judgment", "evidence": "def _is_plural_reference_images(prompt: str, end_pos: int) -> bool:", "line_end": 128, "line_start": 121, "recommended_fix": "Use a proper NLP lemmatizer or structured metadata to determine the plurality and intent of prompt subjects.", "severity": "P1", "why_problematic": "Uses a regex check on a fixed-length substring slice to determine if a noun is plural, which is used to decide semantic intent."}, {"category": "semantic_string_judgment", "evidence": "def _has_generic_instruction_signal(prompt: str, phrase_start: int) -> bool:", "line_end": 153, "line_start": 131, "recommended_fix": "Use structured instruction tags or an LLM to classify the nature of the instruction.", "severity": "P1", "why_problematic": "Determines prompt intent ('generic instruction') using substring checks and keyword presence, which is a fragile way to handle open-world prompt semantics."}], "path": "backend/app/core/ref_contract_validator.py", "scan_kind": "python", "sha256": "93a34075f568486f74bd384f96f18208bab54f8e5777726c6d7beef955a7b504"}
{"candidate_reason": "python scope discovery", "chunk_end": 369, "chunk_start": 1, "chunk_summary": "no actionable findings", "duration_ms": 15384, "findings": [], "path": "backend/app/core/steps/background_planner_step.py", "scan_kind": "python", "sha256": "53d33dcd02ffd8685de2d501c41a75e5556ef64a21aa374b9bfa3ed61ecdbfb8"}
{"candidate_reason": "python scope discovery", "chunk_end": 519, "chunk_start": 1, "chunk_summary": "The file orchestrates the background master plan step, primarily handling data flow between checkpoints and calling external modules for LLM interaction and ID assignment without performing open-world semantic string judgment.", "duration_ms": 17226, "findings": [], "path": "backend/app/core/steps/background_master_plan_step.py", "scan_kind": "python", "sha256": "f316adda85aac755b904d1cc15c8d3ff7ed749c8c9584a6289f4d9cd457d646c"}
{"candidate_reason": "python scope discovery", "chunk_end": 264, "chunk_start": 1, "chunk_summary": "no actionable findings", "duration_ms": 19254, "findings": [], "path": "backend/app/core/steps/background_classify_step.py", "scan_kind": "python", "sha256": "1ffb71738766db06c062a8702c057a161718c17d4b721d15138ca50a742f33eb"}
{"candidate_reason": "python scope discovery", "chunk_end": 995, "chunk_start": 1, "chunk_summary": "The file contains scenario-dependent prompt instructions regarding character possession and hardcoded string mutations for entity ID resolution fallbacks.", "duration_ms": 34300, "findings": [{"category": "semantic_string_judgment", "evidence": "int(e.get(\"importance\", 50)) < 10 and int(e.get(\"appearances\", 0)) < 2", "line_end": 132, "line_start": 129, "recommended_fix": "Move importance filtering logic to the LLM prompt or allow project-specific thresholds in ProjectSettings.", "severity": "P2", "why_problematic": "Hardcoded numeric thresholds are used to decide the semantic importance of entities, which should be handled by the LLM or a more flexible world-rule configuration rather than fixed code logic."}, {"category": "scenario_dependent_prompt", "evidence": "인물: 자기 물리적 몸으로 존재하는 인물만. 대사를 하더라도 빙의/원격접속 중이면 제외. ... 앞쪽 씬에서 인물A가 인물B의 몸에 접속/빙의/라이드했다면", "line_end": 817, "line_start": 810, "recommended_fix": "Move scenario-specific logic (like possession rules) into a dynamic 'world rules' block or a specialized prompt template for specific genres.", "severity": "P1", "why_problematic": "The prompt contains specific scenario-dependent logic (possession/remote control) that may not apply to all projects, polluting the general analysis pipeline with specific story tropes."}, {"category": "blind_string_mutation", "evidence": "raw_id.replace(\"CHAR_\", \"\").replace(\"BG_\", \"\").replace(\"PROP_\", \"\")", "line_end": 863, "line_start": 845, "recommended_fix": "Use a more robust fuzzy matching or strictly enforce the schema via the LLM client. If fallbacks are needed, use the structured SOT (sid_info) without hardcoded string prefixes.", "severity": "P1", "why_problematic": "The code uses blind string replacement and hardcoded prefixes ('CHAR_', 'BG_', 'PROP_') to resolve entity IDs when the LLM fails to follow the requested short-ID format. This is scenario-dependent nomenclature and fragile."}], "path": "backend/app/core/steps/analysis_steps_legacy.py", "scan_kind": "python", "sha256": "5309486c802bcbd8c46d563320b1b652467837e59f9414993412e25513d9e3a3"}
{"candidate_reason": "python scope discovery", "chunk_end": 366, "chunk_start": 1, "chunk_summary": "The file is a pipeline step orchestrator that manages background prompt generation using structured IDs and checkpoints; no actionable findings related to semantic string judgment or scenario pollution were found.", "duration_ms": 15210, "findings": [], "path": "backend/app/core/steps/background_prompt_step.py", "scan_kind": "python", "sha256": "9318a00456c3f8bb0eaf22772efac3dc477b8daba55892adea74588c6be52359"}
{"candidate_reason": "python scope discovery", "chunk_end": 214, "chunk_start": 1, "chunk_summary": "The StepRunner orchestrates background chain planning by loading multiple checkpoints and preparing floor plan prompts, but relies on string-based markers to control logic gates.", "duration_ms": 33243, "findings": [{"category": "semantic_string_judgment", "evidence": "FLOOR PLAN 블록 있으면 outdoor skip 금지", "line_end": 101, "line_start": 18, "recommended_fix": "Introduce a structured boolean field (e.g., 'has_floor_plan') in the location/floor_plan schema and pass this flag explicitly to the planning logic instead of relying on string-based markers in the prompt context.", "severity": "P1", "why_problematic": "The pipeline logic (specifically the 'outdoor skip' rule) is triggered or suppressed based on the presence of a specific string block '[FLOOR PLAN]' within the prompt text. This is a semantic judgment based on a string pattern rather than a structured boolean flag or SOT attribute."}], "path": "backend/app/core/steps/background_chain_planning_step.py", "scan_kind": "python", "sha256": "0fed58a03fae448319826f84ba117fe61266792824f6f7500ea105d649ad5d39"}
{"candidate_reason": "python scope discovery", "chunk_end": 766, "chunk_start": 1, "chunk_summary": "The file is a pipeline step runner for background rendering and database synchronization, using structural ID validation and schema-based data mapping without actionable semantic string judgment or scenario pollution.", "duration_ms": 21032, "findings": [], "path": "backend/app/core/steps/background_render_step.py", "scan_kind": "python", "sha256": "61f88d19935e0b1efe05d6cdf48d004f8901bd7705a8577ad4c8f9516bf5d7b9"}
{"candidate_reason": "python scope discovery", "chunk_end": 95, "chunk_start": 1, "chunk_summary": "The file implements a step runner for entity relation extraction, constructing a context string from beats and shots with hardcoded Korean labels.", "duration_ms": 17620, "findings": [{"category": "scenario_dependent_prompt", "evidence": "f\"[씬 {si}]\" and \"(beat/shot 데이터 없음)\"", "line_end": 95, "line_start": 85, "recommended_fix": "Move context labels to a configuration file or use a localized template system for building LLM context strings to decouple logic from scenario language.", "severity": "P2", "why_problematic": "Hardcoded Korean labels are used to structure the context string passed to the LLM. This introduces language-specific scenario pollution into the prompt construction logic, making the pipeline dependent on specific domain nomenclature baked into the code."}], "path": "backend/app/core/steps/entity_relation_step.py", "scan_kind": "python", "sha256": "b1a46b315c5130b07a2b6a7569645a2432bfda4bcba2a1d7a27a110c321ebe71"}
{"candidate_reason": "python scope discovery", "chunk_end": 304, "chunk_start": 1, "chunk_summary": "no actionable findings", "duration_ms": 10531, "findings": [], "path": "backend/app/core/steps/floor_plan_prompt_step.py", "scan_kind": "python", "sha256": "3d0f71b1f1d384eeb6b46210236a9910106809dec6aed137cf941ee4f45ab8be"}
{"candidate_reason": "python scope discovery", "chunk_end": 107, "chunk_start": 1, "chunk_summary": "The step runner extracts a character list using an LLM, but contains hardcoded Korean section headers in the prompt construction logic.", "duration_ms": 31528, "findings": [{"category": "scenario_dependent_code", "evidence": "f\"[씬 {seg['scene_index']}]\\n{text}\" and f\"[시각적 규칙]\\n{visual_rules}\\n\\n{user_prompt}\"", "line_end": 71, "line_start": 54, "recommended_fix": "Move the section headers into the prompt template (e.g., using Jinja2 variables in the user prompt) or pass the data as a structured dictionary to the LLM client if it supports template-based assembly.", "severity": "P2", "why_problematic": "Hardcoding Korean section headers ('[씬]', '[시각적 규칙]') directly in the Python logic couples the step runner to a specific language and prompt formatting strategy. This makes the pipeline less flexible for other languages or different prompt engineering approaches."}], "path": "backend/app/core/steps/character_list_step.py", "scan_kind": "python", "sha256": "ee4727eab61e855b22a5898064a08f163d1b34b4b45a0d79472a6489ae8b5254"}
{"candidate_reason": "python scope discovery", "chunk_end": 329, "chunk_start": 1, "chunk_summary": "The SceneDirectorStep contains hardcoded genre-specific semantic filters and embedded prompt instructions in Korean, which should be managed via structured templates.", "duration_ms": 30330, "findings": [{"category": "scenario_dependent_code", "evidence": "if r.get(\"rule_type\") in (\"possession\", \"projection\", \"ghost\"):", "line_end": 127, "line_start": 126, "recommended_fix": "Move the filtering logic to a configuration-driven approach or allow the LLM to filter relevant rules based on the full set of provided guidelines.", "severity": "P1", "why_problematic": "The logic filters visual rules based on a hardcoded list of genre-specific semantic categories ('possession', 'projection', 'ghost'). This makes the pipeline dependent on specific story domains and prevents it from being a general-purpose director agent."}, {"category": "scenario_dependent_prompt", "evidence": "visual_rules = \"시각적 존재 판단 참고사항 (단, 카메라에 보이면 무조건 포함):\\n\" + \"\\n\".join(f\"- {n}\" for n in notes)", "line_end": 156, "line_start": 121, "recommended_fix": "Relocate these hardcoded strings to the prompt template system and pass the content as structured variables to the LLM client.", "severity": "P2", "why_problematic": "Prompt instructions and section headers (e.g., '[Shot 분석 참고]', '## 기획서: 전체 줄거리') are hardcoded in Korean within the Python StepRunner. This scatters prompt logic across the codebase and hinders localization and template management."}], "path": "backend/app/core/steps/director_steps.py", "scan_kind": "python", "sha256": "4832353f84c57b443d43af643af033dcd71d175ad607caeb2853e389bd1c3265"}
{"candidate_reason": "python scope discovery", "chunk_end": 537, "chunk_start": 1, "chunk_summary": "The FloorPlanRenderStep implementation is clean, focusing on ID-based orchestration, status validation, and file path management without performing open-world semantic string judgment.", "duration_ms": 9079, "findings": [], "path": "backend/app/core/steps/floor_plan_render_step.py", "scan_kind": "python", "sha256": "95bc28000e5f70e5ac3f3830e7679e68d6c937b7ce401cc96a21bafcc17a6b06"}
{"candidate_reason": "python scope discovery", "chunk_end": 561, "chunk_start": 1, "chunk_summary": "The file contains hardcoded narrative tropes for logic branching and significant amounts of hardcoded Korean prompt instructions within the Python step logic.", "duration_ms": 36596, "findings": [{"category": "scenario_dependent_code", "evidence": "r.get(\"rule_type\") in (\"possession\", \"projection\", \"ghost\")", "line_end": 115, "line_start": 114, "recommended_fix": "Move trope-based filtering to a configuration file or use a generic attribute (e.g., 'is_visual_guideline') in the rule schema.", "severity": "P1", "why_problematic": "The pipeline logic is hardcoded to specific narrative tropes (possession, ghost) to filter visual rules. This prevents the system from being genre-agnostic and scatters domain nomenclature in the core logic."}, {"category": "scenario_dependent_code", "evidence": "r.get(\"rule_type\") in (\"possession\", \"projection\", \"ghost\")", "line_end": 324, "line_start": 323, "recommended_fix": "Abstract the rule filtering logic into a shared method that uses metadata flags rather than string-matching on tropes.", "severity": "P1", "why_problematic": "Duplicate of the trope-based filtering logic in the shot extraction step, coupling core logic to specific supernatural story elements."}, {"category": "scenario_dependent_prompt", "evidence": "\"[물리적 존재 판단 기준 — ... ]\"", "line_end": 462, "line_start": 110, "recommended_fix": "Move all hardcoded prompt strings into the prompt template files (system/user) or a dedicated localization/config SOT, passing them as variables to the template.", "severity": "P1", "why_problematic": "Multiple natural language prompt instructions, headers, and state descriptions (e.g., lines 110, 117, 124, 210, 318, 326, 335, 437, 454, 462) are hardcoded in Korean within the Python logic. This pollutes the code with scenario-specific prose and makes prompt maintenance or localization difficult."}], "path": "backend/app/core/steps/beat_shot_steps.py", "scan_kind": "python", "sha256": "c25b9c5afa08d9056c9af9387bfc34788f5960423bb426df0b7445ce8a2d8321"}
{"candidate_reason": "python scope discovery", "chunk_end": 3146, "chunk_start": 1, "chunk_summary": "The file contains logic for entity variant grouping based on name parsing, heuristic grammar fixes on LLM output via blind regex, and developer meta-commentary within production prompts.", "duration_ms": 47156, "findings": [{"category": "semantic_string_judgment", "evidence": "re.split(r'\\s*[\\(（]', name)[0].strip()", "line_end": 1940, "line_start": 1939, "recommended_fix": "Use a structured SOT field (e.g., base_entity_id) to link variants instead of parsing name strings.", "severity": "P1", "why_problematic": "Determines entity variant relationships by parsing parentheses in character names, which is open-world scenario text. This assumes a specific naming convention to decide identity logic."}, {"category": "blind_string_mutation", "evidence": "re.sub(r\"focus on\\s+'s\", \"focus on the figure's\", prompt)", "line_end": 2869, "line_start": 2867, "recommended_fix": "Improve the system prompt to ensure correct grammar or use a structured validator that doesn't rely on blind regex replacement.", "severity": "P1", "why_problematic": "Heuristically mutates open-world LLM output (t2i_prompt) based on substring patterns to fix grammar. This is a blind mutation that can lead to unintended semantic changes in the generated prompt."}, {"category": "scenario_dependent_prompt", "evidence": "production code / prompt 어디에도 hardcoded ethnicity 어휘 list 를 강요하지 마세요 (시나리오/region 마다 다름, A-prime binding)", "line_end": 2268, "line_start": 2268, "recommended_fix": "Remove developer notes and meta-justifications from the prompt text.", "severity": "P2", "why_problematic": "Contains meta-commentary and developer notes about system design and 'A-prime binding' philosophy inside the production prompt sent to the LLM."}], "path": "backend/app/core/steps/detail_steps.py", "scan_kind": "python", "sha256": "13d5494eabda57fd91fc817a7b5df837146b46e1ba33ed168da309d019eb88ed"}
{"candidate_reason": "python scope discovery", "chunk_end": 978, "chunk_start": 1, "chunk_summary": "The file uses string patterns within the 'prompt_used' field to manage image asset sub-types and character states, and employs hardcoded keyword matching to detect semantic character conditions from LLM-generated text.", "duration_ms": 22373, "findings": [{"category": "semantic_string_judgment", "evidence": "COALESCE(prompt_used, '') NOT LIKE '[outfit:%' AND COALESCE(prompt_used, '') NOT LIKE '[composite:%'", "line_end": 148, "line_start": 147, "recommended_fix": "Add a structured 'asset_sub_type' column or a dedicated metadata table to ImageAsset to track whether an image is a base reference, outfit, or composite.", "severity": "P1", "why_problematic": "The system distinguishes between base references and composite/outfit assets by checking for string prefixes in a text field intended for the prompt, rather than using a structured sub-type or relationship."}, {"category": "semantic_string_judgment", "evidence": "ImageAsset.prompt_used.like(\"%composite:%\") ... re.search(r'composite:([a-f0-9-]+):([a-f0-9-]+)', ca.prompt_used or \"\")", "line_end": 331, "line_start": 301, "recommended_fix": "Store character_id and outlook_id in structured columns or a junction table for composite assets.", "severity": "P1", "why_problematic": "Composite asset identification and entity ID extraction rely on parsing a string pattern ('composite:{char_id}:{outlook_id}') inside the prompt field. This is fragile and bypasses database integrity."}, {"category": "blind_string_mutation", "evidence": "ca.prompt_used = ca.prompt_used.replace(\"composite:\", \"composite_old:\")", "line_end": 307, "line_start": 307, "recommended_fix": "Use a 'status' or 'is_active' boolean column to handle asset invalidation/versioning.", "severity": "P1", "why_problematic": "Invalidating assets by prepending 'old' to a string prefix in a text field is a blind mutation that makes state management dependent on string manipulation."}, {"category": "semantic_string_judgment", "evidence": "if gaze in self.STATE_DESCRIPTIONS:", "line_end": 666, "line_start": 666, "recommended_fix": "The upstream 'shot_staging' step should output a structured 'character_state' enum rather than relying on the 'gaze_target' text field for this logic.", "severity": "P1", "why_problematic": "The code judges the semantic state of a character (dead, injured, etc.) by checking if the 'gaze_target' string (generated by an LLM in a previous step) matches a closed list of keywords. This is an open-world semantic judgment via string matching."}, {"category": "blind_string_mutation", "evidence": "ov.prompt_used.replace(\"state_variant:\", \"state_variant_old:\")", "line_end": 711, "line_start": 707, "recommended_fix": "Implement a formal versioning or status system for ImageAssets.", "severity": "P1", "why_problematic": "Similar to the composite step, state variants are invalidated via blind string replacement in the prompt field."}, {"category": "semantic_string_judgment", "evidence": "ImageAsset.prompt_used.like(\"%state_variant:%\") ... _re.search(r'state_variant:([a-f0-9-]+):(\\w+)', r.prompt_used or \"\")", "line_end": 946, "line_start": 942, "recommended_fix": "Use structured columns for entity_id and state_type on the ImageAsset model.", "severity": "P1", "why_problematic": "Verification of character state variants depends on parsing metadata (UUIDs and state types) embedded in the prompt string."}], "path": "backend/app/core/steps/image_steps.py", "scan_kind": "python", "sha256": "a1404a37cdd9b2ef5529d5936f0663a5a51e07a44f5c6ed057442bd14ce2491c"}
{"candidate_reason": "python scope discovery", "chunk_end": 315, "chunk_start": 1, "chunk_summary": "The step uses substring matching for location-to-scene mapping and pattern-based key lookups for entity details, along with semantic filtering instructions in the LLM prompt.", "duration_ms": 24505, "findings": [{"category": "semantic_string_judgment", "evidence": "if name in heading:", "line_end": 105, "line_start": 104, "recommended_fix": "Rely exclusively on the ID-based mapping (present_entity_ids) from the scene_director step. If a fallback is needed, it should be performed by an LLM with the full context or via a strict ID registry.", "severity": "P1", "why_problematic": "Uses a substring check to associate a location with a scene based on the scene's heading text. This is an open-world semantic judgment that can lead to incorrect mappings if location names are substrings of each other (e.g., 'Lab' matching 'Laboratory')."}, {"category": "scenario_dependent_code", "evidence": "detail_key = f\"{name}:location\" ... k.startswith(f\"{name}:\")", "line_end": 226, "line_start": 217, "recommended_fix": "Use a structured ID (e.g., short_id) as the primary key for entity lookups instead of relying on name-based string patterns.", "severity": "P1", "why_problematic": "The code attempts to resolve entity details by constructing keys with a specific string format ('name:type') and using startswith as a fallback. This depends on the specific way a previous LLM step formatted its output keys, making the pipeline brittle to changes in LLM behavior."}, {"category": "semantic_string_judgment", "evidence": ".startswith(\"실패\")", "line_end": 179, "line_start": 122, "recommended_fix": "Introduce a structured 'status' field (e.g., 'SUCCESS', 'FAILED') in the schema and use it for logic branching instead of parsing the summary string.", "severity": "P2", "why_problematic": "Control flow (determining success/failure) is decided by checking for a specific string prefix in a human-readable summary field ('analysis_summary'). This mixes status signaling with descriptive text."}, {"category": "llm_closed_list_instruction", "evidence": "\"- 포함: 크기·형태·재질... - 제외: 날씨·시간·조명...\"", "line_end": 266, "line_start": 259, "recommended_fix": "Move the filtering logic to a more robust semantic analysis step or provide a more generalized definition of 'fixed visual traits' that doesn't rely on a closed list of exclusions.", "severity": "P2", "why_problematic": "The prompt instructs the LLM to perform semantic classification and filtering based on a hardcoded list of categories. This is scenario-dependent and may not cover all edge cases in open-world scene descriptions."}], "path": "backend/app/core/steps/location_consistency_step.py", "scan_kind": "python", "sha256": "6dadf83395598ee2fa477e1a5d0b06313b910f0a9baf0707c6dafd727dde1dab"}
{"candidate_reason": "python scope discovery", "chunk_end": 217, "chunk_start": 1, "chunk_summary": "The step uses hardcoded string keys and length heuristics to judge document availability and section presence, creating schema drift and semantic judgment risks.", "duration_ms": 20572, "findings": [{"category": "schema_or_enum_drift", "evidence": "computed_available = []\\n        if result.get(\"characters\"):\\n            computed_available.append(\"characters\")\\n        if result.get(\"world_setting\", \"\").strip():", "line_end": 208, "line_start": 195, "recommended_fix": "Derive 'available_sections' dynamically by iterating over the keys in _ANALYSIS_SCHEMA or trust the LLM's output as requested in the prompt.", "severity": "P1", "why_problematic": "The code manually recalculates 'available_sections' by checking for non-empty strings against a hardcoded list of keys. This duplicates the keys defined in _ANALYSIS_SCHEMA and overrides the LLM's own determination, leading to maintenance debt if the schema evolves."}, {"category": "semantic_string_judgment", "evidence": "if not has_pdf and (not planning_text or len(planning_text.strip()) < 100):", "line_end": 139, "line_start": 139, "recommended_fix": "Use a more robust check for document validity or allow the LLM to attempt analysis and return an empty result if the content is insufficient.", "severity": "P2", "why_problematic": "Uses a magic number (100) and string length to decide if a planning document is semantically 'present' or 'valid' for analysis. This is a heuristic judgment on open-world content."}, {"category": "llm_closed_list_instruction", "evidence": "role: 역할 (주인공, 조연, 악역 등) ... 관계 (가족, 연인, 적대 등)", "line_end": 89, "line_start": 74, "recommended_fix": "Remove specific examples or move them to a separate 'guidelines' section of the SOT to avoid biasing the extraction process.", "severity": "P2", "why_problematic": "The prompt provides specific domain-specific examples to guide LLM extraction. While not a strict closed list, it biases the LLM towards these specific categories for open-world character roles and relationships instead of allowing pure extraction from the source."}], "path": "backend/app/core/steps/planning_doc_step.py", "scan_kind": "python", "sha256": "30736eedb414dddb5c432e9cde26f18ac1a90b52404ba4abdd013f8df3485131"}
{"candidate_reason": "python scope discovery", "chunk_end": 553, "chunk_start": 1, "chunk_summary": "no actionable findings", "duration_ms": 29913, "findings": [], "path": "backend/app/core/steps/outlook_steps.py", "scan_kind": "python", "sha256": "7cff875a4cc7dd2b8d2a3bc0961d812f257d97eba72e149a570f06b4e35548d2"}
{"candidate_reason": "python scope discovery", "chunk_end": 781, "chunk_start": 1, "chunk_summary": "The audit identified heuristic-based semantic summary extraction from LLM output and blind string mutation on open-world text used for prompt construction.", "duration_ms": 33942, "findings": [{"category": "semantic_string_judgment", "evidence": "_summarize_prompt", "line_end": 734, "line_start": 714, "recommended_fix": "Modify the upstream prompt generation (generate_floor_plan_prompt) to return a structured JSON containing both the full prompt and an explicit summary field, rather than parsing the raw string.", "severity": "P1", "why_problematic": "Uses string patterns (find('. '), find('.\\n')) and a magic length limit (200) to extract a semantic summary from open-world LLM output. This heuristic is fragile and attempts to derive structured meaning from unstructured text to be used as context for subsequent LLM calls."}, {"category": "blind_string_mutation", "evidence": "cleaned = summary.replace(\"\\n\", \" \").strip()", "line_end": 746, "line_start": 746, "recommended_fix": "Use a structured prompt template or ensure the summary is generated as a single line by the LLM.", "severity": "P2", "why_problematic": "Blindly replaces newlines in open-world summary text to force it into a single-line prompt format. This can corrupt intended formatting or semantic breaks in the summary text."}], "path": "backend/app/core/steps/location_floor_plan_step.py", "scan_kind": "python", "sha256": "86bc164f3624c24cfe7edc1704db823d4f3d9122c7f53fd39106b85d11dcb146"}
{"candidate_reason": "python scope discovery", "chunk_end": 242, "chunk_start": 1, "chunk_summary": "The code contains hardcoded Korean strings in prompt construction and output data, which introduces scenario-dependent pollution and risks downstream semantic string judgment.", "duration_ms": 29220, "findings": [{"category": "scenario_dependent_code", "evidence": "\"flow_summary\": \"선택된 샷 없음 — 플로우 없음\"", "line_end": 126, "line_start": 126, "recommended_fix": "Use a structured status field (e.g., 'status': 'no_shots') and move the display string to a localization layer or the UI.", "severity": "P2", "why_problematic": "Hardcoding natural language status messages in the data block instead of using structured status codes or booleans. This pollutes the checkpoint and risks downstream logic relying on string matching."}, {"category": "scenario_dependent_prompt", "evidence": "\"(없음)\", \"(엔티티 정보 없음)\"", "line_end": 196, "line_start": 187, "recommended_fix": "Move these fallback strings into the prompt template (e.g., using Jinja2 default filters) or handle them within the prompt loader.", "severity": "P2", "why_problematic": "Hardcoding language-specific fallback strings in the Python logic for prompt construction. This makes the pipeline dependent on specific language/scenario context and harder to maintain or localize."}, {"category": "scenario_dependent_code", "evidence": "\"flow_summary\": f\"분석 실패: {exc}\"", "line_end": 239, "line_start": 239, "recommended_fix": "Use a dedicated 'error' field and a structured error code, keeping 'flow_summary' for successful semantic summaries.", "severity": "P2", "why_problematic": "Storing raw exception messages in a data field intended for summary. This can lead to inconsistent data shapes and encourages downstream string-based error handling."}], "path": "backend/app/core/steps/scene_camera_flow_step.py", "scan_kind": "python", "sha256": "657e39ab6fa5b6c7e0a3f31c562ee529d72342fdf1be7e23ac81d058dbcef351"}
{"candidate_reason": "python scope discovery", "chunk_end": 3657, "chunk_start": 1, "chunk_summary": "The file contains multiple instances of open-world semantic judgment using hardcoded phrase lists and token patterns, including bilingual detection lists for spatial and continuity rules.", "duration_ms": 34891, "findings": [{"category": "llm_closed_list_instruction", "evidence": "_ID_BODY_PART_TRIGGERS, _ID_CLOSE_FACE_FORBIDDEN_PHRASES, _ID_ETHNICITY_COMPONENTS, _ID_AGE_BANDS", "line_end": 190, "line_start": 114, "recommended_fix": "Move these descriptors and trigger phrases to a structured Source of Truth (SOT) or allow the LLM to determine semantic focus based on high-level principles rather than exact phrase matches.", "severity": "P1", "why_problematic": "These constants provide closed lists of phrases and descriptors (e.g., 'focus on', 'face filling the entire frame', ethnicity lists) to the LLM to classify or generate open-world semantic content. This restricts the LLM's ability to handle nuanced or novel scenario text."}, {"category": "semantic_string_judgment", "evidence": "_CONTINUITY_GENERIC_PERSON_NOUNS, _CONTINUITY_FURNITURE_LAYOUT_TOKENS", "line_end": 231, "line_start": 207, "recommended_fix": "Replace regex/substring detection with LLM-based semantic validation or use a more robust NLP approach that doesn't rely on fixed token lists.", "severity": "P1", "why_problematic": "These lists are used as 'ground-truth' for detecting semantic violations (double-description, layout imports) in open-world text. Using hardcoded phrase lists for detection is brittle and fails to capture semantic equivalents."}, {"category": "semantic_string_judgment", "evidence": "_SPATIAL_CAMERA_LOW_TOKENS, _SPATIAL_CAMERA_HIGH_TOKENS, _SPATIAL_INTERACTION_VERBS, _SPATIAL_SHARED_ANCHOR_KEYWORDS", "line_end": 408, "line_start": 336, "recommended_fix": "Use structured metadata from the staging/shot_info objects to determine spatial properties instead of parsing natural language descriptions with token lists.", "severity": "P1", "why_problematic": "These constants contain bilingual (English/Korean) token lists used to detect camera positions and spatial interactions in open-world shot descriptions. This is a direct violation of the ban on using string patterns to judge open-world semantic meaning."}, {"category": "semantic_string_judgment", "evidence": "if (perception_mode or \"\").lower() in (\"reflection\", \"mirror\", \"through_device\", \"projection\")", "line_end": 1222, "line_start": 1220, "recommended_fix": "Ensure perception_mode is a strict enum validated at the schema level, or use a structured flag in the input metadata to trigger this logic.", "severity": "P1", "why_problematic": "The code uses a substring/membership check on a potentially open-world string ('perception_mode') to decide whether to append specific constraints to the ID policy. This is a pattern-based semantic judgment."}, {"category": "llm_closed_list_instruction", "evidence": "\"do not use any of: 'the existing X', 'from the reference', ... 'matching the reference camera'\"", "line_end": 1335, "line_start": 1326, "recommended_fix": "Provide high-level semantic constraints (e.g., 'do not refer to the background reference image') rather than specific forbidden substrings.", "severity": "P1", "why_problematic": "The prompt instructions explicitly list forbidden phrases for the LLM based on a closed list of examples. This forces the LLM to perform string-level classification of its own output against a fixed list of phrases."}], "path": "backend/app/core/steps/render_prompt_card.py", "scan_kind": "python", "sha256": "d7960bcbc097897cfdc1205fc20d0a448feb49c786901023e7461f33438f221d"}
{"candidate_reason": "python scope discovery", "chunk_end": 779, "chunk_start": 1, "chunk_summary": "The file implements semantic validation of LLM outputs using regex-based keyword matching on descriptions and element IDs to infer framing types, and contains scenario-specific examples in the prompt instructions.", "duration_ms": 31269, "findings": [{"category": "semantic_string_judgment", "evidence": "_ELEMENT_ID_CLOSE_REGEX, _DESCRIPTION_CLOSE_KEYWORDS, and _classify_framing", "line_end": 121, "line_start": 65, "recommended_fix": "Modify the LLM response schema to include an explicit 'framing_type' enum field. The LLM should categorize the framing during generation rather than the code attempting to infer it from text.", "severity": "P0", "why_problematic": "The code determines the semantic 'framing' (close-up vs full-body) of a scene element by matching open-world descriptions and element IDs against a hardcoded list of body parts and photography terms. This judgment is used to fail validation and block downstream processing via STATUS_VALIDATOR_VIOLATIONS."}, {"category": "scenario_dependent_prompt", "evidence": "사망/부상/의식불명 인물의 자세와 위치... 깨진 창문, 열린 문, 혈흔 등", "line_end": 751, "line_start": 749, "recommended_fix": "Replace specific examples with abstract categories (e.g., 'static environmental states', 'character physical conditions') or move them to a world-specific configuration file.", "severity": "P1", "why_problematic": "The prompt uses specific scenario-dependent examples (e.g., 'bloodstains', 'broken windows', 'unconscious people') to define what the LLM should look for. This couples the pipeline logic to specific genres or story types."}, {"category": "semantic_string_judgment", "evidence": "summary.startswith(\"분석 실패\") or summary.startswith(\"분석 차단\")", "line_end": 356, "line_start": 355, "recommended_fix": "Ensure all checkpoints are migrated to use the explicit 'status' field constants, and remove the string-prefix fallback logic.", "severity": "P2", "why_problematic": "The code relies on localized string prefixes in the 'analysis_summary' field to determine the success or failure status of previous runs for backward compatibility. This is a fragile way to handle state compared to structured status enums."}], "path": "backend/app/core/steps/scene_consistency_step.py", "scan_kind": "python", "sha256": "a07ea382bb71f5f22aa07c2f67ccd3ca32b598675d75e6e1b8731e0664ec9267"}
{"candidate_reason": "python scope discovery", "chunk_end": 889, "chunk_start": 1, "chunk_summary": "The SceneContextLoader centralizes checkpoint loading but contains a brittle semantic check using a Korean string prefix to filter location visuals.", "duration_ms": 31161, "findings": [{"category": "semantic_string_judgment", "evidence": "if summary.startswith(\"실패\"):", "line_end": 266, "line_start": 266, "recommended_fix": "Update the location_consistency checkpoint schema to include a structured status field and use it for filtering instead of the summary text.", "severity": "P1", "why_problematic": "The code uses a hardcoded Korean string prefix ('실패', meaning 'failure') to determine if a location analysis failed. This is an open-world semantic judgment on a natural language field, which is fragile and should be replaced by a structured status enum."}], "path": "backend/app/core/steps/scene_context_loader.py", "scan_kind": "python", "sha256": "1cd6a7d7b27f853057101287ef0c718d2337d8e7e89a3f65096e6395b01ed410"}
{"candidate_reason": "python scope discovery", "chunk_end": 162, "chunk_start": 1, "chunk_summary": "The file contains hardcoded prompt instructions and labels in Korean, which should be moved to the prompt template system to avoid scenario-dependent logic in the code.", "duration_ms": 21178, "findings": [{"category": "scenario_dependent_prompt", "evidence": "user_prompt += (\\n[다양성 규칙] 같은 씬의 연속 shot에 동일한 촬영 기법을 배정하지 마세요. ...)", "line_end": 134, "line_start": 131, "recommended_fix": "Move the diversity rule instruction into the 'shot_cinematography' user prompt template stored in the database.", "severity": "P1", "why_problematic": "Hardcoding behavioral constraints and cinematography-specific rules directly in the Python logic instead of the prompt template. This splits the prompt definition and makes the pipeline rigid and harder to localize or tune."}, {"category": "scenario_dependent_prompt", "evidence": "beat = f\"beat:{sh['based_on_beat']}\" if sh[\"based_on_beat\"] else \"원문\"", "line_end": 120, "line_start": 120, "recommended_fix": "Pass the label as a variable to the prompt template or use a configuration-driven constant for the 'original text' label.", "severity": "P2", "why_problematic": "Hardcoded Korean label '원문' (Original Text) is injected into the LLM prompt. This creates a language and scenario dependency in the code that should be handled by the template or a configuration constant."}], "path": "backend/app/core/steps/shot_cinematography_step.py", "scan_kind": "python", "sha256": "553ac84f1a94ced477d79e1c148187e782717b8332a723c27413affb05b1d128"}
{"candidate_reason": "python scope discovery", "chunk_end": 312, "chunk_start": 1, "chunk_summary": "The scene segmentation logic contains scenario-dependent prompt instructions and hardcoded formatting assumptions used to validate and retry LLM-generated regex patterns.", "duration_ms": 31359, "findings": [{"category": "scenario_dependent_prompt", "evidence": "PDF에서 추출한 텍스트라 줄바꿈 없이 본문과 씬 헤딩이 붙어있을 수 있습니다.", "line_end": 83, "line_start": 83, "recommended_fix": "Move source-specific context (like PDF artifact handling) to a project-level configuration or a pre-processing step rather than hardcoding it in the core step prompt.", "severity": "P2", "why_problematic": "The prompt hardcodes assumptions about the source material (PDF extraction artifacts), which may not apply to all input types and pollutes the general segmentation logic."}, {"category": "llm_closed_list_instruction", "evidence": "씬 내부의 장소 전환('- 장소명')이 아니라 씬 번호('숫자.') 패턴으로 분리해야 합니다.", "line_end": 131, "line_start": 131, "recommended_fix": "Instead of hardcoding specific patterns, provide the LLM with a few examples from the actual 'cleaned_text' or allow the user to define the expected heading style in ProjectSettings.", "severity": "P1", "why_problematic": "This retry instruction forces the LLM to look for specific string patterns ('숫자.') and avoid others ('- 장소명'), which is scenario-dependent and prevents the system from correctly parsing scripts that use different heading conventions."}, {"category": "semantic_string_judgment", "evidence": "if avg_len < min_avg:", "line_end": 134, "line_start": 127, "recommended_fix": "Use the LLM to validate the segmentation quality or make the threshold a soft warning rather than a hard routing trigger.", "severity": "P1", "why_problematic": "Uses a blind character-length heuristic to invalidate and retry semantic segmentation. This can cause infinite retries or failures for scripts with legitimate short scenes (e.g., montage or rapid cuts)."}], "path": "backend/app/core/steps/scene_steps.py", "scan_kind": "python", "sha256": "3aa185bd3a049018d749c91b437e39ae0ace024d800f64d0e2fab32a510d9015"}
{"candidate_reason": "python scope discovery", "chunk_end": 46, "chunk_start": 1, "chunk_summary": "No actionable findings; this chunk is a standard StepRunner orchestrator passing data between pipeline modules without semantic string judgment.", "duration_ms": 3815, "findings": [], "path": "backend/app/core/steps/shot_staging_step.py", "scan_kind": "python", "sha256": "e86b746584f044e566326671dc8be66843ca343ea2d8b7e09356a77dcf07d935"}
{"candidate_reason": "python scope discovery", "chunk_end": 126, "chunk_start": 1, "chunk_summary": "The file defines a StepRunner for shot direction that orchestrates data between checkpoints and delegates semantic logic to external modules and LLM prompts, with no local string-based semantic judgment.", "duration_ms": 13111, "findings": [], "path": "backend/app/core/steps/shot_director_step.py", "scan_kind": "python", "sha256": "9a337c33f08a7b50b1b2344fde61ef56d39e90922639effa8e06e34012114abe"}
{"candidate_reason": "python scope discovery", "chunk_end": 302, "chunk_start": 1, "chunk_summary": "The file implements a shot dependency analysis step using an LLM, but contains hardcoded prompt instructions and unstable schema access patterns.", "duration_ms": 26200, "findings": [{"category": "scenario_dependent_prompt", "evidence": "f\"장소: {loc_name} ({loc_id})\\n\\n\" ... f\"  Description: {s['description']}\\n\"", "line_end": 265, "line_start": 248, "recommended_fix": "Move the entire user prompt construction to a Jinja2 template loaded via load_prompt, passing the shots and location info as a context dictionary.", "severity": "P1", "why_problematic": "Prompt instructions, labels (Korean and English), and shot data formatting are hardcoded directly in the Python logic. This prevents clean separation of prompt engineering from code and complicates localization or prompt iteration."}, {"category": "schema_or_enum_drift", "evidence": "detail_cp.get(\"data\", detail_cp).get(\"scenes\", detail_cp.get(\"scenes\", []))", "line_end": 187, "line_start": 187, "recommended_fix": "Enforce a strict schema for the 'scene_detail' checkpoint and use a single, deterministic path (e.g., detail_cp['data']['scenes']).", "severity": "P2", "why_problematic": "The redundant and defensive access pattern for 'scenes' suggests that the checkpoint schema for 'scene_detail' is inconsistent or has drifted, leading to fragile code that tries to guess the data location."}], "path": "backend/app/core/steps/shot_dependency_t2i_step.py", "scan_kind": "python", "sha256": "4e5bf81c1d9040583870f1e8f14636e453dad73093f557bbbaf9b1511ec9cc8a"}
{"candidate_reason": "python scope discovery", "chunk_end": 180, "chunk_start": 1, "chunk_summary": "The code performs shot dependency analysis using entity overlap scoring based on structured IDs and scene locations, with no actionable findings regarding open-world semantic string judgment or scenario pollution.", "duration_ms": 33816, "findings": [], "path": "backend/app/core/steps/shot_dependency_step.py", "scan_kind": "python", "sha256": "6db76375617c27cd5402dc6c8b9a6177fc84261bf73fe26f54e2293104acf226"}
{"candidate_reason": "python scope discovery", "chunk_end": 54, "chunk_start": 1, "chunk_summary": "No actionable findings; the file contains standard orchestration logic for text extraction without semantic string judgment or scenario-dependent pollution.", "duration_ms": 4463, "findings": [], "path": "backend/app/core/steps/text_steps.py", "scan_kind": "python", "sha256": "dee0d749567714398fc791817e5732c769049f44dd42787e16593f185023fd08"}
{"candidate_reason": "python scope discovery", "chunk_end": 40, "chunk_start": 1, "chunk_summary": "no actionable findings", "duration_ms": 3850, "findings": [], "path": "backend/app/core/task_registry.py", "scan_kind": "python", "sha256": "ff6afe0ab1644dfaa038b42db7fadb121b972c02662ff83b685bfd4aa28ff348"}
{"candidate_reason": "python scope discovery", "chunk_end": 258, "chunk_start": 1, "chunk_summary": "no actionable findings", "duration_ms": 14266, "findings": [], "path": "backend/app/core/steps/t2i_review_step.py", "scan_kind": "python", "sha256": "7f44d1520245d3690182a66bd7e2d6892be7b59cc8939e0b4c85abeb75ae6baf"}
{"candidate_reason": "python scope discovery", "chunk_end": 309, "chunk_start": 1, "chunk_summary": "The file implements a step to extract essence, peripheral, and atmospheric elements from shot descriptions using an LLM, but contains hardcoded prompt instructions and domain nomenclature in the Python logic.", "duration_ms": 35366, "findings": [{"category": "llm_closed_list_instruction", "evidence": "user_prompt = ( \"아래 샷들의 description을 essence/peripheral/atmospheric 3분류로 나누세요.\\n\" \"각 샷마다 scene_index/shot_index/essence/peripheral/atmospheric 모두 출력.\\n\" \"\\n[분석 대상 샷]\" + \"\\n\".join(shots_lines) )", "line_end": 195, "line_start": 190, "recommended_fix": "Move the instruction text to a template in the prompt library and use load_prompt or a template formatter to inject the shot data. Ensure domain nomenclature is consistent with the loaded schema.", "severity": "P1", "why_problematic": "The user prompt contains hardcoded task instructions and domain-specific nomenclature ('essence', 'peripheral', 'atmospheric') in Korean. This couples the runner logic to a specific language and task definition that should be managed in the external prompt library (SOT) rather than scattered in code."}], "path": "backend/app/core/steps/shot_essence_extraction_step.py", "scan_kind": "python", "sha256": "bd3cf91c9b396438817a9061f82d5b2f1d5414d825f17ae6641c171361bf1820"}
{"candidate_reason": "python scope discovery", "chunk_end": 287, "chunk_start": 1, "chunk_summary": "The code contains scenario-dependent heuristics for shot truncation and hardcoded Korean/English strings for selection metadata.", "duration_ms": 39634, "findings": [{"category": "scenario_dependent_code", "evidence": "r[\"selected_shots\"][:quota], max_r[\"selected_shots\"].pop(), deduped[:effective_max]", "line_end": 248, "line_start": 177, "recommended_fix": "Ensure the LLM provides an importance ranking for all shots, and use that ranking to perform truncation when budget limits are exceeded.", "severity": "P1", "why_problematic": "The code uses blind list slicing and popping to enforce shot count limits (both per-scene and per-episode). This assumes that list order correlates perfectly with narrative importance, which is a scenario-dependent heuristic that overrides the LLM's semantic selection without narrative context."}, {"category": "scenario_dependent_code", "evidence": "\"reason\": \"SHOT_SELECTION_ENABLED=false\", \"reason\": \"씬 내 유일 샷\", \"reason\": \"fallback: LLM 결과 없음\"", "line_end": 257, "line_start": 73, "recommended_fix": "Use a structured SOT or a constants file to manage selection reasons and status messages.", "severity": "P2", "why_problematic": "Hardcoded strings (including Korean text) are used as selection reasons within the logic. This introduces scenario-dependent pollution into the codebase and makes the system harder to maintain or localize."}], "path": "backend/app/core/steps/shot_selection_step.py", "scan_kind": "python", "sha256": "0b3d7f87247c5760f000b0ed7d6031a2ed5722df3d4e2c087fc2ef055d16a254"}
{"candidate_reason": "python scope discovery", "chunk_end": 197, "chunk_start": 1, "chunk_summary": "The file defines StepRunners for the summary phase, coordinating LLM modules and checkpoint persistence, with some hardcoded prompt headers in the logic.", "duration_ms": 29421, "findings": [{"category": "scenario_dependent_prompt", "evidence": "pctx.inject_if_available(\"world_setting\", \"## 기획서: 세계관\") + pctx.inject_if_available(\"tone_mood\", \"## 기획서: 톤/분위기\") + pctx.inject_if_available(\"visual_concepts\", \"## 기획서: 비주얼 컨셉\")", "line_end": 100, "line_start": 98, "recommended_fix": "Move these headers into the get_planning_context helper or use a structured template/SOT to define how planning document sections are labeled in prompts.", "severity": "P2", "why_problematic": "Hardcoded Korean headers for prompt injection are scattered domain nomenclature. This couples the StepRunner logic to specific prompt structure and language, making it harder to maintain or localize."}], "path": "backend/app/core/steps/summary_steps.py", "scan_kind": "python", "sha256": "68cadacda2b891546478fce5e24ebdc5117e38140cbe27bd27070780fcb4a949"}
{"candidate_reason": "python scope discovery", "chunk_end": 20, "chunk_start": 1, "chunk_summary": "No actionable findings; the file implements a standard i18n string loader using file-system locales and JSON templates.", "duration_ms": 8174, "findings": [], "path": "backend/app/i18n/loader.py", "scan_kind": "python", "sha256": "1df58ae74fe67baf31a49079fd1cbddf2dabf03081e7ba321cd802e024319af9"}
{"candidate_reason": "python scope discovery", "chunk_end": 42, "chunk_start": 1, "chunk_summary": "No actionable findings; the activity logger performs standard structured data persistence without semantic string judgment or scenario-dependent logic.", "duration_ms": 3102, "findings": [], "path": "backend/app/logging/activity_logger.py", "scan_kind": "python", "sha256": "997d9bc02147260f7d5770a934833ee80b3a205db9e22867aa632dd587764d10"}
{"candidate_reason": "python scope discovery", "chunk_end": 16, "chunk_start": 1, "chunk_summary": "The file defines a standard SQLAlchemy model for activity logging with no actionable findings regarding semantic string judgment or scenario pollution.", "duration_ms": 4152, "findings": [], "path": "backend/app/logging/models.py", "scan_kind": "python", "sha256": "5e479a106bc17dc09334a3b3871991ec6787e18bdf9d372fe805dc4c9a19d6c8"}
{"candidate_reason": "python scope discovery", "chunk_end": 51, "chunk_start": 1, "chunk_summary": "No actionable findings; the file contains standard SQLAlchemy model definitions with closed-world system constants.", "duration_ms": 3007, "findings": [], "path": "backend/app/models/catalog.py", "scan_kind": "python", "sha256": "9d4a647c9247795de8bc255e2c91bbe888ea3fb924797258c1032c869d7f5ac7"}
{"candidate_reason": "python scope discovery", "chunk_end": 130, "chunk_start": 1, "chunk_summary": "The file contains standard FastAPI application setup, lifespan management, and health check logic with no actionable findings regarding semantic string judgment or scenario pollution.", "duration_ms": 5129, "findings": [], "path": "backend/app/main.py", "scan_kind": "python", "sha256": "49472efd68728a06703b1264499b0c8c0dc8dc57e7b020a1fec14e155b4587bc"}
{"candidate_reason": "python scope discovery", "chunk_end": 524, "chunk_start": 1, "chunk_summary": "The shot validator step uses substring checks on exception messages for retry routing and contains hardcoded domain-specific prompt instructions in the Python code.", "duration_ms": 37396, "findings": [{"category": "semantic_string_judgment", "evidence": "if any(k in msg for k in (\"content_filter\", \"content filter\", \"safety\", \"recitation\", \"empty response\")):", "line_end": 345, "line_start": 340, "recommended_fix": "Use specific exception classes or status codes provided by the LLM client SDK, or centralize error classification into a dedicated handler that maps provider-specific errors to internal status enums.", "severity": "P2", "why_problematic": "The code judges the semantic nature of an LLM provider error (e.g., whether it is a safety block or a transient error) using substring patterns to decide retry routing. This is fragile as error messages are open-world strings that can change across providers or versions, and it includes a check for 'empty response' which is an internally raised string from line 330."}, {"category": "scenario_dependent_prompt", "evidence": "f\"각 shot의 description이 '한 찰나' 원칙에 맞는지 검증하고, 위반 시 재작성하세요. character_ids 는 entity map 의 short_id 사용.\"", "line_end": 289, "line_start": 287, "recommended_fix": "Move the 'one moment' principle definition and ID mapping instructions into the system prompt loaded via load_prompt, or use a template engine for the user prompt.", "severity": "P1", "why_problematic": "Core domain nomenclature ('한 찰나' principle) and task instructions are hardcoded in the Python string instead of being managed in the system prompt or a structured template. This scatters prompt logic and makes it difficult to tune the validation criteria without code changes."}], "path": "backend/app/core/steps/shot_validator_step.py", "scan_kind": "python", "sha256": "4c28fc69a4a9e706c558d73f8bae3e411531373f1529fa9d2b553333ec3a1b86"}
{"candidate_reason": "python scope discovery", "chunk_end": 156, "chunk_start": 1, "chunk_summary": "The file is a central version and prompt dependency registry; no actionable findings were identified as it contains only configuration metadata and descriptive comments regarding version history.", "duration_ms": 19526, "findings": [], "path": "backend/app/core/version_registry.py", "scan_kind": "python", "sha256": "b4454265a347ae52a317b1277af5b517365c83f12e58fc952e6630c916bda0c1"}
{"candidate_reason": "python scope discovery", "chunk_end": 1016, "chunk_start": 1, "chunk_summary": "The file contains several instances of scenario-dependent logic and semantic string judgments, particularly in image filtering and status reporting.", "duration_ms": 255195, "findings": [{"category": "semantic_string_judgment", "evidence": "returns images whose prompt_used contains 'outlook_id:{outlook_id}'", "line_end": 108, "line_start": 107, "recommended_fix": "Store outlook_id as a structured metadata column or a foreign key in the ImageAsset table instead of embedding it in the prompt string.", "severity": "P0", "why_problematic": "The API uses a substring check on a prompt string (open-world text) to determine a relational mapping (outlook_id). This is a fragile way to handle metadata and violates the rule against using substring checks to decide semantic meaning."}, {"category": "scenario_dependent_code", "evidence": "if entity.entity_type == \"character\": ... elif entity.entity_type == \"outlook\":", "line_end": 368, "line_start": 353, "recommended_fix": "Implement a more generic relationship registry or use a polymorphic approach where entities define their own composite dependencies.", "severity": "P1", "why_problematic": "The logic for finding affected composites is hardcoded to specific entity types ('character', 'outlook'). This makes the system rigid and scenario-dependent; adding new entity types requires code changes in the API layer."}, {"category": "semantic_string_judgment", "evidence": "enforce_shot_binding=(custom_prompt is None)", "line_end": 419, "line_start": 419, "recommended_fix": "Use an explicit boolean flag in the request body to control binding enforcement rather than inferring it from the prompt's presence.", "severity": "P1", "why_problematic": "The presence or absence of a string (custom_prompt) is used to bypass a safety gate (catalog freshness/shot binding). This uses string state to decide pipeline routing logic."}, {"category": "scenario_dependent_code", "evidence": "EntityCanon.entity_type != \"location\"", "line_end": 739, "line_start": 739, "recommended_fix": "Move this logic to a configuration or a method on the Entity model that determines if an entity type requires reference images.", "severity": "P2", "why_problematic": "Hardcoded exclusion of 'location' from reference image counts assumes that locations never have reference images. This is a scenario-specific rule that may not hold for all project types."}, {"category": "scenario_dependent_code", "evidence": "ImageAsset.variant_type == \"angle_fal\"", "line_end": 819, "line_start": 819, "recommended_fix": "Use a more generic variant category or move vendor-specific counts to a separate metadata-driven statistics service.", "severity": "P2", "why_problematic": "The status API contains a hardcoded string referencing a specific vendor ('fal'). This pollutes the general API with vendor-specific implementation details."}], "path": "backend/app/api/v1/images.py", "scan_kind": "python", "sha256": "5b74e60458b29711ced0f37357bf4401cf87a5b19b1b52a81fe798e58a34f6b9"}
{"candidate_reason": "python scope discovery", "chunk_end": 139, "chunk_start": 1, "chunk_summary": "No actionable findings; the module serves as a data tracking and aggregation layer using closed-world status constants.", "duration_ms": 6633, "findings": [], "path": "backend/app/modules/generation_tracker.py", "scan_kind": "python", "sha256": "635516c04a47c0e75bbe854a51872482c73a7b9b55195c7d16f4df442b04c6fa"}
{"candidate_reason": "python scope discovery", "chunk_end": 122, "chunk_start": 1, "chunk_summary": "The image checkpoint manager uses closed-world identifiers and stage enums for state persistence, with no actionable findings regarding semantic string judgment or scenario pollution.", "duration_ms": 4893, "findings": [], "path": "backend/app/modules/image_checkpoint.py", "scan_kind": "python", "sha256": "b253475dfa2921c3c105af37d0654414f3438c8f17adef04d4f4e31fcc8081a4"}
{"candidate_reason": "python scope discovery", "chunk_end": 1, "chunk_start": 1, "chunk_summary": "No actionable findings; the file chunk contains only a standard module docstring.", "duration_ms": 2062, "findings": [], "path": "backend/app/modules/llm/__init__.py", "scan_kind": "python", "sha256": "154c2f746a44ce90bdb76dee3da6fbcef2f58bd62941312a6fdb533cc3f4882d"}
{"candidate_reason": "python scope discovery", "chunk_end": 18, "chunk_start": 1, "chunk_summary": "No actionable findings; the file defines a generic abstract base class for LLM clients using structured schema validation.", "duration_ms": 3438, "findings": [], "path": "backend/app/modules/llm/base.py", "scan_kind": "python", "sha256": "e80908ebc287a538add83673361d6a5652e7508a279e0c95d3940885d1c8b3a9"}
{"candidate_reason": "python scope discovery", "chunk_end": 953, "chunk_start": 1, "chunk_summary": "The validator implements semantic enforcement and exemption logic using brittle substring matching and hardcoded keyword regex against open-world prompt text.", "duration_ms": 42192, "findings": [{"category": "semantic_string_judgment", "evidence": "_FACE_CLOSE_UP_PATTERNS", "line_end": 145, "line_start": 139, "recommended_fix": "Move semantic classification to the LLM producer as a structured boolean field in the prompt card (e.g., is_face_close_up).", "severity": "P1", "why_problematic": "Hardcoded keyword list (face, eye, expression, etc.) and regex used to classify prompt semantics to decide validation routing."}, {"category": "semantic_string_judgment", "evidence": "t.lower() in prompt_lower", "line_end": 217, "line_start": 217, "recommended_fix": "The producer should explicitly flag the exemption reason in the id_policy rather than relying on substring matching of phrases.", "severity": "P1", "why_problematic": "Judges semantic exemption based on whether a trigger phrase (open-world text) exists as a substring in the prompt."}, {"category": "semantic_string_judgment", "evidence": "prompt_lower.find(name_lower, offset)", "line_end": 553, "line_start": 553, "recommended_fix": "Require the LLM to output structured entity references or use a dedicated semantic parser instead of substring matching on names.", "severity": "P1", "why_problematic": "Uses character names (open-world strings) to find occurrences in the prompt to enforce ID proximity, which is a substring check for semantic meaning."}], "path": "backend/app/core/visible_entities_validator.py", "scan_kind": "python", "sha256": "366e2185c5df853112075146a74dc8d81d4c36e5fe4441689b7cc81833b80485"}
{"candidate_reason": "python scope discovery", "chunk_end": 90, "chunk_start": 1, "chunk_summary": "No actionable findings; the file contains standard infrastructure logic for managing API key rotation via environment variables.", "duration_ms": 3930, "findings": [], "path": "backend/app/modules/llm/gemini_key_pool.py", "scan_kind": "python", "sha256": "b0eb3589a24814832578ac71eada7146f061dfa877735336bc58949c7976e48b"}
{"candidate_reason": "python scope discovery", "chunk_end": 214, "chunk_start": 1, "chunk_summary": "The module implements dependency graph construction using hardcoded semantic keywords and positional assumptions for visual reference routing.", "duration_ms": 36676, "findings": [{"category": "semantic_string_judgment", "evidence": "VISUAL_REFERENCE_RULES: Dict[tuple, Set[str]] = { ... \"identity\", \"transformation\", \"possession\" ... }", "line_end": 26, "line_start": 18, "recommended_fix": "Move visual dependency flags into the relation schema (e.g., 'is_visual_reference: bool') or use a centralized SOT for relation types that includes visual metadata, rather than hardcoding semantic strings in the logic module.", "severity": "P1", "why_problematic": "Visual dependency routing is decided by matching hardcoded semantic keywords ('identity', 'possession', etc.) against the 'relation_family' field. This is a string-pattern based judgment of open-world semantic relationships, which is fragile if the analyzer uses different terminology or if new visual-impacting relations are added."}, {"category": "semantic_string_judgment", "evidence": "for i, p1 in enumerate(parts): for p2 in parts[i + 1 :]: ... deps[id2].add(id1)", "line_end": 76, "line_start": 59, "recommended_fix": "Explicitly define roles in the participant schema (e.g., 'role': 'source' vs 'role': 'target') and use those roles to determine dependency direction instead of relying on list order.", "severity": "P1", "why_problematic": "The direction of visual dependency is blindly inferred from the index order of participants in the list (treating the first as 'primary'). This assumes the semantic 'source' of a reference always appears first in the data structure, which is an unreliable way to judge semantic roles in open-world story data."}], "path": "backend/app/modules/entity_dependency.py", "scan_kind": "python", "sha256": "d7d1dc91f1acdd43734d0097877c1e5646006e307759d8e080654f80375d5ab3"}
{"candidate_reason": "python scope discovery", "chunk_end": 221, "chunk_start": 1, "chunk_summary": "no actionable findings", "duration_ms": 4329, "findings": [], "path": "backend/app/modules/llm/gemini_text_client.py", "scan_kind": "python", "sha256": "ab01305191af5c3df4f4b3f5085b038ead3cc5c3a2e918c810577467921fc89b"}
{"candidate_reason": "python scope discovery", "chunk_end": 369, "chunk_start": 1, "chunk_summary": "The Gemini image client contains scenario-dependent metadata keys for logging and a hardcoded prompt fragment for reference images.", "duration_ms": 24245, "findings": [{"category": "scenario_dependent_code", "evidence": "_LOG_COLUMN_KEYS = frozenset({...}), _LOG_METADATA_KEYS = frozenset({...})", "line_end": 59, "line_start": 52, "recommended_fix": "Inject the list of loggable metadata keys via configuration or pass a pre-filtered metadata dictionary to the client instead of hardcoding domain-specific keys.", "severity": "P1", "why_problematic": "The client hardcodes domain-specific keys like 'episode_id', 'scene_index', 'shot_index', and 'beat_title'. This couples the low-level LLM client to the specific scenario structure of the storytelling domain, making it difficult to reuse or adapt to different domain models."}, {"category": "scenario_dependent_prompt", "evidence": "parts.append({\"text\": label or f\"Reference image {i}:\"})", "line_end": 122, "line_start": 122, "recommended_fix": "Move prompt fragments to a template or SOT, or ensure labels are always provided by the caller.", "severity": "P2", "why_problematic": "Hardcoded English prompt fragment 'Reference image' is used as a fallback label. This is scenario-dependent and prevents localization or task-specific labeling from a structured SOT."}], "path": "backend/app/modules/llm/gemini_image_client.py", "scan_kind": "python", "sha256": "da5180d7d8c0c04163c45e1c6f95e1d82d0243af4597aefb939e2dfd6baf3162"}
{"candidate_reason": "python scope discovery", "chunk_end": 351, "chunk_start": 1, "chunk_summary": "The module implements an image-to-image editor using Gemini, but contains hardcoded semantic labels for multi-modal inputs and visual diagrams that couple the code to external prompt logic.", "duration_ms": 45024, "findings": [{"category": "scenario_dependent_prompt", "evidence": "\"Original scene image:\", \"Camera angle diagram:\"", "line_end": 317, "line_start": 244, "recommended_fix": "Define these labels as constants or include them in a structured prompt configuration (SOT) that both the code and the prompt templates reference.", "severity": "P2", "why_problematic": "Hardcoded labels for multi-modal image parts create a hidden coupling between the Python logic and the external prompt files (e.g., angle.md). If the prompt instructions refer to these specific labels to identify images, changing them requires synchronized updates across code and prompt assets."}, {"category": "scenario_dependent_prompt", "evidence": "\"IMAGE\", \"CAMERA\", \"Angle:\", \"Elevation:\", \"Zoom:\"", "line_end": 92, "line_start": 69, "recommended_fix": "Move visual label strings to a centralized configuration or pass them as parameters to the diagram generation function.", "severity": "P2", "why_problematic": "These strings are rendered into a diagram used as a visual prompt for the LLM. They constitute domain nomenclature that the LLM must interpret. Hardcoding them in the drawing logic makes the visual prompting strategy difficult to manage or localize."}], "path": "backend/app/modules/gemini_i2i_editor.py", "scan_kind": "python", "sha256": "3ecd9be0d08929461578a029b148fa228612db99edd966f28d8d46d44b5d3a8a"}
{"candidate_reason": "python scope discovery", "chunk_end": 289, "chunk_start": 1, "chunk_summary": "The image validation module uses GPT-5.4 Vision to verify generated images against scenario data, but employs a blind heuristic to extract visual traits from arbitrary dictionary values.", "duration_ms": 38218, "findings": [{"category": "semantic_string_judgment", "evidence": "if not visual_traits and isinstance(traits_data, dict): for _k, v in traits_data.items(): if isinstance(v, str): visual_traits.append(v)", "line_end": 160, "line_start": 157, "recommended_fix": "Enforce a strict schema for visual traits (e.g., requiring the 'visual_anchor_traits' key) and remove the fallback that scrapes arbitrary string values from the dictionary.", "severity": "P2", "why_problematic": "This logic blindly promotes any string value found in the traits dictionary to a 'visual trait' for the LLM to validate. This is a semantic judgment based on data type rather than schema, which can pollute the prompt with non-visual metadata (e.g., 'age', 'alignment') and lead to incorrect validation results or LLM hallucinations."}], "path": "backend/app/modules/image_validator.py", "scan_kind": "python", "sha256": "e6401c59a6af4492ce33f2619f42f239905bc5551008b0b4e4591a93cfe7948b"}
{"candidate_reason": "python scope discovery", "chunk_end": 282, "chunk_start": 1, "chunk_summary": "The module contains hardcoded domain nomenclature and prompt fragments that create maintenance debt and potential for schema drift between code and external prompts.", "duration_ms": 48280, "findings": [{"category": "schema_or_enum_drift", "evidence": "VISUAL_RELATION_FAMILIES = [\"identity\", \"transformation\", \"possession\", \"containment\"]", "line_end": 22, "line_start": 17, "recommended_fix": "Load the valid relation families from a shared configuration or the prompt metadata itself to ensure the schema and prompt remain synchronized.", "severity": "P1", "why_problematic": "These semantic categories are hardcoded in the Python logic but must stay in sync with the external prompt (v7). This scattered domain nomenclature creates a high risk of runtime validation failures if the prompt is updated without a corresponding code change."}, {"category": "scenario_dependent_code", "evidence": "SUPPORTED_LANGUAGE_MAP.get(language, \"Korean\")", "line_end": 252, "line_start": 252, "recommended_fix": "Use a configuration-driven default or raise an error for unsupported languages.", "severity": "P2", "why_problematic": "Hardcoded default language 'Korean' assumes a specific project context within the logic rather than using a configurable default."}, {"category": "scenario_dependent_prompt", "evidence": "\"(없음 / None)\"", "line_end": 261, "line_start": 261, "recommended_fix": "Move the fallback string into the chunk_user.md template or pass it as a variable from a configuration file.", "severity": "P2", "why_problematic": "Hardcoded string literal used as a prompt fallback. This couples the code to a specific language/format that should be handled within the prompt template or a localization layer."}], "path": "backend/app/modules/entity_extractor_legacy.py", "scan_kind": "python", "sha256": "9edba4cc7829246a732320b6db3828ad8c1692f47fd691cf57b98eb2994998ed"}
{"candidate_reason": "python scope discovery", "chunk_end": 216, "chunk_start": 1, "chunk_summary": "The file implements an image generation tracer for Opik, but contains a substring-based heuristic to determine the model provider.", "duration_ms": 12579, "findings": [{"category": "semantic_string_judgment", "evidence": "provider=\"google_ai\" if \"gemini\" in model or \"imagen\" in model else \"fal.ai\"", "line_end": 138, "line_start": 138, "recommended_fix": "Pass the provider as an explicit argument to the log/span methods, or use a configuration-driven mapping (registry) of model IDs to providers.", "severity": "P1", "why_problematic": "The code uses substring checks on the 'model' string to infer the 'provider' metadata. This is a pattern-based semantic judgment that fails if model naming conventions change or if a new provider is added, defaulting everything else to 'fal.ai'."}], "path": "backend/app/modules/llm/image_tracer.py", "scan_kind": "python", "sha256": "0ebe23c2d9b90987dfc52fac1097d9edc98d866707618f5e881a6d5fece24a91"}
{"candidate_reason": "python scope discovery", "chunk_end": 169, "chunk_start": 1, "chunk_summary": "No actionable findings; the file provides a standard infrastructure wrapper for OpenAI API calls with structural JSON parsing and logging.", "duration_ms": 7279, "findings": [], "path": "backend/app/modules/llm/openai_client.py", "scan_kind": "python", "sha256": "b09b7d78ae5185fd4ffc7c1a080f9c8f4bb45efeca66da594fef3bdfe7212d2f"}
{"candidate_reason": "python scope discovery", "chunk_end": 82, "chunk_start": 1, "chunk_summary": "The LLM logger module is a generic utility for recording model interactions and does not contain semantic string judgments or scenario-dependent logic.", "duration_ms": 15262, "findings": [], "path": "backend/app/modules/llm/llm_logger.py", "scan_kind": "python", "sha256": "bc280499bbdb46679b12250b3904c4c770a5cf5128a42d9172beff71e0d09c30"}
{"candidate_reason": "python scope discovery", "chunk_end": 236, "chunk_start": 1, "chunk_summary": "The module implements safety filter bypass logic using hardcoded string replacements and keyword-based error routing, which introduces semantic judgment debt and scenario-dependent pollution.", "duration_ms": 13262, "findings": [{"category": "blind_string_mutation", "evidence": "_SAFETY_REPLACEMENTS_KO = [...] / _SAFETY_REPLACEMENTS_EN = [...]", "line_end": 73, "line_start": 26, "recommended_fix": "Move safety-related rephrasing to a dedicated LLM pass with a structured SOT defining the 'filming' context, rather than using static string mutation.", "severity": "P1", "why_problematic": "Uses hardcoded lists of nouns and phrases (e.g., '시신', 'blood', 'murder') to perform blind string replacement on open-world scenario text. This is a fragile attempt to alter semantic meaning via pattern matching to bypass safety filters."}, {"category": "scenario_dependent_prompt", "evidence": "SAFETY_SYSTEM_SUFFIX = (...)", "line_end": 102, "line_start": 94, "recommended_fix": "Parameterize the safety framing prompt to accept context-specific examples or use a generic framing that refers to the 'production context' without hardcoding specific props.", "severity": "P1", "why_problematic": "The prompt contains scenario-specific examples like 'aged photograph prop' and 'dark red stage paint pool'. This pollutes the global safety utility with domain-specific nomenclature that should be derived from the scenario's SOT."}, {"category": "semantic_string_judgment", "evidence": "_SAFETY_MESSAGE_KEYWORDS = (...)", "line_end": 175, "line_start": 160, "recommended_fix": "Map provider-specific error codes/enums to internal status constants instead of performing substring searches on the raw error message string.", "severity": "P1", "why_problematic": "Uses substring checks on open-world error messages from external LLM providers (e.g., 'harm', 'safety', 'moderation') to decide critical execution routing (Tier 2/3 fallback). This is pattern-based semantic judgment of unstructured error text."}], "path": "backend/app/modules/llm/safety.py", "scan_kind": "python", "sha256": "bf4dadd9df7fe27d183af62dc19d91d9f173c086e3e9705dac55bce4b42e2b46"}
{"candidate_reason": "python scope discovery", "chunk_end": 51, "chunk_start": 1, "chunk_summary": "The name matcher utility uses string prefix matching to determine character identity, which is a pattern-based judgment of open-world semantic meaning.", "duration_ms": 14027, "findings": [{"category": "semantic_string_judgment", "evidence": "entity_name.startswith(shot_name)", "line_end": 30, "line_start": 29, "recommended_fix": "Use a structured alias mapping in the character SOT or a dedicated entity resolution step instead of heuristic prefix matching.", "severity": "P1", "why_problematic": "Uses a blind prefix check to decide if two names refer to the same semantic entity. This is a string-pattern-based judgment of open-world semantics that can lead to false positives (e.g., 'Kim' matching 'Kimberly') when resolving character names from shot descriptions to canonical entities."}], "path": "backend/app/modules/name_matcher.py", "scan_kind": "python", "sha256": "431470212b80e7a3bd4b33fcb142e5499adc75515995ec4e849041ce8a75e34e"}
{"candidate_reason": "python scope discovery", "chunk_end": 36, "chunk_start": 1, "chunk_summary": "No actionable findings; the code performs standard environment variable resolution for infrastructure configuration.", "duration_ms": 3446, "findings": [], "path": "backend/app/modules/pipeline/_workers.py", "scan_kind": "python", "sha256": "022738fed30d0efce3ee58a0a4ef4bac167dcd1e61c1257148162d97ea593559"}
{"candidate_reason": "python scope discovery", "chunk_end": 58, "chunk_start": 1, "chunk_summary": "no actionable findings", "duration_ms": 6269, "findings": [], "path": "backend/app/modules/pipeline/_dag_levels.py", "scan_kind": "python", "sha256": "50b9da1dd402a115f0a6951f0b978e3ed43213794210fbeec653b06882cc504a"}
{"candidate_reason": "python scope discovery", "chunk_end": 35, "chunk_start": 1, "chunk_summary": "The module uses brittle character-range heuristics to determine the language of open-world text.", "duration_ms": 21105, "findings": [{"category": "semantic_string_judgment", "evidence": "ko_count = sum(1 for c in sample if '\\uAC00' <= c <= '\\uD7A3')", "line_end": 35, "line_start": 26, "recommended_fix": "Replace manual character counting with a robust language detection library or move language identification to a configuration/SOT layer.", "severity": "P1", "why_problematic": "Determines the language of open-world scenario text using hardcoded character ranges and arbitrary count thresholds (50), which is a brittle heuristic for semantic classification."}], "path": "backend/app/modules/pdf_parser.py", "scan_kind": "python", "sha256": "6dae8e6cd873a2b8f8b6bcf8de055862e1c3c0105bbeefe6eff5d75ffc091a79"}
{"candidate_reason": "python scope discovery", "chunk_end": 281, "chunk_start": 1, "chunk_summary": "The PDF validator contains hardcoded scenario-specific nomenclature and semantic requirements within the prompt construction and validation schema.", "duration_ms": 19603, "findings": [{"category": "scenario_dependent_code", "evidence": "\"caption_visible\": {\"type\": \"boolean\"}", "line_end": 27, "line_start": 27, "recommended_fix": "Externalize the validation criteria to a configuration file or the prompt metadata so the schema can be generated dynamically based on the document type.", "severity": "P1", "why_problematic": "The validation schema hardcodes 'caption_visible' as a required semantic check. This is scenario-specific nomenclature (specific to the webbook format) that couples the validator's logic to a particular document structure."}, {"category": "scenario_dependent_prompt", "evidence": "\"text\": \"Validate this webbook PDF page.\"", "line_end": 147, "line_start": 147, "recommended_fix": "Move this instruction into the externalized system prompt or pass the document type as a variable to the prompt loader.", "severity": "P2", "why_problematic": "The prompt fragment hardcodes the term 'webbook', which is domain-specific nomenclature. This scattered prompt logic should be consolidated into the structured prompt loader (validate.md)."}], "path": "backend/app/modules/pdf_validator.py", "scan_kind": "python", "sha256": "5cf00222ec7f7687470cb6a65aa529f7ddb6a29be7cc8e0c2464fdd5518ffaaf"}
{"candidate_reason": "python scope discovery", "chunk_end": 778, "chunk_start": 1, "chunk_summary": "The LLM client contains scenario-dependent prompt pollution and hardcoded domain nomenclature that bypasses the central manifest.", "duration_ms": 42160, "findings": [{"category": "scenario_dependent_code", "evidence": "_PIPELINE_STEP_EXTENSIONS", "line_end": 304, "line_start": 278, "recommended_fix": "Move all extension steps into the central STEP_MANIFEST and have the client build its view solely from the manifest.", "severity": "P1", "why_problematic": "Hardcodes domain-specific step names (e.g., 'scene_detail_owned_judge', 'location_consistency') and Korean UI labels directly in the client. This bypasses the STEP_MANIFEST SOT and pollutes the generic client with scenario-specific nomenclature."}, {"category": "scenario_dependent_prompt", "evidence": "system_prompt + SAFETY_SYSTEM_SUFFIX", "line_end": 765, "line_start": 553, "recommended_fix": "Pass the safety framing as an optional parameter or resolve it via the step manifest rather than hardcoding a global suffix in the client.", "severity": "P1", "why_problematic": "The client blindly appends SAFETY_SYSTEM_SUFFIX to system prompts in Tier 2 fallbacks. Comments (lines 551, 759) indicate this is 'movie framing' (영화 프레이밍), which is scenario-specific instruction injected into a generic LLM client."}, {"category": "scenario_dependent_code", "evidence": "_GEMINI_SAFETY_SETTINGS_OFF", "line_end": 409, "line_start": 404, "recommended_fix": "Move safety threshold configuration to the project config or step manifest instead of applying it globally in the client.", "severity": "P2", "why_problematic": "Disables Gemini safety filters globally based on the assumption that the workload consists of 'fiction scenarios' (movies/dramas). This is scenario-specific logic that should be configurable per project or step."}], "path": "backend/app/modules/llm/llm_client.py", "scan_kind": "python", "sha256": "16dda2eae6b7463c05729ddb95db0d5fbc4a5bacf78e9769c2f11af368df8340"}
{"candidate_reason": "python scope discovery", "chunk_end": 361, "chunk_start": 1, "chunk_summary": "The PDF renderer module is a structural implementation for the Webbook format and does not contain actionable semantic string judgments or scenario-dependent pollution.", "duration_ms": 27351, "findings": [], "path": "backend/app/modules/pdf_renderer.py", "scan_kind": "python", "sha256": "24d6b3c7ccf5da2095f93686ae6c526580f18ac2e3bb0a756c255f7a6b00f5ce"}
{"candidate_reason": "python scope discovery", "chunk_end": 1240, "chunk_start": 1, "chunk_summary": "The file contains several instances of semantic judgment based on string patterns and scenario-dependent prompt pollution, particularly in error handling and prompt reinforcement.", "duration_ms": 29907, "findings": [{"category": "semantic_string_judgment", "evidence": "_MODERATION_KEYWORDS = (\"moderation\", \"safety\", \"content_policy\", \"prohibited\", \"policy\", \"blocked\", \"violates\", \"violation\")", "line_end": 48, "line_start": 45, "recommended_fix": "Use structured error codes from the OpenAI API (e.g., check for specific exception types or 'code' fields in the error response) instead of substring matching on the message.", "severity": "P1", "why_problematic": "Heuristic keyword list used to judge the semantic intent of external API error messages to drive internal retry/sanitization logic."}, {"category": "scenario_dependent_prompt", "evidence": "_BACKGROUND_ONLY_REINFORCEMENT = (\"BACKGROUND-ONLY architectural still — empty space, NO people, NO faces, NO body posture, NO action, NO weapons, NO blood. Photoreal scene without any human figures.\\n\\n\")", "line_end": 61, "line_start": 57, "recommended_fix": "Move these constraints into the prompt template or a configuration object passed to the renderer, rather than hardcoding them as a global constant.", "severity": "P1", "why_problematic": "Hardcoded scenario-specific constraints are injected into the prompt pipeline, creating tight coupling between the rendering logic and a specific content type (backgrounds)."}, {"category": "scenario_dependent_prompt", "evidence": "parent_block = ( ... \"[The parent's already-rendered photo will be the reference image for this node. Reuse parent's atmosphere/material/lighting language. Explicitly mention each shared_visual_anchors_with_parent entry.]\" ... )", "line_end": 92, "line_start": 79, "recommended_fix": "Move the parent and root anchor instructions into the 'user_template' or 'system' prompt files as conditional blocks.", "severity": "P1", "why_problematic": "Scenario-specific instructions are hardcoded in the Python logic instead of being managed within the external prompt template, leading to domain pollution in the pipeline code."}, {"category": "semantic_string_judgment", "evidence": "if _NON_ASCII_TEXT_RE.search(t2i): raise ValueError(f\"node {node.get('id')!r} t2i_prompt contains non-ASCII text (Korean/Hanja/kana detected — universal-noun rule violated)\")", "line_end": 155, "line_start": 151, "recommended_fix": "If the goal is to ensure English-only output, this is a valid technical check, but the 'universal-noun' semantic enforcement should ideally be handled by the system prompt or a dedicated validator.", "severity": "P2", "why_problematic": "Uses a character-set regex to enforce a semantic rule ('universal-noun rule') on open-world LLM output, which is a proxy for judging the 'locality' of the content."}, {"category": "blind_string_mutation", "evidence": "if not sanitized.lstrip().startswith(\"BACKGROUND-ONLY\"): sanitized = _BACKGROUND_ONLY_REINFORCEMENT + sanitized", "line_end": 290, "line_start": 289, "recommended_fix": "The sanitizer should return a structured flag indicating if reinforcement is needed, or the reinforcement should be applied based on the pipeline configuration.", "severity": "P2", "why_problematic": "Uses a blind string prefix check to decide whether to mutate the prompt with scenario-specific reinforcement, which is fragile and bypasses structured prompt management."}], "path": "backend/app/modules/pipeline/background_chain_render.py", "scan_kind": "python", "sha256": "f68591b10e7661b59fffffffde393e9c9208fba64dbe4c7d8e6b3c69116b4da6"}
{"candidate_reason": "python scope discovery", "chunk_end": 224, "chunk_start": 1, "chunk_summary": "The file implements background classification logic using an LLM, but contains hardcoded semantic heuristics and scenario-dependent validation rules for specific classification types.", "duration_ms": 31686, "findings": [{"category": "scenario_dependent_code", "evidence": "if kind == \"chain_bg\": ... if total_shots < 3: ... if not has_indoor: ...", "line_end": 168, "line_start": 151, "recommended_fix": "Move the 'chain_bg' validation criteria (shot count and indoor requirement) into the system prompt or a configuration object passed to the validator.", "severity": "P1", "why_problematic": "The code enforces a rigid semantic definition for 'chain_bg' (minimum 3 shots and at least one indoor location) via hardcoded heuristics. This logic is scenario-dependent and should be part of the LLM prompt or a configurable policy rather than hardcoded in the validator, as it may conflict with LLM semantic clustering or vary between different types of productions."}, {"category": "schema_or_enum_drift", "evidence": "if kind not in {\"chain_bg\", \"prev_shot_ref\"}:", "line_end": 148, "line_start": 147, "recommended_fix": "Define the allowed 'kind' values in a shared schema or configuration file that is used by both the prompt and the validator.", "severity": "P2", "why_problematic": "The classification 'kind' is restricted to a hardcoded set of strings in the code. This creates a dependency between the code and the LLM's semantic output categories which may evolve or vary by project, leading to drift if the schema or prompt is updated without updating this validator."}], "path": "backend/app/modules/pipeline/background_classify.py", "scan_kind": "python", "sha256": "49b9985cf8fb46be2e33e57cd8ed86814808df13e625a334b417f431d1bfeea4"}
{"candidate_reason": "python scope discovery", "chunk_end": 473, "chunk_start": 1, "chunk_summary": "The file implements the background master plan logic, including prompt construction and validation of LLM outputs against schema and structural invariants, with no actionable findings regarding semantic string judgment or scenario pollution.", "duration_ms": 31857, "findings": [], "path": "backend/app/modules/pipeline/background_master_plan.py", "scan_kind": "python", "sha256": "d98a2853ffce308c5404b5652147816a637a2cabd6f9acb6847a521675c00cbe"}
{"candidate_reason": "python scope discovery", "chunk_end": 1136, "chunk_start": 1, "chunk_summary": "The file contains logic for planning background image chains, including entity resolution heuristics and language-based output validation that use pattern matching for semantic judgment.", "duration_ms": 48758, "findings": [{"category": "semantic_string_judgment", "evidence": "if len(ename) >= 2 and (loc_raw in ename or ename in loc_raw):", "line_end": 154, "line_start": 151, "recommended_fix": "Ensure the upstream scene_director provides the short_id directly, or use a strict mapping dictionary without fuzzy substring fallbacks.", "severity": "P1", "why_problematic": "This logic uses substring containment to resolve entity identity by mapping raw location names to short_ids. Deciding if two strings refer to the same entity based on substring patterns is a heuristic semantic judgment that should be handled by exact ID lookups or a dedicated resolution step in the SOT."}, {"category": "semantic_string_judgment", "evidence": "_NON_ASCII_TEXT_RE = re.compile(r\"[ㄱ-ㆎ가-힣一-鿿㐀-䶿豈-﫿぀-ヿ]\")", "line_end": 39, "line_start": 35, "recommended_fix": "Downgrade this check to a soft diagnostic (warning) that does not fail the pipeline, or use an LLM-based reviewer to verify language compliance.", "severity": "P2", "why_problematic": "This regex is used to detect non-ASCII characters (Hangul/CJK) to enforce the 'universal-noun' rule. It is used in validate_plan (lines 390, 487, etc.) to trigger hard failures (status='failed'). Judging the language or character set of open-world text via regex is a pattern-based semantic judgment used to fail validation, which is forbidden."}], "path": "backend/app/modules/pipeline/background_chain_planning.py", "scan_kind": "python", "sha256": "0f30de23db87c06e6e9c30c8344679eeb02c49a6348e0f9df6dd19984af8714d"}
{"candidate_reason": "python scope discovery", "chunk_end": 166, "chunk_start": 1, "chunk_summary": "The module uses keyword-based substring checks to classify API error semantics and performs blind prompt mutations based on hardcoded string prefixes.", "duration_ms": 21686, "findings": [{"category": "semantic_string_judgment", "evidence": "any(k in msg for k in (\"moderation\", \"safety\", \"content_policy\", \"prohibited\", \"policy\", \"blocked\", \"violates\", \"violation\"))", "line_end": 137, "line_start": 124, "recommended_fix": "Use official error codes, exception types, or structured error objects provided by the OpenAI SDK (e.g., checking for specific error codes in the API response) instead of parsing the human-readable message string.", "severity": "P1", "why_problematic": "Determines the semantic nature of an external API error (OpenAI moderation block) using a hardcoded list of substrings. This is fragile as error messages are open-world text and subject to change by the provider, leading to incorrect routing or missed retries."}, {"category": "blind_string_mutation", "evidence": "if not sanitized.lstrip().startswith(\"BACKGROUND-ONLY\"): sanitized = _BACKGROUND_ONLY_REINFORCEMENT + sanitized", "line_end": 149, "line_start": 148, "recommended_fix": "The sanitizer should return structured metadata indicating if reinforcement was already applied, or the prompt should be managed by a structured builder that handles layers (base prompt + reinforcement) separately.", "severity": "P2", "why_problematic": "Uses a hardcoded string prefix check to decide whether to prepend a reinforcement block. This assumes semantic state from a string pattern and performs a blind concatenation, which is brittle if the sanitizer changes its output format."}, {"category": "scenario_dependent_prompt", "evidence": "_BACKGROUND_ONLY_REINFORCEMENT = ( \"BACKGROUND-ONLY architectural still — empty space, NO people, NO faces, ... )", "line_end": 26, "line_start": 22, "recommended_fix": "Move reinforcement prompt fragments to a centralized configuration or a dedicated prompt-card system where domain-specific constraints are managed as data.", "severity": "P2", "why_problematic": "Hardcodes a specific list of negative constraints and domain-specific instructions (architectural still, NO people, etc.) directly in the pipeline code. This is scenario-dependent pollution that should be decoupled from the execution logic."}], "path": "backend/app/modules/pipeline/background_render.py", "scan_kind": "python", "sha256": "497a49b7050919807a148e8fa32be91353e756dea2991ba82b21f4691213d15d"}
{"candidate_reason": "python scope discovery", "chunk_end": 280, "chunk_start": 1, "chunk_summary": "The file contains scenario-dependent nomenclature in error messages, uses a technical string property (ASCII) to enforce an open-world semantic constraint, and handles schema drift through redundant template replacements.", "duration_ms": 53454, "findings": [{"category": "semantic_string_judgment", "evidence": "item.encode(\"ascii\") ... owned MUST be English canonical common nouns", "line_end": 210, "line_start": 203, "recommended_fix": "If semantic validation of 'common nouns' is required, use an LLM-based validator or a proper NLP library. If the goal is just character set safety, separate the technical check from the semantic requirement in the error message.", "severity": "P1", "why_problematic": "The code uses a character set check (ASCII) as a proxy to enforce an open-world semantic constraint ('English canonical common nouns'). This is a string-pattern-based judgment of semantic meaning."}, {"category": "scenario_dependent_code", "evidence": "round 4 Q2=B / round 6 BLOCKING 3", "line_end": 209, "line_start": 208, "recommended_fix": "Remove internal project references from the error message. Use a generic description of the requirement.", "severity": "P1", "why_problematic": "The error message contains internal project-specific milestone references ('round 4 Q2=B') and blocking rules ('round 6 BLOCKING 3') which pollute the code with scenario-specific history and nomenclature."}, {"category": "schema_or_enum_drift", "evidence": ".replace(\"{state_label_raw}\", state_label_raw).replace(\"{state_label}\", state_label_raw)", "line_end": 146, "line_start": 145, "recommended_fix": "Standardize the template placeholders and the input bg_spec schema to avoid redundant replacements for the same semantic value.", "severity": "P2", "why_problematic": "The code performs redundant replacements to handle 'pack drift' (legacy schema versions), which is scenario-dependent logic that should be resolved at the data ingestion or schema migration layer."}], "path": "backend/app/modules/pipeline/background_prompt.py", "scan_kind": "python", "sha256": "de5f48a96c5a19b9b07b12506ddaa30556c2053f398f2284c175be142c76a3b2"}
{"candidate_reason": "python scope discovery", "chunk_end": 526, "chunk_start": 1, "chunk_summary": "The file contains instances of semantic filtering based on LLM string outputs and hardcoded prompt instructions within the pipeline logic.", "duration_ms": 39372, "findings": [{"category": "semantic_string_judgment", "evidence": "e.get(\"importance\") == \"none\"", "line_end": 196, "line_start": 193, "recommended_fix": "Instruct the LLM to perform the filtering directly or use a numeric score with a configurable threshold.", "severity": "P1", "why_problematic": "The code performs a semantic filter on open-world entities by checking if an LLM-generated 'importance' field matches the string literal 'none'. This is a semantic decision based on a string pattern rather than structured data or LLM-native filtering."}, {"category": "scenario_dependent_prompt", "evidence": "\"기존 요소가 이번 에피소드에도 등장하면 이름을 동일하게 유지하세요.\"", "line_end": 288, "line_start": 288, "recommended_fix": "Move this instruction into the external markdown prompt templates.", "severity": "P2", "why_problematic": "Hardcoded prompt instruction for entity consistency is embedded in the Python logic, creating scenario-dependent pollution."}, {"category": "scenario_dependent_prompt", "evidence": "system_prompt=\"시나리오 분석 전문가. 요소별 시각적 상세 정보...\"", "line_end": 374, "line_start": 370, "recommended_fix": "Move the system prompt to a dedicated markdown file in the prompts directory.", "severity": "P2", "why_problematic": "The system prompt for entity detail extraction is hardcoded in the Python file, violating the separation of prompt and code."}], "path": "backend/app/modules/pipeline/entity_extractor_v2_legacy.py", "scan_kind": "python", "sha256": "491a07d32324e43bc3d0823cf4b20d94b9635a232efd731ff7f180a324f8c123"}
{"candidate_reason": "python scope discovery", "chunk_end": 499, "chunk_start": 1, "chunk_summary": "The background planner module contains brittle schema injection logic and uses regex-based character set detection to validate semantic fields, creating risks for schema drift and pattern-based semantic judgment.", "duration_ms": 65850, "findings": [{"category": "semantic_string_judgment", "evidence": "_NON_ASCII_TEXT_RE = re.compile(...)", "line_end": 33, "line_start": 28, "recommended_fix": "Remove character-set based semantic validation. If specific fields must be ASCII for technical reasons (e.g., file paths), enforce this via schema 'pattern' or 'format' constraints rather than manual regex checks on semantic categories like 'kind'.", "severity": "P1", "why_problematic": "The code uses a regex to detect Korean, Chinese, and Japanese characters to validate 'universal nouns' and semantic 'kind' fields (as seen in line 310). This constitutes judging open-world semantic validity based on character-set patterns rather than content or a structured SOT."}, {"category": "schema_or_enum_drift", "evidence": "def _inject_runtime_enums(schema: Dict[str, Any], ...)", "line_end": 416, "line_start": 361, "recommended_fix": "Use a schema-aware utility or a library like Pydantic to dynamically inject enums into the model definition, or use a more robust path-based injection that validates the existence of the target nodes.", "severity": "P1", "why_problematic": "This function manually traverses and mutates the JSON schema dictionary using hardcoded nested keys (e.g., props.get('floor_plans').get('items').get('properties')). This is highly brittle and will fail silently or cause runtime errors if the schema structure is updated or uses references ($ref)."}, {"category": "schema_or_enum_drift", "evidence": "_ID_FIELDS_FLOOR_PLAN = (\"id\", \"building_group\", \"primary_location_id\")", "line_end": 127, "line_start": 125, "recommended_fix": "Derive these field lists dynamically from the schema metadata or use a shared configuration that defines both the schema and its validation requirements.", "severity": "P2", "why_problematic": "Hardcoded lists of field names used for validation are disconnected from the actual schema definition. Renaming fields in the schema will not update these lists, leading to validation drift."}], "path": "backend/app/modules/pipeline/background_planner.py", "scan_kind": "python", "sha256": "647640786fa18df4962693aba3275f4d1a2c69b11d1cb8821117c14f74bb1d32"}
{"candidate_reason": "python scope discovery", "chunk_end": 402, "chunk_start": 1, "chunk_summary": "The entity lister uses hardcoded Korean string replacements and substring checks to semantically mutate prompt templates in code.", "duration_ms": 10366, "findings": [{"category": "blind_string_mutation", "evidence": ".replace(\"등장하는 씬 수\", \"등장하는 샷(스틸컷) 수\")", "line_end": 199, "line_start": 199, "recommended_fix": "Use separate prompt templates for shot-based extraction or use a templating engine (e.g., Jinja2) with variables instead of hardcoded string replacement.", "severity": "P1", "why_problematic": "The code performs a semantic transformation of the prompt by searching for a specific natural language phrase in Korean. This couples the Python logic to the exact wording of the prompt files, leading to silent failures if the prompt is edited."}, {"category": "blind_string_mutation", "evidence": "type_prompt_shot.replace(\"등장하는 씬 수\", \"등장하는 샷(스틸컷) 수\") ... if \"shot_count\" not in type_prompt_shot:", "line_end": 240, "line_start": 238, "recommended_fix": "Move the conditional instruction logic into the prompt template or use a structured configuration to determine which instructions to include.", "severity": "P1", "why_problematic": "Similar to line 199, this uses blind replacement and a substring check to decide whether to append additional semantic instructions. This logic should reside in the prompt management layer, not the pipeline execution code."}], "path": "backend/app/modules/pipeline/entity_lister.py", "scan_kind": "python", "sha256": "a0c86dcd35b83d29571024e2362155f34e605e48919512678f9fce63f7a693fb"}
{"candidate_reason": "python scope discovery", "chunk_end": 101, "chunk_start": 1, "chunk_summary": "The entity filtering logic relies on fragile string matching and regex-based cleaning to identify entities for removal, rather than using structured IDs.", "duration_ms": 25861, "findings": [{"category": "blind_string_mutation", "evidence": "re.sub(r'^[CLP]\\d{2,3}\\s*', '', raw_name).strip()", "line_end": 83, "line_start": 71, "recommended_fix": "Modify the LLM response schema to return the 'short_id' (e.g., C01) instead of the 'name'. Use the 'short_id' as the primary key for matching and removal logic.", "severity": "P1", "why_problematic": "The code uses regex to strip ID prefixes from LLM-returned strings to resolve entity identity. This is a fragile heuristic that fails if the LLM alters the name or formatting (e.g., adding punctuation or changing case) during its semantic judgment process."}, {"category": "semantic_string_judgment", "evidence": "e[\"name\"] not in rnames", "line_end": 90, "line_start": 90, "recommended_fix": "Perform filtering based on 'short_id' which is a stable, closed-world identifier already present in the entity objects.", "severity": "P1", "why_problematic": "Entity filtering is performed using exact string matching on open-world names. This is unreliable for identifying entities across LLM boundaries where minor textual variations or hallucinations can occur."}], "path": "backend/app/modules/pipeline/entity_filter.py", "scan_kind": "python", "sha256": "6942f826c21f51b159a159a2e14e22d2dd3826fee9b826bf1a8623cadca757d3"}
{"candidate_reason": "python scope discovery", "chunk_end": 522, "chunk_start": 1, "chunk_summary": "The file contains several instances of hardcoded prompt instructions and scenario-dependent formatting within the Python logic, as well as a module name mismatch that affects prompt loading.", "duration_ms": 43730, "findings": [{"category": "scenario_dependent_prompt", "evidence": "system_prompt=\"시나리오 분석 전문가. 추출된 요소 목록의 정확성을 평가한다.\"", "line_end": 315, "line_start": 315, "recommended_fix": "Move the system prompt string to a dedicated prompt file (e.g., system_review.md) and load it using _load_prompt.", "severity": "P1", "why_problematic": "LLM system instructions are hardcoded in Korean directly in the Python logic. This bypasses the prompt management system and makes localization or prompt iteration difficult."}, {"category": "scenario_dependent_prompt", "evidence": "system_prompt=\"시나리오 분석 전문가. 요소별 시각적 상세 정보를 최대한 많이 추출한다.\"", "line_end": 367, "line_start": 367, "recommended_fix": "Move the instruction to an external prompt template.", "severity": "P1", "why_problematic": "Hardcoded Korean system instructions for the detail extraction step. This is scenario-dependent pollution in the code."}, {"category": "scenario_dependent_prompt", "evidence": "system_prompt=\"시나리오 분석 전문가. 요소별 시각적 상세 정보를 최대한 많이 추출한다.\"", "line_end": 393, "line_start": 393, "recommended_fix": "Move the instruction to an external prompt template.", "severity": "P1", "why_problematic": "Duplicate hardcoded system instruction for the retry logic, further polluting the code with scenario-specific strings."}, {"category": "scenario_dependent_prompt", "evidence": "\"\\n\\n[이전 에피소드에서 추출된 기존 요소]\\n\" ... + \"\\n기존 요소가 이번 에피소드에도 등장하면 이름을 동일하게 유지하세요.\\n\"", "line_end": 254, "line_start": 251, "recommended_fix": "Move the instruction text to the prompt template and use a placeholder for the JSON data.", "severity": "P1", "why_problematic": "Prompt fragments and specific instructions regarding entity naming consistency are constructed via string concatenation in code. This logic should reside in the prompt template."}, {"category": "schema_or_enum_drift", "evidence": "_MODULE = \"entity_extractor_v2\"", "line_end": 25, "line_start": 25, "recommended_fix": "Update _MODULE to \"entity_extractor_v3\" and ensure corresponding prompt files exist in the correct directory.", "severity": "P2", "why_problematic": "The module constant refers to 'v2' while the file is 'v3'. This causes the prompt loader to fetch templates from the v2 directory, leading to potential version drift or use of outdated instructions."}, {"category": "scenario_dependent_code", "evidence": "f\"{e['name']}({e['appearances']}회)\"", "line_end": 301, "line_start": 301, "recommended_fix": "Move the formatting logic or the suffix into the prompt template or a localization-aware utility.", "severity": "P2", "why_problematic": "Hardcoded Korean suffix '회' (times/appearances) used for formatting entity data before sending it to the LLM. This is language-dependent formatting in the pipeline logic."}], "path": "backend/app/modules/pipeline/entity_extractor_v3.py", "scan_kind": "python", "sha256": "c986035718b7f91b3dc9915305b09b38345443d8f0dfc9fcb7f037ea64709b33"}
{"candidate_reason": "python scope discovery", "chunk_end": 37, "chunk_start": 1, "chunk_summary": "The module hardcodes a Korean user prompt instruction within the Python logic instead of using the prompt loader, creating scenario-dependent debt.", "duration_ms": 13125, "findings": [{"category": "scenario_dependent_prompt", "evidence": "f\"아래 시나리오를 {max_length}자 이내로 요약해주세요.\\n\\n\"", "line_end": 26, "line_start": 23, "recommended_fix": "Move the user prompt template to the prompt repository and load it using load_prompt(_MODULE, 'user', max_length=max_length).", "severity": "P1", "why_problematic": "The user prompt instruction is hardcoded in Korean and assumes the input is a 'scenario'. This bypasses the prompt management system (prompt_loader) used for the system prompt, leading to scattered domain nomenclature and making the pipeline less flexible for non-scenario text or different languages."}], "path": "backend/app/modules/pipeline/episode_summarizer.py", "scan_kind": "python", "sha256": "ab7011132784352e2ada55950c411304521d1c9cde4bb6212c2f263baedef8b5"}
{"candidate_reason": "python scope discovery", "chunk_end": 284, "chunk_start": 1, "chunk_summary": "The file implements floor plan prompt generation and validation using structured schema injection and ID-set verification, with no actionable findings of semantic string judgment or scenario pollution.", "duration_ms": 15279, "findings": [], "path": "backend/app/modules/pipeline/floor_plan_prompt.py", "scan_kind": "python", "sha256": "2b5573a261a1a27f638ee9a2d6e993d0b343f3e50ac66ccf20fdf19f37421029"}
{"candidate_reason": "python scope discovery", "chunk_end": 43, "chunk_start": 1, "chunk_summary": "The module performs entity review by formatting extracted entities and scenario text into a prompt for an LLM, with no actionable findings regarding semantic string judgment or scenario pollution.", "duration_ms": 15623, "findings": [], "path": "backend/app/modules/pipeline/entity_reviewer.py", "scan_kind": "python", "sha256": "fcef947cadd9d7c4eb00bbed3df22fa77435056e08c2d59951dbad87a87a800e"}
{"candidate_reason": "python scope discovery", "chunk_end": 115, "chunk_start": 1, "chunk_summary": "no actionable findings", "duration_ms": 9131, "findings": [], "path": "backend/app/modules/pipeline/floor_plan_render.py", "scan_kind": "python", "sha256": "ec4b73480bd0a2ebd76fd42e029f2144bc5201052901c84bd988f77964836551"}
