Source-linked AI summary
Beyond Starry Night: Shortcut-Aware Control-State Planning for Artist-Grounded Text to Image Generation
Kuan Xing, Ye Wang, Changyi Gan, Yuheng Li, Thao Nguyen, Yi Chang, Yilin Wang
TL;DR
Artist-grounded image generation은 요청한 장면을 정형화된 작가 shortcut으로 대체할 수 있다. Atelier는 명시적이고 근거에 기반한 control을 계획하며, 평가된 생성기 전반에서 style fidelity, source-structure preservation, shortcut avoidance를 향상한다.
문제
작가명 prompt만으로는 장면 보존, 양식적 결정, 부적절한 모티프 회피가 충분히 구체화되지 않아 artist-grounded generation에서 정형화된 shortcut 대체가 체계적으로 일어난다.
방법
Atelier는 구조화된 control state를 추론하고, 작가 근거를 장면 역할에 결부하며, generation plan을 컴파일하고, 전역 및 국소 authenticity feedback을 사용해 출력을 수정한다.
결과
Atelier는 평가된 artist-grounded generation 설정 전반에서 작가 수준의 style fidelity와 source-painting resemblance를 향상하는 동시에 shortcut 대체를 줄인다.
시사점 및 한계
Artist-grounded generation은 의도한 장면 구조를 양식적 변환 및 shortcut 회피와 구분하는, 명시적이고 역할에 결부된 artistic control의 이점을 얻는다.
시사점 및 한계
ArtIntentBench는 신뢰할 수 있는 artist-grounded evaluation에 선별된 근거와 인간 검증이 필요하기 때문에 현재 Van Gogh와 Qi Baishi만 다룬다.
Abstract
from arXiv · showhide
Artist-grounded image generation requires more than appending an artist name to a prompt. Image models often respond to artist names through canonical shortcuts, such as recurring motifs, generic palettes, or overrepresented period signatures, rather than preserving the user's intended scene. We introduce Atelier, a shortcut-aware control-state planning framework for artist-grounded image generation. Atelier translates underspecified artistic intent into an explicit control state that separates scene anchors, preserve/transform decisions, style-regime hypotheses, role-bound artist evidence, and shortcut-avoidance constraints. It grounds this state using artist-level knowledge and local patch references, compiles backend-aware generation plans, and iteratively refines candidates through global and local authenticity feedback. We further introduce ArtIntentBench, a benchmark covering Van Gogh and Qi Baishi across artwork re-rendering, period/style-controlled generation, historically unseen subjects, shortcut auditing, and human preference evaluation. Across open-weight and closed-source generators, Atelier improves artist-level style fidelity, preserves source structure more faithfully, and substantially reduces shortcut substitution compared with prompt-engineered, retrieval-augmented, and general-purpose agent baselines. These results suggest that artist-grounded generation is bottlenecked not only by image synthesis, but by the upstream inference of explicit, evidence-grounded artistic controls.
1 서론
Artist-grounded generation은 모델이 장면별 예술적 결정을 반영하는 대신 정형화된 artist shortcut으로 대체할 때 실패한다. Atelier는 이를 language-to-control translation 문제로 다루며, synthesis 전에 명시적이고 근거에 기반한 control을 추론하고 global 및 local authenticity feedback으로 출력을 정제한다.
- 문제: Canonical shortcutting은 요청된 장면이 뒷받침하지 않는 고빈도 motif나 처리를 artist conditioning이 삽입하는 체계적 실패다.이 논문은 이러한 shortcut을 타당한 canonical association과 구분하고, 이를 맥락에 민감한 visual language가 아니라 허위 단서에 의존하는 현상으로 규정한다.
- 문제: 이 framework는 synthesis만으로는 안정적으로 추론되지 않는 artistic decision을 대상으로 하며, 여기에는 시대, 구도, motif 선택, palette, brushwork, material treatment, local texture가 포함된다.의도한 출력은 요청된 장면을 보존하면서 이러한 artist-specific decision을 명시적으로 반영해야 한다.
- 기여: Atelier는 불충분하게 명세된 artist-grounded 요청을 scene anchor, preserve/transform policy, style-regime hypothesis, retrieval target, anti-shortcut constraint를 포괄하는 명시적 control state로 변환한다.artist-level knowledge와 local patch evidence를 검색하고, patch를 scene role에 결속하며, backend-aware plan을 컴파일하고, candidate를 반복적으로 정제한다.
- 평가: 평가는 intermediate control-state quality와 final-image quality를 분리하고, global assessment와 local authenticity diagnosis를 결합한다.이는 그럴듯한 전체 이미지에 가려질 수 있는 실패, 즉 brushwork, motif treatment, material texture, period-specific surface language의 오류를 포착한다.
- 평가: 서론은 Van Gogh와 Qi Baishi에 대해 frozen human-validated shortcut taxonomy를 제안하여 open-weight 및 closed-source generator 전반의 shortcut substitution을 측정한다.감사는 neutral prompt, 세 명의 annotator에 의한 독립 검토, majority voting을 사용하며, 최소 두 명의 annotator가 해당 pattern을 선택했을 때 이를 계수한다.
2 방법
Atelier는 Shortcut-Aware Artistic Control Planning (SACP)을 폐루프 시스템으로 구현해, 불충분하게 명시된 예술적 요청을 명시적이고 근거에 기반한 제어로 변환한다. 역할이 부여된 전역 및 국소 artist evidence를 검색하고 backend-aware plan을 컴파일하며, 전역 및 patch-level authenticity feedback을 사용해 후보를 반복적으로 개선한다.
- 구조화된 control state: Control state는 장면 해석, 보존-변환 결정, 역사적 또는 style-regime 의도, 검색 및 shortcut 방지 제약, backend 실행 선호를 분리한다.장면 해석은 객체, 영역, 관계, 분위기를 포착하며, style-regime 필드는 Van Gogh의 시기나 Qi Baishi의 모티프, 구도, 수묵, 제문, 낙관 단서를 인코딩할 수 있다.
- SACP와 Atelier: Atelier는 불충분하게 명시된 요청을 명시적 control state로 변환하고, 역할이 부여된 artist evidence를 검색하며, backend-aware plan을 컴파일하고, critic feedback을 통해 후보를 반복적으로 개선한다.폐루프는 장면 해석, evidence retrieval, planning, execution, evaluation, memory updates, final selection으로 구성된다.
- 폐루프 refinement: 후보 이미지는 global critic과 학습된 patch-level authenticity critic인 AuthCritic이 평가하며, 이들의 feedback은 제약, 선호, memory, 이후 refinement round를 갱신한다.AuthCritic은 먼저 낮은 점수의 patch를 식별하고, holistic evaluation은 후보 선택과 수용 결정에 기여한다.
- Artist evidence retrieval: Knowledge base는 artist-level style prior, 선별된 역사 요약, 의미적 장면 역할에 결부된 주석이 달린 국소 patch를 결합하며, 이를 일반적인 style example로 덧붙이지 않는다.전역 reference는 시기, 모티프, 구도, palette prior를 제공하고, 국소 patch는 subject, background, material, texture와 같은 역할을 구체화한다.
- Planning과 execution: Planner는 scene anchor, style target, reference binding, negative shortcut constraint, backend-specific parameter를 포함하는 backend-aware generation plan을 생성한다.Executor는 지원되는 경우 이러한 제어를 reference, adapter, 또는 LoRA-style resource에 매핑하고, general-purpose backend에서는 간결한 textual realization으로 매핑한다.
3 실험
작품 재렌더링, 역사적으로 미등장한 주제 생성, 아티스트 간 전이에 걸쳐 Atelier는 소스 또는 현대적 주제의 구조를 보존하면서 더 높은 스타일 충실도를 달성한다. 감사 결과 control state grounding이 개선되고 shortcut substitution이 감소했지만, patch-bank coverage는 여전히 핵심 한계로 남는다.
- 작품 재렌더링과 아티스트 간 전이: Atelier는 open-weight와 closed-source 설정 모두에서 Van Gogh 재렌더링과 Qi Baishi 전이에 대해 가장 낮은 IntroStyle W2를 달성했으며, Qi Baishi의 DINO-cosine과 LPIPS에서도 선두를 차지한다.Van Gogh 재렌더링 점수는 open-weight에서 73.52, closed-source에서 64.22이며, Qi Baishi 전이 점수는 각각 68.33과 57.09다.
- 역사적으로 미등장한 주제 생성: 역사적으로 미등장한 현대적 주제에서 Atelier는 두 설정 모두 가장 낮은 IntroStyle W2를 달성하며, direct 및 agent-based baseline보다 식별 가능한 주제 구조를 더 일관되게 보존한다.W2 점수는 closed-source에서 86.61, open-weight에서 91.28이며, agent baseline은 현대적 요소를 시대에 적합한 등가물로 대체하는 경우가 많다.
- Control-State 감사: Control-state audit는 benchmark의 각 task 설정에서 시대 회복, preserve/transform decomposition, local patch-role binding, identity protection, anti-shortcut check를 검증한다.시대 정보가 구조화된 경우 일치율은 perfect이며, caption-based inference는 open-weight에서 40.0%, closed-source에서 38.0%의 top-1 recovery에 도달하고, top-2 recovery는 66.0%와 62.0%다.
- Human Evaluation: Human evaluation에서 역사적으로 미등장한 주제 composite가 모든 패널에서 가장 높은 평가를 받는 반면, Van Gogh 재렌더링에서는 소스 보존이 중요하고 Qi Baishi 성능은 주로 모티프와 구성에 의해 제한된다.미등장 주제 composite 점수는 open-weight에서 8.30–9.34, closed-source에서 7.64–8.98이며, Van Gogh 소스 보존 점수는 최소 8.4인 반면 Qi Baishi composite는 5.81–6.65와 4.69–6.10 범위다.
- Shortcut 감사와 한계: Shortcut audit는 baseline의 shortcut rate가 높음을 보여주지만, Atelier의 남은 약점은 binding이나 perceiver failure가 아니라 knowledge-base coverage다.Qi Baishi baseline SSR은 closed-source 시스템에서 55.00%–78.75%, open-weight direct baseline에서 40.00%–53.75% 범위이며, patch-evidence usefulness가 human-rated dimension 중 가장 약하다.
4 관련 연구
기존 연구는 style transfer, personalization, agentic generation, evaluation, shortcut mitigation을 다루지만, 제공된 본문은 모호한 예술적 의도를 구조화된 control로 변환하고 국소적 예술적 진정성을 진단하는 데 아직 해결되지 않은 과제가 있음을 지적한다. 기존 접근법은 대체로 고정된 style 통계, 전역 평가 신호, 또는 특정 training-stage confound를 대상으로 하며, 완전한 artist-grounded generation 문제를 다루지는 않는다.
- 예술적 style transfer: Style-transfer 방법은 exemplar feature 통계 또는 작가 작품 컬렉션을 사용해 콘텐츠를 렌더링하며, ChipGAN 은 중국 수묵화에 medium-specific 제약을 추가한다.이러한 접근법은 주로 고정된 style target의 표면 통계를 전이한다.
- Personalization 및 style-conditioned Image Editing: Personalization 및 style-conditioned generation은 지정된 subject, concept, 또는 style에 맞게 모델을 조정한다 [42]. 그러나 모호한 예술적 의도를 구조화된 artist-grounded control로 변환하지는 못한다.Instruction-based editing은 자연어 instruction과 구조적 조건도 사용한다.
- Agentic 및 tool-using generation system: Language agent는 decomposition, tool, memory, iterative feedback을 사용하며 [34] [41], image-generation 및 editing agent는 perceive–plan–execute–evaluate loop를 적용한다 [48] [51].이러한 framework는 복잡한 image-generation 및 editing task를 위한 closed-loop 지원의 동기를 제공한다.
- Image generation 및 artistic authenticity 평가: 전역 image-generation metric은 품질, alignment, 또는 preference를 평가하지만, brushwork, motif, material texture와 관련된 국소적 실패에는 적합성이 낮다.강력한 multimodal model에서도 예술 특화 판단은 여전히 어렵고, patch-level VLM style concept는 미술사학자의 판단과 부분적으로 부합한다 [26].
- Text-to-image generation의 shortcut: Shortcut learning은 알려진 deep-network 실패 양상이며, Goyal et al. 은 confounding-attribute pathway를 노출해 adapter training 중 personalization shortcut을 완화한다.이들의 설정은 target identity와 함께 pose, expression, lighting과 같은 confound를 대상으로 한다.
5 논의
Artist-grounded 평가를 확장하기 어려운 이유는 신뢰할 수 있는 검증이 자동 metric이 충분히 포착하지 못하는 시대 충실도, 모티프 적절성, 붓질, 재료 처리, shortcut 대체까지 평가해야 하기 때문이다.
- 5 논의: ArtIntentBench는 artist-grounded 검증에 일반적인 자동 metric이 신뢰성 있게 포착할 수 없는 정교한 판단이 필요하기 때문에 깊이 우선 평가 설계를 채택한다.평가는 이미지 유무에만 의존하지 않고 시대 충실도, 모티프 적절성, 붓질, 재료 처리, shortcut 대체를 대상으로 한다.
한계
이 벤치마크는 신뢰할 수 있는 artist-grounded 평가에 필요한 상당한 큐레이션과 전문가 검증을 반영해 두 예술가만 다룬다.
- 벤치마크의 두 예술가 범위는 의도적이다. 신뢰할 수 있는 평가에는 엄선된 예술가 지식, patch-level evidence, 그리고 전문가 또는 미술 교육을 받은 사람의 검증이 필요하기 때문이다.
윤리 및 저작권 성명
이 연구는 퍼블릭 도메인 또는 미술관이 보유한 미술 자료를 사용하며, 저작권 이미지를 재배포하지 않고 코드, 메타데이터, 프롬프트, 스키마, 평가 프로토콜을 공개해 재현성을 지원한다.
- 윤리 및 저작권 성명: Van Gogh corpus는 퍼블릭 도메인 복제본을 사용하고, Qi Baishi corpus는 중국 전통 미술의 공공 미술관 디지털 컬렉션을 사용한다.자료 출처에는 WikiArt, National Gallery of Art, 공공 미술관 컬렉션이 포함된다.
- 윤리 및 저작권 성명: 저자들은 코드, 메타데이터, benchmark 프롬프트, control-state schema, 평가 프로토콜을 공개하되, 저작권이 있는 원본 이미지는 재배포하지 않고 링크를 통해 참조한다.공개된 메타데이터는 재현성을 지원하기 위한 것이다.
Control-State Walkthrough
부록에서는 기록된 두 요청을 통해 Atelier의 control state를 살펴본다. 한 요청은 다섯 구성요소 모두를 드러내고, 다른 요청은 세 차례의 refinement round와 memory-mediated critic feedback을 추적한다. 또한 state가 하나의 JSON object로 직렬화되는 대신 runtime artifact 전반에 분산되는 방식도 설명한다.
- Appendix A: 이 walkthrough는 하나의 held-out request에서 z = (s, q, h, r, b)의 다섯 구성요소 모두를 제시하고, 다른 요청에서는 세 차례의 refinement round를 따른다.두 번째 trace에서는 critic feedback이 reflective memory를 통해 이후 plan으로 전달되는 과정을 보여준다.
- Runtime representation: Runtime에서 z는 하나의 JSON object로 직렬화되지 않고 scene specification, world model, knowledge_evidence block, compiled plan에 분산되어 있다.Appendix D, Table 14에서는 이러한 artifact 간 field-level mapping을 제시한다.
- Runtime representation: 발췌문은 control-state component별로 기록된 field를 묶되, global-critic score sg를 나타내는 heavy_judge_score와 같은 implementation name은 유지한다.생략된 부분은 ....로 표시한다.
A.1 단일 라운드 Control State
Atelier은 정물 요청을 장면 역할, 보존 구조, 스타일 변환, 시대 라우팅, 검색 근거, shortcut 제약을 분리한 명시적 control state로 변환한다. 그런 다음 구조적 anchor를 보존하고 역할 수준의 스타일 지침을 전파하면서, 이 state를 backend별 generation call 3개로 컴파일한다.
- A.1 단일 라운드 Control State: Scene reader는 요청을 object still life로 분류하고, 전경 질량, 장 표면, 배경 halo 역할에 서로 다른 translation axis를 할당한다.axis는 표면 부조, palette 관계, contour 압력, brush 리듬, 방향성 motion을 포괄한다.
- A.1 단일 라운드 Control State: Control state는 object-silhouette integrity와 tabletop support plane을 보존하는 동시에, 역할별 surface, palette, contour, brush, motion 속성을 변환한다.
- A.1 단일 라운드 Control State: 요청이 Paris period를 명시적으로 지정하므로 period policy는 단일 Paris-period route를 사용하며, retrieval은 hard filtering이 아니라 scoring을 통해 인접 period family를 선택할 수 있다.선택된 bound family는 Nuenen period 아래에 색인되지만, generation route는 Paris로 유지된다.
- A.1 단일 라운드 Control State: Retrieval은 foreground-object role을 patch family에 결속하지만, field-surface와 background-halo role은 work-level evidence에 의존하도록 둔다.foreground family에는 대표 patch patch_03958가 포함되며, 다른 role은 unbound-family fallback을 사용한다.
- A.1 단일 라운드 Control State: Planner는 canonical Van Gogh shortcut에 대한 negative prompt를 포함해, 공유 구조 및 스타일 제약과 함께 state를 N = 3 parallel calls로 컴파일한다.계획은 remote_flux_lora, qwen_t2i, longcat_t2i를 사용하며, preserve target과 역할별 translation axis를 backend text에 전달한다.
A.2 3라운드 Refinement Trace · B Prompt Library
Qi Baishi Refinement Trace는 고정된 backend에서 memory-mediated prompt revision을 통해 critic score가 향상되지만 acceptance threshold에는 도달하지 못하는 과정을 보여준다. Prompt Library는 구조화된 scene extraction, clarification, intent inference, shortcut-avoidance rules를 통해 Atelier의 evidence-grounded control state를 구현한다.
- A.2 3라운드 Refinement Trace: Global-critic score는 47에서 49, 57로 상승하고 AuthCritic aggregate는 42, 67, 67에 도달한다. FinalSelect는 Round 3을 선택하지만, episode는 여전히 τ ⋆= 85 미만이며 recommendation revise로 종료된다.Round 1 이후 Hunyuan이 고정되므로, 이 trace는 backend 변경이 아니라 iterative execution-prompt refinement를 분리해 보여준다.
- A.2 3라운드 Refinement Trace: 첫 번째 round의 critic은 약한 calligraphic energy, 제한적인 ink-density transition, generic decorative-illustration drift, 지나치게 균일한 brushwork를 식별하고, canonical failure tags와 국소 repair target을 저장한다.보존된 강점은 subject count, 왼쪽의 inscription과 seal, blank-paper reserve다.
- A.2 3라운드 Refinement Trace: 후속 prompt는 이러한 결과를 dry-brush fraying, graded 또는 translucent wash, ink underpainting, pressure-varying calligraphy, abbreviated duck form과 같은 material and stroke instructions로 변환하면서 subject와 layout을 유지한다.Round 2는 composition을 유지하면서 구체적인 repair instruction을 추가하고, Round 3은 ink gradation과 translucent petal color를 더해 이를 이어간다.
- B Prompt Library: Prompt Library는 Atelier의 perceiver, planner, global critic, local AuthCritic, control-state audit judge에 배포된 deployed templates를 재현하며, 두 audit listing은 보고된 analysis에 포함된 field로 제한된다.이 listing은 구현 prompt text를 보존하면서 deployed component interface를 문서화한다.
- B.1 Perceiver Prompt: Perceiver prompt는 raw request를 regime, subject, required and forbidden entities, scene roles, period preference, confidence, reasoning note를 포함하는 minimal JSON scene specification으로 변환한다.Clarification prompt는 추가로 follow-up question이 유용한지 판단하고, 요청된 subject를 대체하지 않으면서 암묵적인 reference intent를 추론한다.
- B.1 Perceiver Prompt: Perceiver rule은 artist-associated motifs가 아니라 user request를 따르도록 output을 요구하고, artist name이나 canonical resemblance에서 period를 추론하는 것을 금지하며, forbidden entity가 positive cue가 되지 않도록 한다.Scene regime과 role은 명시적인 framing과 compositional organization에 따라 closed catalog에서 선택해야 한다.
- B.1 Perceiver Prompt: Clarification policy는 mood, setting, expressive goal, life context, explicit period preference처럼 artistic control을 개선하는 누락 요인에 대해서만 질문하며, period signal은 artist-specific preference로 취급한다.이는 low-intrusion personalization을 지원하고 implicit work-family reference와 사용자의 실제 subject를 구분한다.
C 평가 기준 및 선택 정책
이 부록은 수치 임계값, 중단 및 선택 규칙, critic 입력 계약을 포함한 Atelier의 폐루프 평가 정책을 정의한다. 이 규칙은 §3.3에 보고된 모든 실험에서 고정된다.
- 이 부록은 Algorithm 1에서 사용하는 폐루프 중단 조건, 라운드별 winner 규칙, trajectory 수준의 FinalSelect 규칙을 명시한다.또한 Atelier의 두 critic에 대한 입력 계약을 정의한다.
- 모든 임계값과 정책 규칙은 §3.3에 보고된 모든 실험에서 고정된다.
- Algorithm 1과 이러한 규칙은 함께 Atelier의 전체 구성을 이룬다.
C.1 Runtime Constants · C.2 Global-Score Aggregation
Runtime policy는 Section 3.3의 모든 episode에 동일한 threshold, budget, routing cardinality를 적용하며, backend locking은 round-one quality에 따라 결정된다. Global score는 8개 critic dimension과 artifact penalty를 결합한 뒤 recommendation cap과 confidence-aware calibration을 거친다.
- C.1 Runtime Constants: Runtime threshold, budget, routing cardinality는 Section 3.3의 모든 episode에 일관되게 적용된다.이 constant들은 pool-and-lock policy와 함께 정의된다.
- C.1 Runtime Constants: Round 1에서는 track별 backend pool을 탐색한 뒤, 이후 round에서 사용할 winner를 lock한다.따라서 이 policy는 초기 track별 탐색 이후 backend 선택지를 좁힌다.
- C.1 Runtime Constants: 선택된 outcome이 low-quality일 때, 즉 sg < τℓ이거나 reject이거나 low-confidence uncertain recommendation일 때 backend lock이 해제된다.추가적인 60-point lock gate는 사용하지 않는다.
- C.2 Global-Score Aggregation: Global critic은 8개의 subscore xj ∈[1, 5]와 artifact penalty a ∈[0, 3]을 반환한다.이 값들이 base-score computation의 입력을 구성한다.
- C.2 Global-Score Aggregation: Score weight는 style authenticity와 period match에 각각 18을 우선 배정하고, intent preservation, brushwork directionality, impasto texture에는 각각 12를 배정한다.Motif match와 composition match는 각각 10을, palette match는 8을 받는다.
- C.2 Global-Score Aggregation: Base score는 reject일 때 59, uncertain일 때 69, revise일 때 84로 cap한 뒤, confidence, blocking tag, actionable gap, request-anchor violation을 사용해 calibration한다.이렇게 얻은 calibrated value가 downstream policy에서 사용하는 score sg다.
C.3 종료 조건 · C.4 후보 선택
배포된 policy는 성공, plateau, rejection, budget, error 조건을 정해진 순서로 적용해 episode를 종료한 뒤, trajectory 전체의 evidence와 fallback ranking rules를 사용해 최종 candidate를 선택한다. 라운드 내 선택에도 global scores, AuthCritic evidence, backend preferences, 가능한 경우 candidate order가 반영된다.
- C.3 종료 조건: policy는 고정된 순서로 terminal conditions를 적용하며, score, recommendation, canonical failure-tag overlap을 사용해 episode 진행을 평가한다.overlap은 |A ∩ B|/|A ∪ B|로 정의되며, empty-set overlap은 1이다.
- C.3 종료 조건: Success termination에는 두 success terms가 모두 필요하며, hard near-threshold plateau는 revise recommendation이 three consecutive인 뒤 감지된다.제공된 passages는 두 success terms가 모두 필요하다고 명시하지만 그 formulas는 제시하지 않는다.
- C.3 종료 조건: 일반적인 four-revision plateau guard는 보고된 budget T = 3에서 도달할 수 없는 반면, two-round near-plateau는 terminal로 처리하지 않고 risk로 기록된다.일반 guard에는 score span at most 2, mean tag overlap at least 0.6, maximum score below τ⋆−3이 필요하며, 기록된 risk는 span at most 1, overlap at least 0.5, maximum score at least τ⋆−3을 사용한다.
- C.3 종료 조건: episode는 current improvement 없이 at least two recent low-score rejects가 발생하거나, t = T에 도달하거나, 복구할 수 없는 retry failure가 발생한 뒤에도 종료되며, 이후 FinalSelect가 y⋆를 반환한다.reject condition은 59 이하의 scores를 사용하며, FinalSelect는 completed trajectory에서 동작한다.
- C.4 후보 선택: 사용 가능한 AuthCritic sidecar가 있으면, 라운드 내 선택은 score, real-patch probability, candidate index를 포함해 global-critic과 AuthCritic의 candidate evidence를 결합한다.해당 passage는 sg(y), sa(y), pa(y), i(y)를 식별하지만, selection formula는 제공된 text에 포함되어 있지 않다.
- C.4 후보 선택: 사용 가능한 sidecar가 없으면, candidates는 global score, accept status, confidence, backend preference, earlier index 순으로 정렬되며 backend-specific tie-breaking이 적용된다.Low-quality exact ties에서는 non-FLUX candidate를 우선하며, 그 외에는 default backend tiebreak가 FLUX를 선호한다.
- C.4 후보 선택: FinalSelect는 stored second-stage scores와 generation rounds를 사용해 full trajectory를 평가하며, 가능한 경우 local evidence를 사용하고 그렇지 않으면 lexicographic fallback을 적용한다.사용 가능한 local evidence가 없으면 (sg(y), s2(y), −t(y))로 순위를 매긴다. winner가 rejected, low-scoring, 또는 uncertain and low-confidence인 경우에도 rounds는 multi-backend exploration에 남는다.
C.5 AuthCritic 출력 계약 및 에스컬레이션
AuthCritic은 grounded reasoning model로 작동하며 각 patch에 대해 엄격한 순서의 JSON record를 출력한다. 점수가 낮은 patch는 crop 및 명시된 원인과 함께 global critique로 에스컬레이션되며, candidate당 최대 네 patch로 제한된다.
- AuthCritic 출력 계약: AuthCritic은 source label, grounding field, visual 및 style description, 짧은 reason, real master와의 차이를 포함하는 순서가 있는 JSON object를 출력해야 한다.필수 key는 source_type, grounded_title, grounded_period, source_global_summary, patch_visual_description, patch_style_description, reason_short, difference_from_real_master다.
- AuthCritic 출력 계약: source_type label은 real-master, synthetic-master-style, other-painter patch를 구분하며, non-real-master record는 brushwork, palette, contour, surface 또는 motion을 통해 이탈 양상을 설명한다.real_vangogh_patch 같은 artist-specific alias는 parsing 과정에서 정규화하며, real-master prediction은 작품 title과 period에 근거를 둔다.
- 에스컬레이션: τp 또는 πp보다 낮은 점수를 받은 patch는 crop 및 reasoning record와 함께 global critique의 focus signal로 에스컬레이션하며, candidate당 최대 Kp = 4개 patch를 score 오름차순으로 정렬한다.이는 2단계 평가에서 의심 영역을 명시된 원인과 연결한다.
C.6 AuthCritic 후보 집계 · C.7 Global-Critic 입력 투영
AuthCritic은 patch-level authenticity evidence를 후보 점수로 집계하고, global critic은 의도적으로 투영된 request-state view와 선택된 저점수 patch evidence를 입력으로 받는다. 이 투영에서는 전체 state와 evidence bundle을 제외하며, 나머지 필드는 retrieval, planning, generation을 조건화하는 데 사용된다.
- C.6 AuthCritic 후보 집계: AuthCritic은 local style-evidence patch가 하나라도 있으면 context-only crop을 제외하고, 그렇지 않으면 샘플링된 모든 patch를 점수화한다.
- C.6 AuthCritic 후보 집계: 후보 점수는 scoring patch에 대한 real-master fraction, mean proximity-to-real, validity-related term을 사용한다.변수에는 fr, mean synthetic-master proximity, valid-proximity fraction, patch count, proximity parse-error rate가 포함된다.
- C.6 AuthCritic 후보 집계: valid proximity를 가진 synthetic-master patch가 없으면 mean proximity-to-real은 모든 valid proximity의 평균으로 falls back한다.
- C.6 AuthCritic 후보 집계: 연관된 candidate probability는 real-master fraction이며, proximity 필드와 parse status가 누락되면 legacy rounded-mean fallback을 트리거한다.legacy fallback은 patch별 source-type score의 평균을 계산한다.
- C.7 Global-Critic 입력 투영: global critic prompt는 raw 및 clarified request, working period hypothesis와 alternatives, 선택된 motif 및 composition cue, 저점수 patch metadata를 투영한다.
- C.7 Global-Critic 입력 투영: 후보 이미지와 flagged patch crop은 시각적으로 첨부되며, remaining state and evidence fields는 critic prompt에 들어가지 않은 채 retrieval, planning, generation을 조건화한다.
D 제어 상태 및 메모리 스키마 참조 … D.3 메모리 구조
부록은 추상 제어 상태 z와 메모리 m의 런타임 스키마를 엄격한 JSON 계약, 정규화, fallback 값, 명시적 필드 매핑을 통해 지정한다. 또한 감사 필드와 계획·평가를 위한 세 구성요소의 라운드 간 메모리를 정의한다.
- D 제어 상태 및 메모리 스키마 참조: 런타임은 프롬프트로 정의된 정규화 및 fallback 대체를 포함하는 strict-JSON contract에 따라 z = (s, q, h, r, b)와 메모리 m을 레코드로 구현한다.Perceiver와 planner 프롬프트는 필수 키를 지정하며, 누락되거나 형식이 잘못된 필드에는 해당 fallback 값이 할당된다.
- D.1 Planner 출력 계약: planner는 여섯 개의 필수 최상위 JSON 키를 출력하고, 런타임은 여기에 파생된 routing_plan을 추가로 기록한다.여섯 키 계약과 정규화 동작은 Table 13에 문서화되어 있다.
- D.2 z = (s, q, h, r, b)의 필드 맵: Table 14는 z의 각 구성요소를 구체적인 런타임 레코드 경로에 매핑하며, knowledge_evidence 아래에는 축약 경로를 사용한다.이 매핑은 추상 상태 구성요소와 런타임 레코드에서 소비되는 필드를 연결한다.
- D.2 z = (s, q, h, r, b)의 필드 맵: 런타임은 고정된 해석 입력으로부터 base s와 q를 재구성하고, 한 라운드 안에서는 planner의 수정을 허용하며, 경로별 병합 동작을 적용한다.Qi Baishi 병합은 권위 있는 scene 및 anti-shortcut 필드를 복원하는 반면, Van Gogh 경로는 존재하는 경우 planner가 출력한 world model을 유지한다.
- D.2 z = (s, q, h, r, b)의 필드 맵: 추가적인 world-model audit fields는 scene-preservation logic, regime hypotheses, role constraints, uncertainty 및 shortcut 관련 제어를 기록하지만, 생성에 직접 입력되지는 않는다.예로는 scene_logic_to_preserve, prompt_scene_regime, allowed_scene_roles, role_coverage_audit, unresolved_uncertainties가 있다.
- D.3 메모리 구조: 영속 메모리 m은 라운드 간에 trajectory, reflective 및 selection 관련 정보를 포함한다.Trajectory memory는 각 라운드의 실행된 계획, 평가 결과, 실패 및 수정 정보, authenticity 점수와 selection record를 저장하며, reflective memory는 다음 Derive 단계에서 읽혀 planner에 제공된다.
- D.3 메모리 구조: Trajectory memory는 각 라운드에 대해 실행된 positive 및 negative prompts, guidance note, critic 결과, failure tags, repair targets, AuthCritic score and yes-probability, 그리고 selection record를 기록한다.이 레코드들은 라운드 간 계획, 평가, 수정, authenticity 판단 및 후보 선택을 추적하는 데 사용된다.