Source-linked AI summary
ITL: Interpretable Document Alignment with Structured Reference Frameworks
Raúl Giráldez, Dayrelis Mena, Jesús S. Aguilar--Ruiz
TL;DR
이질적이고 모호한 텍스트 전반에 증거가 분산되어 있고 기존 접근법의 추적 가능한 개념 수준 측정치가 제한적이므로, 구조화된 프레임워크와 문서의 정렬을 평가하기는 어렵다. ITL은 가중 개념 프로파일과 텍스트 단위–개념 affinity를 구축한다. SDG 일관성 평가에서는 모든 descriptor가 대응하는 개념과 가장 강하게 매칭되었으며, 대각선 밖 affinity는 기준 affinity보다 두 자릿수 낮았다.
문제
구조화된 참조 프레임워크와의 정렬을 평가하기는 어렵다. 문서가 이질적이고, 대체로 비구조적이며, 의미적으로 모호하고, 수작업으로 재현성 있게 평가하는 데 비용이 많이 들기 때문이다.
방법
ITL은 SRD에서 가중된 개념별 용어 프로파일을 유도하고, 모든 텍스트 단위–개념 쌍에 대해 해석 가능한 affinity 값을 계산한다.
결과
SDG 일관성 평가에서 모든 descriptor는 대응하는 개념에 대해 가장 높은 affinity를 받았으며, 대각선 밖 평균 affinity는 평균 기준 affinity보다 두 자릿수 낮았다.
시사점 및 한계
ITL은 SRD가 정의한 개념 프로파일을 구분하면서, 정렬 결과를 해석 가능하게 유지하고 이를 뒷받침하는 용어까지 추적 가능하게 한다.
시사점 및 한계
평가에 SRD를 정의하는 데 사용된 것과 동일한 descriptor가 사용되므로, 보지 못한 표현, 도메인, 언어 전반에 대한 경험적 일반화 가능성은 아직 확립되지 않았다.
Abstract
from arXiv · showhide
Measuring alignment between documents and structured reference frameworks requires identifying conceptual evidence distributed throughout the text and reporting it through measures that are quantitative, interpretable, and traceable. Many commonly used retrieval and classification approaches return either pairwise similarity scores or one or more class labels, whereas fewer methods provide concept-level scores that are directly traceable to the terminological evidence supporting them. We present \emph{Intelligent Target Locator} (ITL), a domain-agnostic and language-portable methodology that estimates the affinity between the textual units of a target document and the concepts defined in a \emph{Structured Reference Document} ($SRD$). From the $SRD$, ITL induces concept-specific terminological profiles built from independent terms, bigrams, trigrams, and co-occurrences. Each term is assigned an importance weight that combines concept membership, term-type specificity and inter-concept discriminability. The output is a textual-unit--concept affinity matrix that can be aggregated at different levels of granularity. We conduct an internal consistency assessment using the 17 Sustainable Development Goals (SDGs), evaluating each official goal statement against the $SRD$ induced from the same set of descriptors. Every statement reached its highest affinity with the corresponding concept, and the mean affinity across the remaining concepts stayed marginal relative to the mean reference affinity. This separation indicates that ITL distinguishes the conceptual profiles of the framework. ITL thus offers a general basis for quantifying document alignment with structured frameworks while keeping each result traceable to the terminological evidence that supports it.
1 서론
이 논문은 구조화된 참조 프레임워크에 대한 문서 정렬을 해석 가능하고 추적 가능하며 비교 가능한 방식으로 측정해야 한다는 요구를 다룬다. 이 과제는 이질적인 비정형 텍스트와 문서 전반에 흩어진 개념적 증거로 인해 복잡해진다. 논문은 SRD를 통해 프레임워크 개념을 모델링하고, 가중 affinity를 계산하며, SDGs를 사용해 접근법을 평가하는 domain- 및 language-independent 방법론인 ITL을 제안한다.
- 문제: 이 방법론은 정렬을 개념적 관계로 다룬다. 그 관계의 증거는 문서 전반에 분산될 수 있으며 감사 가능하고 비교 가능한 상태로 유지되어야 한다.관련 프레임워크 요소를 식별하는 것만으로는 충분하지 않다. 관계를 정량화하고 이를 뒷받침하는 증거까지 추적해야 하기 때문이다.
- 방법: ITL은 Structured Reference Document를 용어 참조로 사용한다. 이는 각 개념이 하나의 참조 텍스트 단위로 기술되는 유한한 개념 집합으로 도메인을 나타낸다.이 접근법은 도메인과 언어에 독립적으로 정식화된다.
- 기여: ITL은 대상 문서가 Structured Reference Document와 중요한 개념을 얼마나 강하게 공유하는지 정량화하는 해석 가능한 affinity measure를 제공한다.이 measure는 가중된 관련 개념에 기반하며, 참조 프레임워크 내에서 해당 개념들의 중요도를 반영한다.
- 방법: 이 논문의 two-phase methodology는 먼저 구조화되고 속성이 부여된 참조 모델을 유도한 뒤, 집계되고 가중된 개념을 기준으로 대상 문서를 정렬한다.이를 통해 참조 모델 유도와 문서 정렬을 분리한다.
- 검증: SDGs는 일관성 연구와 실증 검증을 위한 설정을 제공하며, SRD는 공식 목표 텍스트로 구성된다.이 연구는 유럽 대학의 SDGs에 초점을 둔 Universities for Sustainable Development 프로젝트에서 시작되었다.
2 관련 연구
선행 연구는 검색, 의미 표현, 용어 추출, 형식적 framework 모델링, 주제 분류를 아우르지만, ITL은 SRD에서 유도한 concept profile과 단계적이고 추적 가능한 concept-level affinity를 결합한다. 따라서 이러한 근거 기반 alignment를 동시에 제공하지 않고 framework를 검색·분류·형식화하는 방법들을 보완한다.
- 관련 연구: 관련 연구에는 BM25 와 같은 lexical retrieval, BERT [3], Sentence-BERT [4], DPR [5], ColBERT [6]를 포함한 contextual 및 dense model, 그리고 terminology·ontology·classification 방법이 포함된다.이들 연구 방향은 각각 문서 비교, semantic matching, 비교 가능한 표현, framework 형식화, category 할당을 다룬다.
- 새로움: ITL은 scoring 과정에서 concept별 term-level evidence를 보존한다. 반면 dense retrieval method의 native score는 일반적으로 concept별로 match를 분해하지 못한다.Dense representation은 ITL 내부의 matching component로 활용될 수 있으므로 상호 보완적이다.
- 용어 표현: ITL의 representation은 independent term, bigram, trigram, co-occurrence를 추출해 terminology를 명시적 SRD concept에 연결하고, specificity와 discriminability에 따라 term에 가중치를 부여한다.이는 사전 정의된 framework와의 명시적 대응 없이 latent topic을 유도하는 topic modeling과 대조된다.
- 새로움: ITL은 고정된 학습 category에서 discrete label이나 probability를 산출하는 automatic classification과 달리, 각 SRD concept에 대해 graded affinity를 산출한다.Classification 접근법은 대체로 annotated data를 요구하며, 할당 결과를 이를 뒷받침하는 evidence로 명시적으로 분해하지 않는다.
- 새로움: ITL은 고정된 학습 category가 아니라 임의의 SRD에서 직접 유도한 terminological profile을 사용해 모든 textual-unit–concept pair에 대한 graded affinity를 계산한다.SRD가 reference model을 제공하며, 결과 score는 해당 문서에 정의된 concept에 계속 연결된다.
3 문제 형식화
ITL은 문서 정렬을 각 대상 문서 단위와 Structured Reference Document가 나타내는 각 개념 사이의 정량적 affinity로 형식화한다. 출력에는 개념별 terminology에 기반한 해석 가능한 단위 수준 및 집계된 개념 수준 표현이 포함되며, SRD 의존적 척도를 고려하기 위해 상대 점수도 제공한다.
- 3.1 문제 진술: 문제는 각 대상 문서 단위 S_i와 각 SRD 개념 c_j 사이의 정렬을 정량화하여 해석 가능한 textual-unit–concept affinity matrix를 생성하는 것이다.이 matrix는 paragraph, section, document 수준으로 집계할 수 있으며 textual-unit–concept heatmap과 같은 시각화도 지원한다.
- 3.2 Terms and Reference Term Sets: 각 SRD 개념에는 해당 reference unit에서 추출한 reference term multiset RTS_j가 할당되며, term multiplicity를 보존하여 개념별 terminological profile을 구성한다.이 profile은 각 target unit과 각 개념 사이의 affinity를 평가하는 기반이 된다.
- 3.3 Term Typology: term representation은 independent terms, consecutive bigrams, trigrams, bounded-window co-occurrences를 일관되게 포함하여 고립된 언어 증거와 복합 언어 증거를 포착한다.Co-occurrences에는 순서가 있는 비인접 token pair의 양방향이 모두 포함될 수 있다.
- 3.4 Discriminability, Specificity, Relevance, and Importance: Co-occurrence pattern space는 관찰된 distinct co-occurrences를 기반으로 정량화된다. 이는 proximity window와 text의 term arrangement 양쪽에 의존하기 때문이다.이는 token 수에 의해 결정되는 bigram 및 trigram pattern space와 다르다.
- 3.6 Absolute Affinity, Reference Affinity, and Relative Affinity: Affinity matrix와 document-level vector의 크기는 SRD에 의존하므로, 이 형식화는 보편적 reference value에 의존하는 대신 absolute and relative affinity vector를 제공한다.이러한 의존성은 terminological richness, 개념 간 term distribution, aggregation operator에서 발생한다.
- 3.7 Model Output and Intermediate Elements: Formal model의 출력은 textual-unit–concept affinity matrix와 집계된 absolute 및 relative concept-affinity vector이며, 문서 정렬을 정량적이고 해석 가능한 표현으로 제공한다.이 출력은 Figure 1에 예시되어 있으며, 전체 notation은 Appendix B에 정리되어 있다.
4 방법
ITL은 문서–개념 affinity 값에서 이를 뒷받침하는 용어까지의 traceability를 보존하는 재현 가능하고 domain- 및 language-independent한 방법론이다. Structured Reference Document에서 재사용 가능한 reference model을 유도하고, 공통 linguistic processing을 통해 분할된 target-document unit에 적용한다.
- 방법 개요: ITL은 domain 및 language independence, terminological traceability를 통한 interpretability, 그리고 unstructured document에 대한 적용을 지원한다.이 설계는 각 affinity value의 근거를 보존하면서 모든 concept-structured framework에서 SRD를 구성할 수 있게 한다.
- 방법 개요: 이 방법론은 SRD에서 reference-model induction을 수행한 뒤, 유도된 model에 대해 document alignment를 수행하는 방식으로 구성된다.reference model은 SRD마다 한 번 계산한 후 여러 target document에서 재사용하며, alignment를 통해 affinity matrix와 document-level aggregation을 산출한다.
- 공통 처리: 두 단계는 tokenization, tagging, filtering, stemming, term-typology extraction을 공유하지만, induction은 concept descriptor를 처리하고 alignment는 target-document textual unit을 처리한다.Figure 2는 이 공통 processing core와 전체 ITL flow를 요약한다.
- Input representation: SRD는 각 concept를 descriptive unit과 연결하고, target document는 textual unit으로 분할하며, unstructured input에서는 segmentation에 앞서 text extraction을 수행한다.이러한 representation은 각각 induction과 alignment의 input을 제공한다.
- Reference-model induction: induction pipeline은 descriptor text를 정규화하고 concept-bearing content를 유지하며, independent term, bigram, trigram 및 bounded-window co-occurrence를 추출한다.independent term은 stem과 part-of-speech identity를 보존하는 반면, compound term은 연결된 stem으로 구성한다. 추출된 frequency는 효율적인 weighting을 위해 concept matrix로 정리한다.
5 일관성 연구: 지속가능발전목표
일관성 연구에서는 17개 지속가능발전목표를 사용해 ITL을 적용하고, 각 목표의 공식 서술을 SRD descriptor이자 평가 단위로 사용한다. 각 서술은 대응하는 목표와 가장 높은 affinity를 보이며 교차 affinity는 미미하게 유지되지만, 통제된 설계로 인해 일반화와 aggregation에 관한 결론은 제한된다.
- Reference instantiation: 17개 공식 SDG 서술 각각은 SRD의 concept descriptor이면서 우세한 concept이 알려진 평가 단위로 기능한다.이 통제된 configuration은 동일한 서술에서 유도된 framework에 대해 textual-unit affinity를 평가한다.
- Reference-model induction: Term importance는 profile membership, term-type specificity, inter-goal discriminability를 결합하며, 더 희귀한 term에는 더 높은 discriminability를, 복합 pattern에는 더 높은 specificity를 부여한다.Term이 특정 goal에 특화되고 다른 goal과 해당 goal을 구별할수록 높은 importance를 갖는다.
- Alignment results: 0.18–0.38의 diagonal affinity와 0.020의 maximum cross-affinity는 모든 서술이 대응하는 goal과 가장 높은 affinity를 보이며, mean off-diagonal affinity가 두 자릿수 낮음을 보여준다.가장 큰 cross-affinity도 최저 reference affinity인 0.18보다 낮게 유지된다.
- Alignment results: Reference affinity의 차이는 terminological richness를 반영한다. 더 긴 descriptor와 더 구체적인 compound pattern은 일반적으로 더 짧은 서술보다 높은 affinity를 산출한다.이 reference affinity는 외부 target document와의 alignment를 정규화할 때 해석의 기준점으로 기능한다.
- Scope and limitations: 서술이 model 유도에 사용된 descriptor를 재현하므로 이 연구는 internal consistency only를 평가하며, 보지 못한 wording, 외부 document 또는 경쟁 method는 검증하지 않는다.외부의 이질적인 document, 체계적 비교, unit이 여러 concept와 align되는 경우는 향후 과제로 남는다.
- Scope and limitations: 17개 서술은 응집된 document가 아니라 이질적인 unit을 이루므로 affinity는 document level로 aggregation하지 않고 per textual unit으로 평가한다.따라서 이 consistency study에서는 absolute 및 relative affinity vector를 계산하지 않는다.
6 결론
ITL은 대상 문서와 구조화된 참조 문서 간 alignment를 정량화하는 해석 가능하고 추적 가능한 affinity metric을 제공한다. 내부 SDG 평가는 대응 개념을 분리했지만, 도메인 간 및 언어 간 일반화는 아직 확립되지 않았다.
- 기여: ITL은 각 textual-unit–concept 쌍에 대해 graded affinity metric을 형식화하여 대상 문서와 Structured Reference Document 간 alignment를 정량화한다.이 방법론은 참조 문서에서 직접 concept-specific terminological profile을 유도한다.
- 방법: Term importance는 profile membership, term-type specificity, inter-concept discriminability를 결합하여 추적 가능한 textual-unit–concept affinity matrix를 생성한다.Concept self-affinity는 해석의 기준점을 제공하며, 값은 서로 다른 수준의 granularity로 집계할 수 있다.
- 내부 일관성 평가: 평균 reference affinity보다 Two orders of magnitude lower한 평균 off-diagonal affinity는 17개 SDG descriptor 전반의 완벽한 대응 개념 순위와 함께 나타났다.각 descriptor는 대응하는 개념에서 가장 높은 affinity를 받았으며, 이는 보지 못한 텍스트에 대한 일반화가 아니라 유도된 profile의 분리를 평가한다.
- 범위와 한계: ITL은 원칙적으로 normative, strategic, regulatory 또는 programmatic framework와 서로 다른 언어로 확장 가능하지만, 경험적 도메인 간 및 언어 간 일반화는 아직 확립되지 않았다.다른 환경에서의 구현에는 해당 언어 처리 자원이 필요하다.
- 적용 가능성: 도메인 및 언어 독립성은 추적 가능한 affinity 값과 함께, 명시적 참조 framework를 기준으로 이질적인 문서 간 alignment를 설명하는 데 기여한다.이러한 적용 가능성은 alignment의 정량화와 설명을 모두 요구하는 시나리오를 대상으로 한다.
- 향후 연구: Future work에서는 외부 문서, 추가 도메인 및 언어에서 ITL을 검증하고, 문서 수준 aggregation strategy를 테스트하며, 다른 alignment approach와 비교할 예정이다.계획된 비교에는 semantic-representation 및 alignment method가 포함된다.
A 17개 지속가능발전목표
유엔 지속가능발전을 위한 2030 의제는 17개 목표로 구성되며, 각 목표의 공식 서술문은 참조 인스턴스화에서 개념 기술자로 원문 그대로 사용된다.
- 지속가능발전을 위한 2030 의제는 17개 목표로 구성된다.
- 각 목표의 공식 서술문은 참조 인스턴스화에서 해당 개념 기술자 R_j로 원문 그대로 사용된다.
- 열거된 목표는 빈곤, 기아와 식량 안보, 보건, 교육, 성평등, 물과 위생을 포함한 주제를 다룬다.
B 표기법 용어집
이 용어집은 문서 정렬을 형식화하는 데 사용되는 표기법을 정의하며, 구조화된 참조 문서, 개념, 텍스트 단위, 용어 및 용어 프로파일을 다룬다. 또한 해당되는 경우 각 요소를 도입하는 방정식이나 절을 제시한다.
- 문서와 텍스트 단위: 표기법은 SRD, 유한한 개념 집합 C, 개념별 참조 단위 R_j, 그리고 대상 문서 D를 중심으로 구성된다.SRD는 대상 의미 영역을 인코딩하는 Structured Reference Document를 나타내며, C = {c_1, ..., c_m}은 m개의 개념을 포함하고 각 개념은 R_j로 기술된다.
- 용어, 어휘집 및 용어 프로파일: 용어는 텍스트 단위에서 추출된 어휘 단위로, 단일어, n-gram 및 동시출현을 포함하며, 빈도와 어휘집 표기법은 중복 출현과 서로 다른 용어를 구분한다.T(X)는 추출된 용어의 멀티셋이고, f(t | X)는 용어 빈도를 나타내며, V(X)는 서로 다른 용어를 포함한다.
- 용어, 어휘집 및 용어 프로파일: 개념별 용어 프로파일은 RTS_j = T(R_j)로 나타내며 TP에 수집된다. V_SRD와 유형별 어휘집은 유도된 용어 전체 집합과 용어 범주를 정의한다.V_IND(R_j)와 V_COC(R_j)는 각각 개념 c_j에서 관찰된 서로 다른 독립 용어와 서로 다른 동시출현을 포착한다.
Term Typology.
ITL은 전처리된 각 textual unit을 independent terms, bigrams, trigrams, bounded co-occurrences의 네 가지 term type으로 나타낸다. Independent-term identity는 stem과 part-of-speech tag를 결합하며, co-occurrences에는 제한된 수의 중간 token이 허용된다.
- Term Typology.: ITL은 normalized token sequence에서 independent terms, consecutive bigrams, consecutive trigrams, bounded co-occurrences를 추출한다.Co-occurrences는 최대 Nsw개의 intermediate token으로 분리된 두 개의 non-consecutive token으로 구성된다.
- Term Typology.: 각 textual unit은 filtering과 linguistic normalization을 거쳐 얻은 normalized tokens의 ordered sequence다.Sequence length LX는 남은 token의 개수를 세며, multiset union은 term multiplicity를 보존한다.
- Term Typology.: Independent-term identity는 token의 stem과 part-of-speech tag의 pair로 정의된다.Term type은 κ(t) ∈ {IND, 2G, 3G, COC}로 나타내며, specificity는 이 집합에 대해 정규화된다.
- Term Typology.: Term typology는 ITL의 reference-model weighting에 반영되며, relevance, type-dependent specificity, cross-concept discriminability가 term importance를 결정한다.Importance φ(t | RTSj)는 intra-concept relevance와 inter-concept discriminability의 곱으로 정의된다.