Source-linked AI summary
InfoLaw: Information Scaling Laws for Large Language Models with Quality-Weighted Mixture Data and Repetition
Fengze Liu, Weidong Zhou, Binbin Liu, Ping Guo, Zijun Wang, Bingni Zhang, Yifan Zhang, Yifeng Yu, Xiaohuan Zhou, Taifeng Wang
TL;DR
데이터가 제한된 LLM pretraining에서는 기존 scaling law가 mixture recipe와 repetition을 가로질러 안정적으로 외삽하지 못한다. InfoLaw는 information accumulation을 모델링해 loss를 예측하며, scale과 overtraining 전반에서 외삽할 때 0.15% 평균 및 0.96% 최대 절대 오차를 달성한다.
문제
데이터가 제한된 LLM pretraining에서 기존 scaling law는 data-mixture recipe와 repetition 전반에 걸쳐 신뢰성이 제한적이다.
방법
InfoLaw는 data quality, mixture weight, train token, model size와 repetition으로 인한 diminishing return에서 학습된 information을 모델링해 loss를 예측한다.
결과
보지 못한 recipe와 더 큰 scale의 run을 예측할 때 0.15% 평균 및 0.96% 최대 절대 오차를 달성했으며, 25× overtraining까지 안정적으로 외삽했다.
시사점 및 한계
InfoLaw는 광범위한 추가 실험 없이 computational budget 전반에서 효율적인 data-recipe 선택을 지원한다.
시사점 및 한계
이 framework는 더 높은 품질의 bucket이 더 큰 quality-density 값을 가진다고 가정하며, fitted quality function에 감소하는 형태를 부과한다.
Abstract
from arXiv · showhide
Upweighting high-quality data in LLM pretraining often improves performance, but in datalimited regimes, especially under overtraining, stronger upweighting increases repetition and can degrade performance. However, standard scaling laws do not reliably extrapolate across mixture recipes or under repetitions, making the selection for optimal data recipes at scaling underdetermined. To solve this, we introduce InfoLaw (Information Scaling Laws), a data-aware scaling framework that predicts loss from consumed tokens, model size, data mixture weights, and repetition. The key idea is to model pretraining as information accumulation, where quality controls information density and repetition induces scaledependent diminishing returns. We first collect the model performance after training on datasets that vary in scale, quality distribution, and repetition level. Then we build up the modeling for information so that information accurately predicts those model performance. InfoLaw predicts performance on unseen data recipes and larger scale runs (up to 7B, 425B tokens) with 0.15% mean and 0.96% max absolute error in loss, and it extrapolates reliably across overtraining levels, enabling efficient data-recipe selection under varying compute budgets.
1. 서론
InfoLaw는 이질적인 데이터 혼합과 반복에 걸친 정보 축적을 모델링해 데이터 제약이 있는 LLM scaling을 다룬다. 보지 못한 recipe, 더 큰 scale, 25× overtraining에서 loss를 예측하며, 최대 7B parameters와 425B tokens에서 0.15% 평균 및 0.96% 최대 절대 오차를 달성한다.
- 서론: 고품질 데이터는 희소한 반면, 이를 과도하게 upweighting하면 특히 overtraining에서 지나친 반복으로 인해 성능이 저하될 수 있다 (Muennighoff et al., 2023; Touvron et al., 2023; Yang et al., 2025).Overtraining은 compute-optimal regime에 비해 inference 비용을 낮추지만 (Hoffmann et al., 2022), 데이터와 반복 사이의 tradeoff를 심화한다.
- 서론: InfoLaw는 학습을 정보 축적으로 모델링하고, mixture-weight scaling과 repetition-dependent diminishing returns를 결합해 데이터 recipe를 결정한다.고품질 데이터는 처음에는 더 큰 이득을 제공하지만, 반복 노출되면 marginal benefit이 보지 못한 저품질 데이터가 제공하는 수준으로 감소한다.
- 서론: 이 연구는 scale, quality, repetition이 서로 다른 LayerMix 데이터셋으로 학습한 252M부터 1.2B parameters 범위의 from-scratch model 9개를 사용해 InfoLaw를 fitting한다.LayerMix는 source dataset을 quality bucket으로 나누고, target scale에 맞게 downsampling한 뒤 서로 다른 weight로 sampling한다.
- 서론: InfoLaw가 보지 못한 recipe, 더 큰 scale, 25× overtraining에서 loss를 예측할 때 0.15% 평균 및 0.96% 최대 절대 오차를 달성하며, 여기에는 425B tokens로 학습한 7B model이 포함된다.평가는 보지 못한 LayerMix mixture weight, 더 큰 compute scale, 더 높은 overtraining ratio를 포괄한다.
2. 관련 연구
선행 연구는 모델 크기, 학습 데이터, compute allocation에 대한 scaling law를 정립한 뒤, 이를 overtraining, 데이터 품질, repetition, mixture 설계로 확장했다. 이러한 연구는 제한적이고 반복되는 데이터 환경에서 data-aware scaling에 초점을 맞추는 InfoLaw의 동기를 제공한다.
- Scaling Laws: Transformer language model은 모델 크기와 학습 데이터에 따라 예측 가능한 power-law scaling을 보이며, dense 및 mixture-of-experts 시스템의 개발을 촉진했다.Compute-based law는 고정된 compute budget에서 모델 용량과 token allocation을 정식화한다.
- Scaling Laws: 더 많은 token으로 소형 모델을 overtraining하는 방식이 일반화된 한편, 이후의 scaling law는 데이터 품질, inference 요구사항, overtrained regime을 포함하도록 확장되었다.Hoffmann et al. (2022)은 compute-optimal training을 규명했으며, Sardana et al. (2024)와 Gadre et al. (2024)는 이 설정을 넘어 분석을 확장했다.
- Data-Aware Scaling: 반복되거나 upsample된 데이터는 diminishing returns와 결국 성능 저하를 초래할 수 있지만, repetition에 대한 학습을 계속하는 것이 조기 중단보다 더 나은 경우도 있다.이러한 효과는 repetition이 단순하지 않으며 고전적인 scaling law로 포착되지 않음을 보여준다.
- Data-Aware Scaling: Scaling-law 방법은 mixture-weight 효과를 예측하고, proxy model로 비율을 탐색하며, model-scale-dependent mixing을 분석함으로써 data recipe를 점점 더 optimize한다.관련 연구는 scaling insight를 continued pre-training과 domain-mixture 설계에도 적용한다.
3. 기존 Scaling Law의 한계
기존 scaling law는 quality-weighted data mixture와 repetition을 함께 사용해 학습할 때, 특히 extrapolation 상황에서 loss를 모델링하기에 불충분하다. 이 절에서는 data scale, quality composition, repetition을 변화시키는 LayerMix를 소개한 뒤, repetition 유무에 따른 scaling behavior를 비교한다.
- 3.1 LayerMix sampling: LayerMix는 두 classifier의 quality score를 사용해 문서를 여섯 개 percentile bucket으로 나누고, quality가 높은 bucket이 더 많이 반영되도록 한다.가장 낮은 quality bucket을 제외한 다섯 가지 preset mixture를 사용하며, 별도 언급이 없으면 K = S로 두어 mixture weight로 유도된 repetition 효과를 분리한다.
- 3.1 LayerMix sampling: LayerMix는 sampling weight, source size, training-set size를 조절해 scale, quality mixture, repetition이 달라지는 packed dataset을 구성한다.bucket 수준 repetition을 R_d = K_d/M_d로 정의하며, sampled token이 모두 unique이면 R_d = 1, 그렇지 않으면 R_d > 1이다.
- 3.2 Loss–compute scaling: 실험에서는 compute-optimal regime과 overtrained regime에서 high-quality sample이 더 많지만 repetition도 더 큰 HQ data와, 더 diverse하고 repetition이 더 작은 sample이 많은 MLQ data를 비교한다.Loss는 loss–C_m 관점에서 다섯 downstream task에 대한 average perplexity로 측정한다.
- 3.2 Loss–compute scaling: 기존 scaling law는 quality-weighted mixture data with repetition 환경에서의 scaling behavior를, 특히 extrapolation에 대해 특성화하기에 불충분하다.이러한 관찰은 data-quality distribution과 repetition degree를 명시적으로 포함하는 수정 scaling law의 필요성을 뒷받침한다.
4. Information Scaling Laws
InfoLaw는 pretraining을 information accumulation으로 모델링해 data quality, repetition, model scale, training tokens를 결합하고 validation loss를 예측한다. repetition과 quality를 반영한 Information metric은 다양한 training configuration을 하나의 power-law curve로 통합해 더 큰 model과 token budget으로의 extrapolation을 가능하게 한다.
- Information-based formulation: Information은 data quality, repetition, model scale, training tokens에서 누적된 뒤 power law를 통해 최종 validation loss로 단조롭게 매핑된다.이 metric은 LayerMix weights, training tokens, fitted quality-density function과 repetition-rate function으로 계산되므로 새로운 run 전에 loss를 예측할 수 있다.
- Repetition modeling: 반복 노출에서는 exponential decay model이 총 learned information을 각 document의 full information content에 점차 포화시키므로 information gain이 감소한다.Decay rate λ(N)는 model non-embedding FLOPs/token에 의존하며, repetition effect를 model scale과 연결한다.
- Quality and repetition integration: 더 높은 quality의 bucket에는 더 높은 information density가 부여되며, 각 bucket의 learned information은 packed-data content와 model의 repetition-dependent learning ability를 결합한다.Bucket d에 대해 formulation은 unique-token count Md = min(wdK, BdS), density fd, average repeat count Rd를 사용한다.
- Unified scaling relationship: InfoLaw는 LayerMix weights, model non-embedding FLOPs/token, training tokens를 달리한 실험을 하나의 통합된 power-law loss curve로 통합한다.Traditional compute-based scaling laws는 repetition이 있는 quality-weighted mixture에서 신뢰하기 어려우므로, effective data signal로서 Information을 도입한다.
- Power-law fitting: Fitted loss–Information power law는 α = 3.7373과 β = 0.0441을 사용하며, log-log slope는 −β이고 intercept는 log(α)다.이 framework는 small-model comparison과 더 많은 tokens로 학습한 larger model로의 extrapolation을 지원한다.
5. FITTING 실험
fitting 실험은 model size, data-quality mixture, overtraining을 아우르는 27개 run에서 InfoLaw의 quality-density 함수와 model-capacity 함수를 추정한다. 적합된 capacity curve는 작은 model에서 빠르게 증가한 뒤 logarithm적으로 포화하며, 더 큰 model로 강하게 extrapolate된다.
- 실험 설정: 27개 run은 3.6x over-trained ratio에서 HQ, MQ, LQ LayerMix weight를 사용해 252M–1.2B 규모의 9개 model을 학습한다.실험에는 SwiGLU activation, RoPE embedding, 250k-vocabulary tokenizer를 사용하는 transformer가 쓰인다.
- Parameter fitting: fitting objective는 modeled information과 evaluation loss 사이의 Spearman correlation을 최대화하는 동시에, 더 높은 품질의 data bucket에 더 높은 quality density를 부여하도록 제약한다.quality-density parameterization은 bucket index가 증가할수록 감소하도록 제약되며, θ > 0이다.
- Parameter fitting: θ와 λ(N)의 100,000개 sampled combination으로 optimal parameter를 식별한 결과, quality-density 함수의 fitted θ*=0.922를 얻는다.이후 fitted density를 capacity function과 결합해 임의의 mixture, token budget, model size에 대한 information을 계산한다.
- Parameter fitting: a*=0.140과 b*=0.018은 λ(N)-N curve에 적합하며, 이 logarithmic form은 더 큰 model로의 강한 extrapolation을 보여준다.이 관계는 nonlinear하며, 작은 N에서는 빠르게 증가하고 N이 증가할수록 점진적으로 포화된다.
6. 외삽
InfoLaw는 관측하지 않은 데이터 recipe, 최대 7B까지의 더 큰 모델 규모, 더 높은 overtraining 수준에서 validation loss를 외삽하며, 다양한 training budget에서 recipe 선택을 가능하게 한다. 또한 모델 규모와 token 수에 따라 최적의 품질–다양성 trade-off가 어떻게 변하는지도 포착한다.
- 관측하지 않은 recipe: InfoLaw는 관측하지 않은 LayerMix recipe를 직접 예측하는 반면, 기존 scaling law는 새로운 recipe별 곡선을 fitting하기 위해 추가 실험이 필요하다.예측된 loss는 관측하지 않은 MLQ 및 MHQ mixture와 추가로 샘플링한 25개 weight에서의 실험 결과와 일치한다.
- 더 큰 모델 규모: InfoLaw는 1.5B–7B 규모를 포함한 더 큰 모델로 loss를 정확하게 외삽하는 반면, 기존 scaling law는 높은 compute에서 지나치게 낙관적인 예측을 보인다.모델 규모 외삽에는 252M–1.2B 모델의 training data가 사용되었으며, 해당 범위를 넘어선 경우에도 정확도가 유지된다.
- 관측하지 않은 recipe와 규모: 관측하지 않은 LayerMix weight와 모델 규모 전반에서, 7B 모델까지의 결합 외삽을 포함해 0.15% validation-loss error를 달성했다.이 결과는 MLQ, MHQ, 무작위로 샘플링한 25개 weight set, 그리고 관측하지 않은 모델 규모를 포함한다.
- 더 높은 overtraining: m′ = 25에서 fitting한 parameter를 사용하면, InfoLaw는 원래 scaling curve와 nearly parallel한 새로운 higher-overtraining regime를 예측하며, 이는 주로 intercept shift가 발생함을 나타낸다.Cm′ prediction은 Cm data에만 fitting한 parameter로 생성된다.
- Recipe 최적화: 25개의 held-out LayerMix configuration에서 Pearson correlation reached 0.76을 기록해, 추가 search experiment 없이 recipe 순위를 매기는 데 InfoLaw를 사용할 수 있음을 뒷받침한다.모델 규모 또는 전체 training token 수가 증가할수록 최적 recipe는 고품질 강조에서 더 높은 다양성으로 이동한다.
7. 결론
이 논문은 data-constrained 환경에서 quality-weighted mixing을 적용할 때 downstream performance를 예측하는 정교한 scaling law인 InfoLaw를 제안한다. InfoLaw는 더 큰 computational scale에서 관측되지 않은 recipe로 정확하게 외삽하며, 광범위한 추가 실험 없이 효율적인 data-recipe 탐색을 가능하게 한다.
- 결론: InfoLaw는 quality-weighted mixing을 적용한 data-constrained 환경에서 downstream model performance를 예측한다.이러한 환경에 초점을 둔 정교한 scaling law로 제시된다.
- 결론: 0.15% 평균 absolute error와 0.96% maximum error는 더 큰 computational scale에서 관측되지 않은 data recipe에 대한 예측이 정확함을 보여준다.이 결과는 평가된 data recipe와 scale을 넘어선 외삽을 뒷받침한다.
- 결론: InfoLaw는 광범위한 추가 실험 없이 최적 data recipe를 효율적으로 탐색할 수 있게 한다.이러한 예측 능력은 recipe 선택 시 폭넓은 실험적 탐색의 필요성을 줄인다.
8. 영향 성명 · A. 학습 데이터셋
이 논문은 비용이 많이 드는 recipe 실험을 줄이면서 LLM pretraining에서 data mixing과 repetition을 더 잘 이해하는 방법으로 InfoLaw를 제시한다. 학습 데이터는 96개 snapshot과 3.7T tokens에 걸친 deduplicated English Common Crawl이며, 더 광범위한 LLM 위험도 여전히 중요하다.
- 8. 영향 성명: InfoLaw는 서로 다른 data mixing 및 repetition 전략에서 LLM 성능을 더 잘 이해하는 것을 목표로 한다.
- 8. 영향 성명: InfoLaw는 data recipe에 대한 비용이 큰 시행착오를 줄여 pretraining을 더 효율적으로 만들 수 있다.
- 8. 영향 성명: 저자들은 이 기여에서만 고유하게 발생하는 직접적인 부정적 사회적 결과를 예상하지 않는다.
- 8. 영향 성명: Bias, misuse, unsafe deployment는 LLM과 관련된 더 광범위한 중요한 윤리적 문제로 남아 있다.
- A. 학습 데이터셋: 학습 corpus는 CC-MAIN-2013-20부터 CC-MAIN-2024-18까지 96개 snapshot에 걸친 Common Crawl의 English 부분을 사용한다.
- A. 학습 데이터셋: 모든 snapshot에 대해 수행한 global fuzzy deduplication으로 3.7T tokens를 포함하는 dataset을 구축했다.
B. 정규화 항 log(K)의 근거 · C. LayerMix 샘플링 함수
저자들은 상수 또는 power-law 대안과 달리 log(K)가 token budget 전반의 repetition decay를 가장 잘 포착하는 정규화라고 근거를 제시한다. 또한 compute-optimal run과 overtrained run의 차이를 포함해 LayerMix 샘플링을 구체화한다.
- B. 정규화 항 log(K)의 근거: Equation 3은 decay function에 log(K)를 포함해 repetition decay와 전체 token budget 간 상호작용을 모델링한다.저자들은 상수 및 power-law 정규화 대안을 평가한 뒤 이 형식을 선택했다.
- B. 정규화 항 log(K)의 근거: 상수 정규화는 더 큰 token budget으로 학습한 대형 모델에서 누적 Information을 과대추정해, 지나치게 낙관적인 loss 예측을 생성한다.이 형식은 information density의 scaling 특성을 반영하지 못하며 실험 결과에서 크게 벗어난다.
- B. 정규화 항 log(K)의 근거: Power-law 정규화는 Information과 Validation Loss 간 관계를 적합할 수 없다. 데이터 포인트가 필요한 power-law 상관관계 없이 흩어져 있기 때문이다.따라서 저자들은 이 형식으로부터 유효한 scaling law를 도출할 수 없었다.
- B. 정규화 항 log(K)의 근거: log(K)는 다양한 (w, K, S) 설정을 하나의 power-law curve로 유일하게 수렴시키면서, 252M에서 7B model scale까지 낮은 extrapolation error를 유지한다.이 현상은 Figure 3f에 제시되어 있으며, 전체 training budget에 비해 반복 데이터의 한계 효용이 로그적으로 감소함을 보여준다.
- C. LayerMix 샘플링 함수: LayerMix의 sampling function은 Algorithm 1에 자세히 제시되어 있다.이 문단은 해당 알고리즘이 LayerMix sampling procedure의 명세임을 밝히지만, 추가적인 구현 세부사항은 제공하지 않는다.
- C. LayerMix 샘플링 함수: m = 1은 compute-optimal training run을 의미하는 반면, m > 1은 compute budget에 비해 overtraining을 의미한다.이 구분은 compute-optimal regime을 넘어선 학습을 LayerMix가 규정하는 방식을 정의한다.
D. 학습 … K. Refinedweb으로의 일반화
학습, 반복 분석, 품질 평가, 일반화 테스트 전반에서 InfoLaw는 데이터 품질, 반복, 모델 규모, 토큰 혼합이 loss와 downstream 성능에 미치는 영향을 포착한다. RefinedWeb에서는 관측하지 않은 품질 구성을 extrapolate해 0.24% 평균 및 0.36% 최대 절대 백분율 오차를 보인다.
- D. 학습: 학습에는 2048-token sequences, cosine decay, 초기화된 learning rate lr = round(0.3118 · C^-0.1250, 8), 0.5% warmup, AdamW, 그리고 명시된 β 및 weight-decay 설정이 사용된다.optimizer는 β1 = 0.9, β2 = 0.95, weight decay = 0.1을 사용한다. 반복 분석은 IST와 LST regime을 구분하며, LST는 repetition을 유발하고 더 강한 repetition은 최종 loss를 악화시킨다.
- F. benchmark validation loss와 성능의 관계: Validation loss는 downstream benchmark 성능과 거의 선형적으로 관련되며, 연구한 operating regime 내에서 더 낮은 loss가 일관되게 더 높은 성능에 대응한다.이 관계는 ARC-C, ARC-E, HellaSwag, MMLU-Lighteval, TriviaQA에 걸쳐 나타나며, Spearman 상관계수는 Table 4에 보고된다.
- G. λ의 대안적 Fit: 로그arithmic λ(N) model은 exponential 및 power-law 대안보다 더 잘 fit하고 extrapolate하므로 최종 parameterization으로 채택된다.λ(N)은 parameter λ와 non-embedding FLOPs/token N의 관계를 나타낸다.
- H. 전통적 Scaling Law의 편차: IST와 LST sampling 모두에서 LayerMix loss–C curve는 처음 세 data point에 fit한 전통적 scaling law에서 명확히 벗어난다.이 편차는 서로 다른 LayerMix sampling weight와 repetition regime 전반에서 나타난다.
- I. 품질 점수: 고품질 FineWebEdu-selected subset은 random data보다 우수하며, 30B tokens로 1.2B model을 학습할 때 더 높은 품질의 subset이 더 나은 성능을 보인다.비교에는 Penedo et al. (2023)의 top 5%, top 20%, random subset이 사용된다.
- J. InfoLaw를 이용한 Token Mix 최적화: InfoLaw는 작은 model과 budget에서는 quality를 우선하고, 큰 model과 budget에서는 최적의 LayerMix token mix에서 diversity를 우선한다고 예측한다.이러한 model 및 budget별 token-mix ratio는 Table 6에 보고된다.
- K. Refinedweb으로의 일반화: RefinedWeb에서 InfoLaw가 fit한 quality-density parameter θ는 0.93으로, primary dataset의 0.92에 가깝다.두 dataset이 서로 다른 filtering strategy를 사용했지만 모두 Common Crawl에서 파생되었다는 점이 이러한 유사성의 원인으로 제시된다.
- K. Refinedweb으로의 일반화: InfoLaw가 RefinedWeb에서 관측하지 않은 MLQ 실험의 validation loss를 extrapolate했을 때 0.24% 평균 및 0.36% 최대 절대 백분율 오차를 달성했다.HQ와 LQ를 사용해 fit하고 MLQ는 extrapolation을 위해 hold out했다. 사용 가능한 model scale이 세 개뿐이어서 λ(N) curve는 fit하지 않았다.
L. 한계
이 연구의 데이터 버킷화는 quality tier의 최적 개수나 경계에 대한 ablation 없이 고정된 경험적 휴리스틱에 의존하며, overtrain degree m이 미치는 체계적 영향은 아직 이론적으로 설명되지 않았다.
- L. 한계: 데이터 버킷화는 quality tier의 최적 개수나 경계를 식별하는 ablation study 없이 고정된 경험적 휴리스틱을 사용한다.더 체계적인 데이터 분할은 예측 정확도를 높일 수 있다.
- L. 한계: overtrain degree m에 따른 scaling-law curve의 체계적 이동은 관찰되지만, 아직 이론적으로 설명되지 않았다.이러한 설명을 개발하는 일은 여전히 미해결 과제다.
- L. 한계: 더 체계적인 데이터 분할 접근법은 모델의 예측 정확도를 높일 수 있는 명확한 방향이다.