Source-linked AI summary
Data Mixing as Mixture Experiment: Response Surface Methodology and Optimal Design for Large Language Model Pretraining
Yicheng Mao, Hongru Du
TL;DR
LLM pretraining requires choosing domain shares under a fixed token budget, but proxy-based data-mixing studies do not traditionally frame this choice as an experiment. This paper applies mixture-experiment modeling and robust optimal design, showing that proxy runs can be reduced while retaining useful mixture rankings.
Problem
With a fixed token budget, practitioners must determine how to allocate pretraining data across domains because corpus composition strongly affects model performance.
Method
The paper reframes proxy data-mixing studies as mixture experiments, modeling validation loss over the probability simplex with sparse Scheffé surfaces and robust I-optimal designs.
Results
Model-robust I-optimal designs reduce the proxy runs needed to recover useful mixture rankings, while Scheffé models decompose domain effects and interactions.
Takeaways & Limitations
LLM data-mixing optimization can be treated as an experimental-design problem in which proxy mixtures are selected for statistical efficiency.
Takeaways & Limitations
The conclusions depend on how well the fitted response surface approximates the true training response, especially if higher-order interactions or threshold effects dominate.
Abstract
from arXiv · showhide
Data mixing is a central design problem in large language model pretraining: given a fixed token budget, practitioners must decide how much data to allocate to each domain. Recent proxy-based methods address this problem by training small models on candidate mixtures, fitting a response model, and using the response to select mixtures for larger-scale training. We show that this workflow has the structure of a classical mixture experiment. Under this view, data domains are mixture components, token shares are component proportions, proxy-training runs are experimental design points, and validation loss defines a response surface over the probability simplex. We develop this formulation using sparse second-order Scheffé response-surface models and construct model-robust $\mathcal{I}$-optimal designs for proxy data-mixing experiments. Using RegMix as an empirical case study, we demonstrate how the framework can both interpret observed mixture responses and design more efficient proxy experiments. The Scheffé analysis shows that domain value is strongly relational: several domains that are weak under additive effects become favourable through pairwise interactions, especially through combinations with web-derived text. The sparse Scheffé model preserves mixture rankings across model scales and remains competitive with a flexible machine-learning predictor while providing an explicit decomposition of additive and interaction effects. In a simulation study calibrated to observed proxy-training responses, model-robust $\mathcal{I}$-optimal designs recover the relevant mixture ordering after removing about 25\% of the original proxy runs. These results suggest that LLM data mixing should be treated not only as a prediction problem, but also as an experimental-design problem in which the proxy mixtures themselves can be chosen to improve statistical efficiency.
1 Introduction
The paper recasts LLM data mixing as a constrained mixture experiment on the probability simplex, where domain proportions determine validation loss or downstream performance. It develops interpretable Scheffé models and optimal proxy-run designs to improve mixture understanding and experimental efficiency.
- Motivation: LLM pretraining data mixing is a constrained allocation problem because increasing one domain’s token share reduces the shares available to others.The fixed token budget makes mixture selection consequential for validation loss, downstream performance, compute use, and model capability.
- Mixture-experiment formulation: The framework maps data domains to mixture components, token shares to proportions, proxy-training runs to design points, and model performance to the response surface.All feasible mixtures lie on a probability simplex because proportions are nonnegative and sum to one.
- Response-surface modeling: Sparse second-order Scheffé models respect simplex geometry while decomposing mixture responses into additive domain effects and pairwise interactions.This makes relational domain value interpretable when a domain is weak additively but useful in combination with another.
- Optimal experimental design: Proxy mixtures can be selected as experimental design points rather than sampled randomly, with I-optimality targeting low average prediction variance across the mixture region.Model-robust designs balance performance across additive and interaction-inclusive candidate models when response-surface order is unknown.
- Empirical contributions: Using RegMix data, the paper evaluates whether sparse Scheffé models preserve mixture rankings across model scales and whether model-robust I-optimal designs recover rankings with fewer proxy runs.The framework is intended for proxy-based workflows allocating a fixed training budget across data sources.
- Contributions: The paper’s contribution is to connect LLM data mixing with mixture modelling and optimal experimental design while improving proxy-experiment efficiency.It argues that choosing proxy mixtures deserves as much attention as modelling the resulting data.
2 Proxy Data Mixing as a Mixture Experiment
The paper recasts proxy-based LLM data mixing as a classical mixture experiment over the probability simplex, with proxy runs supplying response-surface data for selecting larger-scale mixtures. This view motivates interpretable Scheffé modelling and informative I-optimal experimental designs rather than purely random or heuristic sampling.
- Mixture-experiment formulation: The fixed token budget constrains feasible mixtures to a probability simplex, so increasing one domain’s share necessarily reduces the proportions available to others.This compositional constraint distinguishes mixture predictors from ordinary unconstrained covariates.
- Mixture-experiment formulation: Proxy studies select candidate mixtures, train small models, record validation loss or downstream scores, and fit predictors to guide larger-scale mixture selection.Data domains are mixture components, token shares are proportions, proxy runs are design points, and fitted predictors are response surfaces over the simplex.
- RegMix case study: Proxy-to-large-scale transfer relies on rank preservation: mixtures performing relatively well for small proxy models are expected to remain relatively strong at larger scales.Accordingly, RegMix evaluates fitted regressions primarily through cross-scale rank correlation rather than only absolute loss prediction.
- Response-surface modelling: Scheffé response-surface models provide a simplex-native, interpretable alternative that represents pairwise domain interactions through explicit coefficients.This is useful when a source’s value depends on both its own share and the sources with which it is combined.
- Optimal experimental design: I-optimal design selects proxy mixtures to reduce average prediction variance across a specified mixture region, potentially achieving comparable ranking performance with fewer runs.Treating proxy mixtures as experimental design points enables their simplex locations to be chosen for prediction of unseen mixtures and lower token and compute costs.
3 Methods
The methods formulate proxy data-mixing experiments as response surfaces on the mixture simplex, using canonical and sparse second-order Scheffé models to capture additive effects and domain interactions. They select proxy mixtures with model-robust I-optimal designs that target prediction precision under both first- and second-order candidate models.
- Mixture response-surface formulation: Data mixtures are represented as proportions on a (K −1)-dimensional simplex, where the ordinary intercept is not separately identifiable.The simplex constraint makes the constant term expressible through the component proportions.
- Scheffé models: Canonical Scheffé models omit the ordinary intercept, and first-order coefficients represent fitted responses at pure-domain vertices.The vertex interpretation is model-based because practical pretraining mixtures need not include pure-domain configurations.
- Scheffé models: Second-order Scheffé terms measure departures from additivity, with negative coefficients lowering fitted validation loss relative to additive expectations.Positive interaction coefficients instead raise fitted loss, conditional on the validation objective.
- Sparse interaction modeling: An L1-penalised second-order Scheffé model shrinks weak coefficients to zero, yielding a smaller set of selected main effects and pairwise interactions.The sparse model is motivated when interactions help collectively but individual coefficients are weak, redundant, or unstable.
- Optimal experimental design: Model-robust I-optimal designs minimise average prediction variance across the simplex under both first-order and full second-order Scheffé models.The criterion addresses uncertainty about interaction structure and favours designs that retain precision under additive and interaction-inclusive surfaces.
4 Empirical Response-Surface Analysis
The Scheffé analysis decomposes RegMix validation-loss responses into additive domain effects and pairwise interactions, revealing that mixture value is strongly relational. Sparse second-order terms improve cross-scale mixture rankings while retaining an interpretable effect decomposition.
- Analysis objective: The analysis uses RegMix proxy-training data to separate additive domain contributions from pairwise departures and assess ranking preservation across model scales.The Scheffé formulation treats validation loss as a response surface over the mixture simplex.
- Additive effects: The additive ordering places enron emails, hackernews, philpapers, nih exporter, ubuntu irc, and pile cc among the most favourable contributors, while github, arxiv, pubmed central, freelaw, and stackexchange rank less favourably.Because the response is validation loss, smaller first-order coefficients indicate lower fitted loss under the additive component.
- Interaction effects: F = 6.5145 on 136 interaction and 359 residual degrees of freedom (p < 0.001) shows that pairwise terms improve the first-order Scheffé surface.The L1-penalized model retained 80 terms: all 17 first-order terms and 63 pairwise interactions.
- Interaction effects: Pile cc has the strongest relational role: its four largest interactions are negative, making additive-weak domains such as arxiv, pubmed central, freelaw, and github favourable when combined with it.Its strongest negative interactions include arxiv (β = −2.62), pubmed central (−2.36), freelaw (−2.31), and github (−2.24).
5 Optimal Design for Proxy Data-Mixing Experiments
Model-robust I-optimal designs concentrate mixture proportions at simplex boundaries, reducing prediction variance and preserving mixture-ranking performance with fewer proxy runs than the original RegMix design.
- Mixture-ranking performance: Both average Spearman rank correlation and pairwise ranking accuracy exceed the original n = 512 reference level with roughly 350 proxy runs.At n = 512, the model-robust I-optimal design exceeds the RegMix reference under both metrics.
- Budget reduction: The model-robust I-optimal designs match or exceed the original n = 512 RegMix design after removing about 25% of the proxy runs.Because each proxy run trains a small model on a fixed token budget, the reduction directly lowers pilot-stage token and compute cost; the exact saving depends on the fitted sparse Scheffé surface.
- Sensitivity analysis: Under a first-order Scheffé data-generating surface, model-robust I-optimal designs continue to outperform the original RegMix reference design.This supplementary sensitivity check supports the efficiency advantage beyond the fitted sparse second-order surface.
6 Discussion
The discussion frames LLM data mixing as an experimental-design problem, highlighting relational domain effects, efficient proxy-run selection, important limitations, and directions for extending the framework.
- Core framing: Data mixing is a structured mixture experiment in which domains are components, proxy runs are design points, and validation loss is a response surface.This reframing turns mixture optimisation from a purely predictive task into an experimental-design problem.
- Empirical interpretation: Sparse Scheffé models expose additive and pairwise interaction effects, showing that several domains become favourable through combinations with pile cc.This indicates that domain value depends on complementarity under a fixed token budget, rather than only on individual-domain quality.
- Design efficiency: About 25% of proxy runs can be removed while model-robust I-optimal designs match or exceed the original RegMix reference in the simulation-based case study.The saving depends on the fitted sparse Scheffé surface and assumed noise structure, so design-aware sampling is not universally guaranteed.
- Limitations: The empirical evaluation relies on one publicly available proxy-training dataset and simulation-based design comparisons, with conclusions depending on the fitted surface approximating the true response.Higher-order interactions, threshold effects, or local irregularities could make a sparse quadratic Scheffé model too restrictive.
- Limitations: Practical studies should incorporate infeasible proportions, data availability, licensing, preprocessing costs, and other restrictions through constrained or cost-sensitive designs.The current optimal designs assume any simplex point is sampleable and every proxy run has equal cost.
- Future directions: Future work could model scale and training conditions as process variables, develop sequential adaptive designs, and apply the framework to other fixed-budget data-mixing problems.The broader principle applies whenever training composition is controllable, including multimodal, domain-adaptive, multilingual, and task-mixed settings.