Source-linked AI summary

A Large Open Multi-Energy Corpus of Soil Compaction Tests, with Machine-Learning Baselines

Sompote Youwai, Chana Phutthananon, Warat Kongkitkul

arXiv:2609.03337v1cs.LG

TL;DR

Published compaction correlations are limited by small, narrow, rarely released datasets, motivating a broader laboratory corpus for estimating maximum dry density and optimum moisture content. The paper assembles and screens 2,854 tests, evaluates predictors under increasingly demanding provenance splits, and supplies physically constrained closed forms. The tabular foundation model performs best under random folds but degrades for grouped and source-held-out evaluation, while compactive energy becomes decisive conditionally.

  • Problem

    Existing compaction correlations commonly rely on small, single-laboratory or single-energy datasets that are rarely released, limiting evaluation across materials and unseen provenance.

  • Method

    The paper assembles and screens 2,854 laboratory tests from six public sources, then estimates MDD and OMC from classification properties and compaction standard information using multiple learners and physically constrained closed forms.

  • Results

    The tabular foundation model reaches R2 0.824 for density and 0.784 for water content under random folds, falling to 0.727 and 0.696 with provenance-grouped folds and 0.520 and 0.614 with a source held out.

  • Takeaways & Limitations

    The corpus provides a broader laboratory baseline, while provenance-aware evaluation shows that random-fold accuracy should not be treated as performance on unseen sources.

  • Takeaways & Limitations

    The corpus has no field density, and laboratory-held-out evaluation cannot be run because public records do not identify the laboratory performing each test.

Abstract

from arXiv · show

Every engineered fill is specified by a maximum dry density and an optimum moisture content. Each determination needs a full Proctor test. Published correlations rest on one to four hundred specimens, usually from one laboratory at one compactive energy, and are seldom released. This paper releases a corpus without those limits. It holds 2,854 laboratory compaction tests from six public sources, across 162 provenance groups and four Proctor energy levels, with fines from 1.5 to 100%. Every record is audited to the Proctor method its source names, and no energy is inferred. Screening on the zero-air-voids condition removed 11.8% of harmonised records, and 5.7% of those with a measured specific gravity. A material share of published compaction data is physically impossible. The optimum degree of saturation over the corpus is 0.815 at a coefficient of variation of 11%. That is a baseline, not a constant. Both parameters are then estimated from one classification suite and the compaction standard. A tabular foundation model reaches R2 0.824 for density and 0.784 for water content under random folds. It reaches 0.727 and 0.696 with folds drawn around provenance, and 0.520 and 0.614 with a whole source held out. Compactive energy is negligible marginally yet decisive conditionally. Density on the 66 modified-Proctor records is predicted at R2 0.740 with it and -0.651 without. Symbolic regression yields closed forms coupled through a phase relation. No predicted pair can then exceed the zero-air-voids line. The predictions are for screening, not acceptance.

I. INTRODUCTION

Proctor tests determine the maximum dry density and optimum moisture content that govern engineered-fill compaction, but existing correlations rely on small, narrow, rarely released datasets. This study frames a broader laboratory corpus as necessary for evaluating estimates across soils, energies, and provenance boundaries.

  • I. INTRODUCTION: Each Proctor determination requires four to six specimens at different water contents and about a working day, while specifications require one for every material change.MDD is the highest dry density under stated effort; OMC is the water content at that maximum.
  • I. INTRODUCTION: The optimum moisture content also governs soil fabric, strength, volume change, permeability, and susceptibility to swelling or collapse.These effects can persist at equal dry density when soils are compacted dry or wet of optimum.
  • I. INTRODUCTION: The zero-air-voids line bounds physically admissible optima, so records must be screened against it rather than treated as automatically valid.No compactive effort can produce a state above this line.
  • I. INTRODUCTION: Randomly shuffled evaluation can measure interpolation within a pooled corpus because related soils, protocols, laboratories, or publications may occur on both sides of the split.The paper therefore treats provenance-aware evaluation as central to judging performance on unseen material sources.
  • I. INTRODUCTION: Published correlations typically use one to four hundred specimens, are rarely deposited, and often restrict inputs to plastic soils or one compactive energy.Existing deposited datasets are generally small and tied to one laboratory, compilation, programme, or standard effort.

D. Scope and objectives

The study addresses the lack of broad, provenance-preserving compaction data by assembling and screening public laboratory records, evaluating MDD and OMC predictors under provenance-aware folds, and providing physically constrained closed forms. Its scope remains laboratory compaction, with exclusions and data-integrity choices materially affecting evaluation.

  • D. Scope and objectives: The practical gap is that engineers cannot determine whether published correlations represent a new borrow material or a laboratory the correlation never saw.Pooling sources is intended to support evaluation across materials and provenance boundaries rather than merely enlarge an existing single-study dataset.
  • D. Scope and objectives: The study compares software-based ensembles with closed-form expressions whose phase relation constrains predicted pairs below the zero-air-voids line.The closed forms trade accuracy for inspectability and physical admissibility.
  • D. Scope and objectives: The task estimates MDD and OMC from one classification suite and the known compaction standard, without using quantities derived from the compaction curve as predictors.The inputs include plastic limit, plasticity index, fines, sand, and specific gravity.
  • D. Scope and objectives: The corpus combines six public sources, four Proctor energy levels, coarse and fine soils, and record-level provenance so models can be tested on unseen sources.It includes 2,854 tests and spans fines contents from 1.5% to 100%.
  • D. Scope and objectives: The retained evidence is laboratory-only and open-access, while integrity and completeness exclusions can substantially change model scores.Admitting the excluded article-and-thesis stream reduced gradient-boosted performance from R2 0.819 to 0.536 for density and from 0.783 to 0.467 for water content.

B. The six retained sources

The retained corpus comprises six audited laboratory sources, with source-specific conventions preserved rather than inferred. Its composition is dominated by LTPP, while non-standard energy coverage comes from a single multi-energy deposit.

  • Every retained record is the peak of a laboratory compaction curve whose origin was audited source by source.
  • All LTPP records are treated as AASHTO T99 at 592.5 kJ/m3 based on programme documentation and physically sensible density interpretation.
  • LTPP contributes 69% of records and is coarser and drier than four of the five deposits, weighting pooled accuracy toward LTPP soils.
  • The multi-energy deposit is a single source, so every non-standard test comes from it and source-held-out evaluation has limited scope.
  • No compactive energy is inferred; records without a named Proctor method are excluded, leaving four stated energy levels.

D. Physical admissibility

Physical admissibility is screened with the optimum degree of saturation and the zero-air-voids condition. The screen removes a substantial share of harmonised records, but its exact rate depends on substituted specific gravity.

  • The optimum degree of saturation directly tests admissibility because values above unity place records above the physically impossible zero-air-voids line.
  • Records are retained only when optimum saturation lies between 0.20 and 1.0, MDD between 0.8 and 2.6 Mg/m3, and OMC between 1 and 70%.
  • 11.8% of 8,242 harmonised records were removed, with 94% of removals above the zero-air-voids line.
  • The rejection rate changes from 19.8% to 8.2% when substituted specific gravity moves from 2.60 to 2.75, while the broader impossibility finding remains robust.
  • The screened dataset contains 2,854 model-ready tests across 162 provenance groups, while two missing values are retained rather than filled.
  • Random partitions can leak near-copies across provenance groups, including exact copies in 40% of cases, so grouped evaluation better tests transfer.

C. The dataset in the compaction plane

The corpus is structured as a physically screened band in the MDD–OMC plane, where compactive effort, saturation, provenance, and incomplete inputs shape the estimation task.

  • The corpus occupies a band beneath the zero-air-voids line, with peaks treated as coupled coordinates of a saturation curve.The band closes above the screening rule and thins below S_r=0.6.
  • Compactive energy cannot be inferred from the plane alone because effort classes overlap, so each level is scored separately.Modified-effort records do not sit above standard records; between 5 and 10% water content, their median density is 0.192 Mg/m3 lower.
  • The model estimates MDD and OMC from five classification properties, specific gravity, and the compaction standard without using compaction-curve quantities.The two regressors use identical inputs and partitions but are fitted independently.
  • Figure 2 presents the estimation pipeline from measurements and derived inputs through mapping, targets, and intended screening uses.The intended uses exclude acceptance decisions because prediction error could consume the specification margin.
  • The modelling set contains 2,854 tests grouped by provenance, while the composition table records effort classes and incomplete columns.The grouping structure matters because records from a group can share deposits, laboratories, campaigns, and reporting conventions.

B. Models

The study compares several learners under fold designs that progressively restrict shared provenance, while qualifying random-fold interpretation and symbolic-regression evaluation.

  • Three principal learners are evaluated on identical folds, with additional architectures testing whether accuracy reflects the data rather than one model class.Symbolic regression additionally searches for closed-form expressions.
  • Models: Nested tuning of XGBRegressor returns R2 0.816 and 0.779 versus 0.819 and 0.783 for fixed values, so tuning yields no improvement.The fixed-value model avoids attributing the comparison to unequal search effort.
  • The three fold designs: Grouped folds measure prediction for new provenance groups, whereas random folds mainly measure interpolation within groups and are not evidence for unseen sources.The random and grouped designs use the same records, inputs, learners, and matched training-set size.
  • The three fold designs: Source-held-out evaluation is the harshest design, and its groupwise errors are skewed from 0.066 Mg/m3 at the median to 0.268 at the tail.Holding out LTPP leaves 878 training records; other held-out sources leave training sets at least 69% LTPP.
  • Evaluation qualifications: Symbolic-regression accuracies from one random split are selection scores because model choices and the saturation constant used held-out information.The expressions are later refitted and rescored under all three designs.

V. RESULTS

Predictive accuracy is strong under random folds but falls when provenance or sources are held out, while physical consistency separates independently trained learners and motivates coupling.

  • Predictive accuracy: TabPFN reaches R2 0.824 for density and 0.784 for water content under random folds, with errors of 0.066 Mg/m3 and 1.87%.These errors are reductions of 63% and 60% against the mean predictor.
  • Predictive accuracy: Grouped folds cost about 0.10 in R2 on both targets and raise water-content error from 1.87% to 2.28% without changing inputs, learners, or training-set size.The difference reflects whether the scored publication was already represented in training.
  • Transfer to an unseen source: With a whole source held out, the learners span 0.19 in density R2 and their ordering reverses, unlike the 0.03 span under random and grouped folds.Source-level difficulty varies substantially across the corpus.
  • Physical consistency: Independent prediction can place pairs above the zero-air-voids line, whereas phase-relation coupling forbids such physically impossible pairs.Even the training-fold mean produces 28 violations, showing that improved fitting alone cannot guarantee admissibility.
  • Physical consistency: TabPFN produces 1 above-line pair versus 7 for the ensemble on random folds, while neither learner is more accurate than the other.The comparison favors TabPFN on this physical-consistency criterion and on having no hyperparameters.
  • Transfer to an unseen source: Leave-one-group-out errors are skewed: the median is 0.066 Mg/m3 and 1.73%, while the worst group reaches 0.268 Mg/m3 and 11.0%.A pooled mean absolute error understates risk for an unlucky material.

C. What the estimate rests on

Prediction depends chiefly on consistency limits and gradation, and continuous inputs add information beyond classification labels, although broad classes and additive models leave residual gaps.

  • Input groups: Consistency limits and gradation each cost 0.096–0.213 in R2 when removed, while sand, energy, and specific gravity each cost at most 0.030.The two leading groups are jointly load-bearing rather than consistently rankable across targets and designs.
  • Consistency limits: The plastic limit and plasticity index set is used throughout because it matches or exceeds the three-limit set under every design and target.The liquid limit is released but excluded from the mapping because its redundancy costs transfer.
  • Value beyond a classification label: The USCS symbol explains R2 0.563 for density and 0.561 for water content, while six continuous inputs add 0.116 and 0.134 under random folds.Class means are computed within each training fold and scored on held-out records.
  • Value beyond a classification label: Under grouped folds, the ensemble’s gains over a class mean narrow but persist at 0.179 for density and 0.157 for water content.The practical advantage remains at about two-thirds of its apparent random-fold size.
  • Value beyond a classification label: Classification groups remain broad, and assuming additivity leaves a residual gap between the linear model and ensemble even with finer classification.The AASHTO block shows both gaps persist under a finer scheme.

D. Effect of compactive energy

Compactive energy has little pooled importance because standard-Proctor records dominate, but it is decisive for the minority of modified-Proctor tests. Encoding energy categorically preserves predictive performance across representations.

  • 96.5% standard-Proctor records make compactive energy appear marginally weak in pooled feature importance.Its pooled R2 importance is only 0.015 to 0.025, understating effects on less frequent effort strata.
  • 0.740 R2 for density on 66 modified-Proctor records with energy falls to -0.651 without it.Mean absolute error rises from 0.058 to 0.162 Mg/m3 when energy is removed.
  • Energy is decisive conditionally because models default to the dominant standard-Proctor relationship when modified effort is absent.The reduced-modified row contains only 7 records and carries no weight.
  • Energy cannot be evaluated under source-held-out folds because all 101 non-standard records come from one source.Holding that source out leaves energy present but untrained at three of its four levels.
  • 0.819 and 0.783 R2 remain unchanged when energy is represented by exact conversion, recomputed values, rounded SI values, or ordered categories.Four indicator variables return 0.819 and 0.780.

E. Closed-form expressions by symbolic regression

Symbolic regression produces a compact density expression and recovers moisture content through a phase-relation identity. The resulting formulation is physically bounded but empirically limited to the corpus ranges used for fitting.

  • E. Closed-form expressions by symbolic regression: The recommended symbolic pair predicts density directly and recovers optimum moisture content through a phase relation.The coupling makes admissibility structural rather than fitting two independent expressions.
  • 1) The recommended pair:: Complexity 20 retains normalized energy at an R2 cost of only 0.001, avoiding the unsound omission of effort.The complexity-18 entry scores 0.693 against 0.692 but has no energy term.
  • 1) The recommended pair:: The density expression uses normalized percentages and energy, specific gravity, and density inputs, with standard effort setting normalized energy to one.Plastic limit, plasticity index, and fines are divided by 100%; density is divided by 1 Mg/m3.
  • 1) The recommended pair:: Equation (4) is an exact phase identity whose sole estimated quantity is the constant 0.820.At fixed saturation it defines MDD–OMC curves for each specific gravity under the zero-air-voids bound.
  • 1) The recommended pair:: Density falls with fines and plasticity measures and rises with specific gravity and effort, but the recurring grouping is empirical rather than mechanistic.The expression was returned because one reciprocal costs fewer symbols than two linear terms.
  • 1) The recommended pair:: Figure 4 shows the identity at saturation 0.820 for three specific gravities and the fold-to-fold range of its constant.The figure places measured records behind the curves and compares the constant against alternative saturation values.
  • 1) The recommended pair:: The symbolic equations apply only across the fitted ranges of the 2,720-record subset and flag inputs outside those ranges.The ranges include fines from 1.5 to 100%, specific gravity from 2.30 to 2.94, and four Proctor energies.

2) Why the coupling, and what it costs:

Coupling OMC to predicted density through the phase relation prioritizes physical admissibility over small accuracy gains. The corpus-wide saturation estimate is a useful baseline, not a soil-independent constant.

  • 0 predicted pairs cross the zero-air-voids line for every coupled form, versus 75 to 98 for the independent OMC expression.The independent form reaches saturation 132.7 in the source-held-out design.
  • Form B is recommended because it is bounded for any specific gravity, whereas the alternative proportional-to-void-ratio form fails above specific gravity approximately 3.26.Form C adds 0.012 R2 but requires a second fitted model.
  • The closed form costs about 0.11 R2 under random folds, but the gap nearly closes under grouped folds and reverses for water content with a source held out.The comparison uses the same records and folds for the closed form and ensemble.
  • Coupling figures are confined to the 2,720-record symbolic subset, and its structure was selected after seeing every source.The fitted constants transfer across folds, but structure discovery was not source-blind.
  • The corpus-wide optimum degree of saturation is 0.815 with an 11% coefficient of variation, but soil group shifts the mean from 0.706 to 0.853.Compactive effort moves it least, while source and specific-gravity treatment also affect the mean.
  • The 0.815 saturation value is a baseline rather than a constant or substitute for testing the material at hand.Its fifth-to-ninety-fifth percentile range is 0.655 to 0.951.

VI. LIMITATIONS

The study’s evidence is limited by laboratory-only data, coarse provenance grouping, and cross-validation without an independent test set. Its corpus and models support screening conclusions, not unrestricted field or out-of-corpus acceptance.

  • Laboratory-only data do not test whether optimum saturation near 0.82 describes field-compacted lifts.The comparison with prior field-related work therefore uses different evidence bases.
  • The provenance grouping is coarser than laboratory-held-out evaluation because public sources do not identify the laboratory for each test.A single programme contributes 69% of the records.
  • No independent test set exists; source-held-out validation remains a selected, harmonised, and screened substitute rather than a genuinely new laboratory.The input set, ablations, and one constant were chosen after inspecting the same 2,854 records.
  • 2,854 laboratory tests from six public sources form the primary contribution, spanning four Proctor energy levels and both sides of the coarse–fine boundary.Every record retains its source-named Proctor method and stated compactive energy.
  • The corpus removed 11.8% of harmonised records under the zero-air-voids screen, compared with 5.7% among records reporting measured specific gravity.The optimum saturation mean is 0.815 with an 11% coefficient of variation and is a laboratory baseline, not evidence of field invariance.
  • Random-fold results overstate interpolation because 96.6% of records share a provenance group with training and 40.1% share an exact input vector.Grouped folds reduce R2 by about 0.10, while source-held-out folds reduce it further.
  • Closed forms reach R2 0.710 and 0.688 under random folds, 0.676 and 0.666 grouped, and 0.579 and 0.606 source held out.Coupling prevents zero-air-voids violations, but the expressions are not recommended over ensembles outside the corpus.

CREDIT AUTHORSHIP CONTRIBUTION STATEMENT

The statement assigns conceptual, methodological, analytical, writing, supervisory, and funding roles across three authors, while also documenting the use and verification of an editing and scripting tool.

  • Sompote Youwai is credited with conceptualization, methodology, software, formal analysis, data curation, writing, visualization, supervision, and funding acquisition.
  • Chana Phutthananon contributed methodology, validation, and writing–review and editing.
  • Warat Kongkitkul contributed conceptualization, methodology, writing–review and editing, and supervision.
  • The authors used Claude for prose editing, tabular restructuring, and routine analysis scripting, then reviewed the resulting content.
  • The corpus, fitted models, command-line predictor, and analysis code are publicly released under CC BY 4.0 with validation materials.
Loading 2609.03337v1…