Source-linked AI summary
MineTRACE: An Evidence-Grounded Interactive Reasoning System for Mineral Prospectivity
Yiran Zhang, Jinwen Liu, Daniel Su, Yisu Chen, Qiang Sun, Chris Gonzalez, Eun-Jung Holden, Marco Fiorentini, Wei Liu, Yihao Ding
TL;DR
Mineral exploration systems often provide opaque scores despite the need to integrate heterogeneous evidence for practical decisions. MINETRACE combines public geoscience data with a transparent expert tree and shared evidence records exposed through interactive maps and natural-language tools. The scorer reaches 0.917 Spatial AUC for Ni, while interaction evaluation reports 92% Good responses and strong grounding.
Problem
Existing prospectivity systems often hide supporting evidence or fail to connect scoring, interactive inspection, and conversational access in one workflow.
Method
MINETRACE integrates public geochemical, geophysical, and geological data with a transparent expert tree and shared evidence records for maps, queries, panels, and conversational tools.
Results
0.917 Spatial AUC is achieved by Ni on realistic tests, while 92% of 150 evaluator-trial responses were rated Good and one response contained a fabricated number.
Takeaways & Limitations
MINETRACE supports evidence-grounded exploration by connecting prospectivity scores to their contributing signals, sources, expert weights, and coverage metadata.
Takeaways & Limitations
MINETRACE is a decision-support system rather than an autonomous or field-validated one, and its performance is uneven across commodities under spatial holdout.
Abstract
from arXiv · showhide
Mineral exploration requires integrating heterogeneous geochemical, geophysical, and geological evidence, yet existing prospectivity systems often provide only opaque scores or heatmaps. We present MineTRACE, a web-based system for evidence-grounded exploration of eight commodities: Cu, Au, Ni, W, Sn, Co, Ta, and Mn. Users can explore prospectivity maps, query locations or regions, inspect supporting evidence, and interact through natural language. A transparent expert tree, informed by geological knowledge and known deposits, combines multi-source evidence into interpretable prospectivity scores. For a new location, the conversational assistant retrieves the score and supporting evidence from the analysis pipeline and presents them in natural language. The scorer achieves spatial AUC values of up to 0.917 across different test scenarios, while end-to-end evaluation assesses query accuracy and response grounding. MineTRACE makes public geoscience data easier to access, interpret, and verify, supporting more efficient and transparent mineral exploration.
1 Introduction
MINETRACE addresses the gap between opaque prospectivity outputs and practical exploration decisions by linking multi-source evidence, interpretable scoring, interactive inspection, and grounded natural-language interaction. Its unified evidence record connects assessments to contributing signals, sources, and coverage.
- Research gap: Existing systems often separate accurate but opaque data-driven maps, transparent but static knowledge-driven methods, and interactive tools that are rarely connected.Users therefore lack a unified workflow for inspecting the evidence behind prospectivity assessments.
- System aim: MINETRACE lets users query locations or regions, inspect supporting evidence, compare targets, and obtain evidence-grounded explanations through natural language.The system is designed to make prospectivity assessments interactive, interpretable, and traceable.
- Core design: The system combines public geochemical, geophysical, and geological data with a transparent expert tree to produce structured prospectivity evidence records.Records include prospectivity scores, contributing signals, expert contributions, evidence sources, and local sample coverage.
- Core design: The same evidence record supports the map workspace, evidence inspection, spatial queries, and conversational interaction.This shared representation keeps system outputs consistent with the underlying prospectivity analysis.
- Evaluation: MINETRACE evaluates prospectivity ranking, spatial generalisation, multi-source fusion, query accuracy, and response grounding across eight commodities and large public datasets.The evaluation uses approximately 9.35 million assays and 3,420 known mineral sites.
2 Related Work
Prior work offers predictive prospectivity maps, interpretable decision rules, or interactive interfaces, but commonly lacks unified multi-source reasoning and quantitative evaluation. MINETRACE combines these elements into a traceable, evidence-grounded exploration system.
- Data-driven approaches: Data-driven prospectivity models capture complex geochemical patterns and can achieve strong predictive performance, but typically output only a score without interpretable evidence.Their anomaly-detection framing does not expose what contributes to each location’s prediction.
- Knowledge-driven approaches: Knowledge-driven models improve interpretability through feature importance and decision rules, yet are often limited to one modality, weak regional generalisation, or offline use.These approaches commonly lack explicit reasoning over multi-source evidence.
- Interactive systems: Existing interactive geoscience systems may integrate multiple sources and provide rankings, but lack publicly available quantitative evaluation.This limits verification of their evidence-supported output claims.
- MINETRACE: MINETRACE links multi-source data, interpretable scoring, and interactive natural-language exploration with outputs traceable to underlying evidence and quantitatively evaluated.Its architecture includes geospatial data storage and tool-mediated external-agent access.
3 System Design
MINETRACE uses a web architecture in which a shared scoring path converts public geoscience data into traceable prospectivity records. Users access those records through maps, evidence panels, spatial queries, and conversational tool use.
- System architecture: A browser workspace routes user requests through a backend to a scoring service that reads processed geoscience data and returns scores with evidence records.The architecture uses Django, PostGIS, and offline preprocessing and training jobs.
- Data and evidence layer: MINETRACE combines three Western Australia data products containing 9.35 million assays, 3,420 known mineral sites, seven geophysical rasters, and five geological vector layers.The mineral-site data provide supervision labels across eight commodities.
- Data and evidence layer: Raw observations are harmonised into a spatially indexed evidence store with cleaned assays, aligned rasters, geological attributes, structural data, and distance features.For each query, the system retrieves nearby observations from this store.
- Interpretable prospectivity engine: The prospectivity engine represents local evidence, scores it with fixed named exploration heuristics, and aggregates active expert scores into an interpretable result.Insufficient nearby observations are marked as low coverage rather than treated as confident predictions.
- Interpretable prospectivity engine: The final prospectivity score is a weighted mean of active expert scores, with weights fitted offline to separate known deposits from background locations.Experts lacking required evidence abstain, while a separate coverage value indicates support for the score.
- Interactive interfaces: Users can click or draw regions for scores, inspect ranked signals and expert contributions, explore toggleable context layers, and ask the assistant to operate the system through tools.The assistant answers from the same evidence records shown in the panels and reports tool-computed outputs.
4 Evaluation
MINETRACE is evaluated through scorer benchmarks across eight commodities and end-to-end human testing of natural-language interactions. Results show strong performance in realistic ranking settings, consistent gains from evidence fusion, and generally grounded system responses.
- Scoring Quality: 0.946 NonMine AUC and 0.917 Spatial AUC make Ni the strongest commodity on the realistic tests.Co also remains strong, reaching 0.929 NonMine AUC and 0.805 Spatial AUC.
- Scoring Quality: Cu and Mn degrade substantially under spatial holdout, indicating weaker regional generalisation.Performance varies across commodities and test settings, while Au and Sn perform well against validated nontarget sites.
- Scoring Quality: Combining geochemical, geophysical, and geological evidence consistently gives the highest AUC for every commodity.The ablation uses the easier far-random setting, and the importance of each evidence family differs by commodity.
- End-to-End Human Evaluation: 92% of responses across 150 question–evaluator trials were rated Good, with 21 of 30 questions receiving Good ratings from all five evaluators.Twenty-nine of 30 questions were rated Good by a majority, and every capability exceeded 87% except interpolation disclosure at 70%.
- End-to-End Human Evaluation: Only one of 150 responses was flagged for a fabricated number, while dominant failures involved map-pin placement and occasional wrong-tool calls.The evaluation reports end-to-end functional success and cross-trial consistency.
5 Case Study: A Blind Test on New Gold Discoveries
A retrospective blind case study tested MINETRACE on six Western Australian gold discoveries announced after the model’s 2021–2022 public-data cutoff. Most discoveries ranked highly, while the Mustang example illustrates how independent geological and geophysical evidence can support an inspectable prospectivity argument.
- Blind Discovery Test: Five of six discoveries fell in Western Australia’s top 13%, three reached the top 10%, and Astro exceeded the 98th percentile.Edjudina Range was the only mid-ranked site, near the 65th percentile.
- Blind Discovery Test: The discoveries were not part of the supervision set, so the results suggest MINETRACE was not merely rediscovering labelled deposits.The model concentrated prospectivity in areas where new gold systems were later reported.
- Mustang Worked Example: Mustang’s moderate-to-high score was supported by gravity-gradient, metasedimentary-host, Yilgarn-Craton, and structural-context signals despite weak local gold-geochemical evidence.Mustang had no effective prior drilling in the model data, and its nearest catalogued gold deposit was approximately 53 km away.
- Mustang Worked Example: MINETRACE surfaced a defensible mineral-systems argument for Mustang before the later drilling result was public, rather than knowing that the site contained gold.The evidence was presented for inspection by a geologist.
- Scope and Interpretation: The six discoveries are early-stage exploration results, and the sample is too small for statistical validation.Absolute scores are moderate in several cases, and high percentile rankings partly reflect generally low predicted gold prospectivity across Western Australia.
6 Conclusion
MINETRACE turns public exploration data into traceable evidence records connecting prospectivity scores to their underlying features, experts, sources, and coverage. This shared structure powers interactive interfaces and grounded natural-language interaction.
- Conclusion: MINETRACE connects final scores to measured features, named experts, expert weights, source layers, and coverage metadata.The system turns public exploration data into traceable evidence records.
- Conclusion: The same evidence structure powers the map interface, heatmaps, evidence panels, APIs, MCP endpoint, and grounded conversational assistant.These outputs remain connected to the underlying prospectivity analysis.
- Conclusion: MINETRACE turns a prospectivity model from a static score generator into an interactive reasoning workflow.The conclusion describes this design as an evidence-grounded interactive reasoning system for mineral prospectivity.
- Conclusion: The system demonstrates that interpretable domain models can support natural-language scientific interaction when their evidence structure is preserved end to end.
Limitations
MINETRACE supports decision-making from public evidence but is not autonomous or field-validated. Its outputs remain constrained by data quality, sampling coverage, model simplification, uneven commodity performance, and limited user-study evidence.
- MINETRACE ranks and explains public-data evidence but does not replace field verification or expert geological judgement.
- Scores are bounded by data coverage, assay quality, spatial sampling bias, and completeness of public reporting.
- Sparse local samples produce weak or interpolated support that the interface must flag.
- The fixed expert tree improves auditability but may miss commodity-specific mineral-system processes.
- Performance is uneven across commodities, with stronger results for Ni and Co but weaker results for Cu and Mn under spatial holdout.
- Grounding the conversational assistant does not guarantee correctness, and larger user studies are needed to assess geologist use in practice.
A Data and Preprocessing Details
The study uses a large, openly licensed Western Australian geoscience corpus and standardizes assay data for commodity-specific supervision. Preprocessing harmonizes heterogeneous records while preserving sampling-medium distinctions.
- 9,352,545 assay samples across five sampling media and 3,420 positive sites across eight commodities form the corpus.Positive-site counts range from 1,057 for Cu to 99 for W.
- Geochemistry, deposit labels, geophysical rasters, and geological vectors come from GSWA CM02, Mineralization Sites, and the 2021 CM08 Basemap.All products are openly licensed and standardized to GDA2020.
- CM02 records are mapped to a common schema with coordinates, sampling medium, and 123 element or oxide fields.
- The pipeline converts −9999 to missing, sets negative below-detection values to zero, separates five media-specific tables, and applies log(1+x).The transformation reduces skew.
- Commodity-specific Mineralization Sites exports are merged into one supervision table per target commodity, with polymetallic sites positive for each associated commodity.
B Neighbourhood Feature Catalogue
MINETRACE computes neighbourhood-based geochemical features at multiple spatial scales from log-transformed assay concentrations. The resulting feature definitions are composed into a prospectivity score by the model described in Section 3.3.
- Every feature summarizes samples within neighbourhood radii s ∈{5, 10, 50} km around the query location.
- Features use log(1+x) concentrations to reduce the effect of heavy censoring in assay data.
- Element-level features cover a fixed 14-element panel of target metals and common pathfinders.Ta and Mn therefore lack element-level features for the target element itself.
- Section 3.3 describes how the neighbourhood features are composed into a prospectivity score.
C Pathfinders and Expert Definitions
MINETRACE combines domain-defined commodity pathfinders with transparent expert families that aggregate geochemical, geophysical, and geological evidence. Pathfinder weights remain knowledge-based, while feature directions and expert weighting are fitted from data.
- Pathfinders: Each commodity uses its target element, exploration-geochemistry pathfinders, and element ratios for scoring.
- Pathfinders: Pathfinder weights encode diagnosticity from domain knowledge, while each feature’s favourable direction is learned during offline fitting.
- Expert definitions: Ten experts are organized into three families: three geochemical, four geophysical, and three geological experts.
- Expert definitions: Each geochemical expert aggregates independently fitted sampling-medium evidence weighted by training ROC-AUC, keeping the geochemical expert count fixed at three.
- Expert definitions: Geophysical and geological experts operate as single-source leaves over raster and vector layers, respectively.
D Human Evaluation Details
The human evaluation used 30 questions spanning seven user-facing exploration capabilities. Responses were organized by capability, with failures categorized systematically for analysis.
- 30 questions evaluated seven user-facing capabilities, each targeting a distinct part of the exploration workflow.The questions are grouped by capability in Table 5.
- Bad responses were assigned to ten failure categories covering data fabrication, disclosure, mapping, workflow, scope, comparison, explanation, ambiguity, and other errors.The categories distinguish factual, interaction, reasoning, and communication failures.
- Table 5 organizes the human-evaluation questions by capability rather than presenting them as a single undifferentiated test set.