Source-linked AI summary
A Simulator-Grounded Framework For Constructing Verifiable Muscle-Grounded QA From 3D Tongue Meshes (extended version)
Seungho Eum, Unsang Park
TL;DR
Existing articulatory corpora lack traceable pairings between observed tongue configurations and controlled muscle states. This paper introduces a simulator-grounded framework that constructs fact-verifiable QA from tongue meshes and evaluates it through unified language and structured prediction. The resulting supervision supports geometry-grounded learning and both natural-language QA and structured readouts, within a simulator-specific scope.
Problem
Existing rtMRI- and EMA-based corpora capture articulatory motion but do not directly pair each tongue configuration with a controlled muscle state, limiting traceable biomechanical supervision.
Method
The framework maps controlled 11-dimensional muscle activations through the ArtiSynth Badin model to screened meshes, structured records, and deterministic QA, with language naturalization separated from factual construction.
Results
The resulting supervision supports sample-specific geometry grounding, unified natural-language QA, and efficient structured prediction across complementary evaluations.
Takeaways & Limitations
3DTongueQA is a reusable, decoder-agnostic resource whose provenance is recoverable from simulator inputs to final QA.
Takeaways & Limitations
The resource remains simulator-specific, using one Badin anatomy, fixed topology, no jaw–lip coupling, and simulator labels that are not uniquely identifiable physiological causes.
Abstract
from arXiv · showhide
Existing articulatory corpora based on real-time MRI and electromagnetic articulography capture tongue shape and motion but do not provide traceable labels for the muscle-driven process that generated an observed configuration. We introduce a simulator-grounded data-construction framework and instantiate it as 3DTongueQA. Controlled 11-dimensional muscle activations are mapped to fixed-topology tongue meshes with the ArtiSynth Badin finite-element model, converted into structured biomechanical records, and rendered as deterministic QA on muscle state, geometry, and target-directed change. We screen 295,157 configurations, retain 295,115 valid meshes, and construct 891,156 QA records per language. Language naturalization changes only surface form and is verified against the source records; English and Korean instantiations demonstrate construction-level portability. A swappable SpiralNet++--Qwen3-8B baseline reaches 62.9 $\pm$ 9.2 Muscle EM, 74.0 $\pm$ 0.2 Value Accuracy, and 65.9 $\pm$ 4.7 Direction EM, while mismatching the paired mesh reduces Muscle EM to 2.2; a dataset-leakage-controlled anchor-held-out model retains 80.4--98.6\% of the full-inventory scores on unseen anchors. Task-specific structured readouts further reach 88.7 $\pm$ 0.7 Muscle EM and 93.3 $\pm$ 1.0 Direction EM. These complementary results show that the constructed supervision supports both efficient structured prediction and heterogeneous natural-language QA rather than being tied to a particular decoder architecture.
1. INTRODUCTION
3DTongueQA addresses the lack of traceable muscle-state supervision for observed tongue configurations by grounding QA construction in controlled biomechanical simulation. The framework preserves provenance while supporting English and Korean language instantiations and complementary evaluation modes.
- Existing rtMRI- and EMA-based corpora capture natural articulatory motion but do not directly pair tongue configurations with controlled muscle states.
- The framework maps controlled 11-dimensional muscle activations to screened 3D meshes, structured biomechanical facts, and deterministic QA on muscle state, geometry, and target-directed change.
- 295,115 valid meshes and 891,156 QA records per language instantiate the construction procedure in English and Korean.
- Language naturalization changes only surface form, while record-based checks preserve the underlying facts.
- The evaluation combines unified natural-language QA with structured readouts and controls that test whether supervision depends on paired geometry.
2. SIMULATOR-GROUNDED DATASET CONSTRUCTION FRAMEWORK
The framework converts controlled simulator inputs into valid tongue meshes, reusable biomechanical records, and fact-preserving QA. Its modular separation of factual construction from language realization supports new tasks and language instantiations without regenerating the biomechanical corpus.
- Simulator-Grounded 3D Tongue Generation: Controlled simulator inputs produce observable tongue geometry that is converted into reusable structured fact records before linguistic realization.
- Simulator-Grounded 3D Tongue Generation: The ArtiSynth Badin model maps 11-dimensional muscle-activation vectors to fixed-topology 3D tongue meshes.
- Simulator-Grounded 3D Tongue Generation: 295,115 valid meshes are retained after automatic checks for solver failure, element inversion, non-positive volume, global volume-ratio bounds, and settling behavior.
- Simulator-Grounded 3D Tongue Generation: The construction uses literature-informed vowel and consonant anchors, perturbations, and interpolations to create target-directed postures.
- Deterministic Grounded QA Construction: Each structured fact record contains simulator-defined muscle state, observable geometry, physical metadata, and intervention relations for Cause, State, and Change objectives.
- Deterministic Grounded QA Construction: Additional deterministic questions and metrics can be defined from stored records without regenerating the biomechanical corpus.Intervention-effect QA was instantiated in 3.2 s with zero new simulation.
- Deterministic Grounded QA Construction: Naturalization modifies linguistic surface form without changing facts, and English and Korean renderers verify outputs against the same structured records.First-generation record checks passed 87.2% of English and 88.6% of Korean outputs.
3. EVALUATION
The evaluation tests whether the supervision is learnable and geometry-grounded across Cause, State, and Change objectives. It uses anchor-balanced sample-held-out meshes and deterministic scoring, while explicitly not measuring extrapolation beyond sampled anchor neighborhoods.
- The evaluation covers active-muscle recovery, geometric-value prediction, and target-directed correction across Cause, State, and Change.
- The primary sample-held-out protocol tests learnability and geometry grounding within sampled anchor neighborhoods.
- The protocol does not measure extrapolation because configurations near training anchors are present by design.
- 400 anchor-balanced test meshes generate 1,200 automatically scored items, with each mesh paired with one question for each evaluation objective.
- The evaluation additionally includes 30 open-ended questions per model.
4. RESULTS
The evaluation finds that 3DTongueQA supervision is learnable, grounded in paired geometry, and effective for both unified language QA and structured prediction.
- 62.9 ± 9.2 Muscle EM, 74.0 ± 0.2 macro Geometric Value Accuracy, and 65.9 ± 4.7 Direction EM are achieved by the 3D Muscle-Aware unified-QA model.
- 2.2 Muscle EM after mesh shuffling shows that unified-QA muscle recovery depends on the paired mesh rather than only question or answer priors.
- 80.4%, 93.7%, and 98.6% of full-inventory Muscle, Geometric Value, and Direction scores are retained by the anchor-held-out model.
- 88.7 ± 0.7 Muscle EM and 93.3 ± 1.0 Direction EM are reached by task-specific structured readouts on the identical anchor-balanced test items.
- The Korean instantiation transfers the construction and verification procedure, although its scores use a separate evaluation set and are not direct cross-language comparisons.
5. CONCLUSION
The paper concludes that 3DTongueQA preserves provenance from controlled simulator inputs through factual records and linguistic realization, while supporting multiple prediction modes.
- 3DTongueQA separates controlled physical generation, structured factual representation, deterministic task construction, and fact-preserving linguistic realization.
- Experiments support both specialized structured prediction and unified natural-language QA, positioning the resource as reusable and decoder-agnostic.
- The resource remains limited to one Badin anatomy with fixed topology, no jaw–lip coupling, and simulator labels that are not identifiable physiological causes.
Compliance with Ethical Standards
Human evaluation involved three consenting adult expert raters assessing anonymized machine-generated outputs without collecting personal or physiological data.
- Three adult expert raters voluntarily assessed anonymized machine-generated outputs and provided informed consent.
- The evaluation collected no personal, sensitive, or identifying information and involved no patients, minors, or physiological data collection.
Supplementary Material A Simulator-Grounded Framework for Constructing Verifiable Muscle-Grounded QA from 3D Tongue
The framework samples controlled muscle activations, simulates and filters tongue meshes, and organizes valid configurations into diverse reasoning-oriented subsets under reproducible controls.
- Mesh representation: All samples share a fixed topology of 370 surface vertices, 736 triangular faces, 948 FEM nodes, and 740 volumetric elements.
- Biomechanical model: The model uses 11 muscles with per-activation bounds of 0.9 and total activation limits of 2.0 for vowels and 2.3 for consonants.
- Biomechanical generation: Each activation vector is applied through a 1.0-second minimum-jerk ramp with 24 knots, followed by adaptive settling for 0.2–1.5 seconds.
- Biomechanical generation: Validity screening records solver status, element inversion, volume ratios, residuals, and nodal velocities before retaining configurations.
- Sampling design: REST, SINGLE, PAIR, and TRIPLE sample increasingly complex muscle states, while ANCHOR, NEIGHBOR, EFFORT, and SPACEFILL target relations and feasible-space coverage.
- Reproducibility: Fixed seeds 0 and 42 govern generation and global shuffling, and duplicate vectors are removed after rounding activations to four decimal places.
A.3. Phoneme Anchor Construction and Acoustic Plausibility
Phoneme anchors provide literature-informed target settings for vowel and consonant configurations, with perturbation and constraint rules defining variation before biomechanical screening. Acoustic checks preserve relative vowel organization but do not establish speaker-independent absolute calibration.
- Phoneme Anchor Construction: Vowels are sampled as Gaussian perturbations around literature-informed 11-D targets, with noise applied only to nonzero entries.The /æ/ target is merged into /a/ because their simulated midsagittal meshes differ by only 0.9 mm vertex RMS.
- Phoneme Anchor Construction: Consonants use defining-muscle intervals, sparse background activation, clipping to [0, 0.9], and rejection sampling after budget checks.The consonant activation budget is P_m a_m ≤2.3, with defining constraints rechecked after non-defining muscles are reduced.
- Acoustic Plausibility: Category-level simulated formants correlate with measurements at r = 0.98 for F1 and r = 0.99 for F2.The relative vowel organization is retained, but absolute F2 is compressed, especially for front vowels, because of fixed-boundary and midsagittal modeling assumptions.
- Biomechanical Screening: 295,157 activation configurations were screened, with 295,115 VALID meshes retained for downstream corpus construction.Retained items store activation vectors, surface vertices, FEM nodes, shared topology, validity metadata, and applicable links.
- Acoustic Plausibility: Target-directed QA uses simulator-defined anchor adjustments, while the factual records and naturalized renderings are checked for consistency.The faithfulness gate checks numerical values, muscle names, phonemes, regions, directions, relation signs, and abstentions against source records.
B.3. Post-Hoc Task Instantiation Without Regeneration
The stored biomechanical records support deterministic post-hoc task construction without rerunning simulation. An intervention-effect task adds complementary supervision because the existing unified model performs below trivial baselines zero-shot.
- Post-Hoc Task Instantiation: A NEIGHBOR record links a base configuration, one changed muscle, and a signed ±0.15 delta for intervention-effect questions.Gold answers compare stored feature values under the same dead bands used for geometric-value scoring.
- Post-Hoc Task Instantiation: 400 balanced intervention-effect instances were sampled from 990 admissible pairs in 3.3 seconds with zero additional simulation.The classes were distributed as 360 no-change, 20 increase, and 20 decrease instances.
- Evaluation: 59.8% zero-shot accuracy fell below the 90.0% majority-class baseline on the 400-instance set.On a class-balanced 120-instance subset, performance was 27.5%, near the 33.3% chance level.
- Evaluation: The post-hoc task was not solved by the existing supervision, indicating complementary rather than redundant task supervision.All automatically scored answers are derived directly from structured records, while undefined facts are skipped or assigned explicit abstentions.
C.1. Models, Training Hyperparameters, and Evaluation Protocol
The evaluation uses a frozen SpiralNet++ mesh encoder and Qwen3-8B language model connected by trainable projection and LoRA components. It measures active-muscle recovery, geometric-value prediction, and target-directed correction on held-out meshes with exact or thresholded scoring.
- Models: SpiralNet++ encodes rest-relative displacement on a fixed-topology 370-vertex surface through a 370 →93 →24 hierarchy.The final level provides 24 local tokens of dimension 16, alongside a 32-D global latent.
- Models: One global and 24 local projected vectors form a 25-token mesh prefix placed before the question in Qwen3-8B.The encoder and Qwen3-8B remain frozen; only the projector and LoRA adapters are updated.
- Training: Training one unified-QA seed requires approximately three days on two NVIDIA H200 GPUs.The reported three-seed results therefore correspond to roughly nine H200-pair-days per model variant.
- Evaluation Protocol: The primary evaluation uses 400 anchor-balanced sample-held-out meshes and 1,200 items spanning Muscle, Value, and Direction tasks.Muscle EM requires exact active-set matching, while Direction EM requires exact matching of unordered (muscle, ↑/↓) pairs.
- Evaluation Protocol: Geometric Value Accuracy uses feature-specific dead bands: 1.0 mm for metric values, 0.1 for normalized quantities, and 0.03 for constriction degree.Region-movement facts use a 0.5 mm displacement threshold, and the score is an unweighted macro-average.
C.2. Additional Automatic and Human Analyses
Additional analyses test decoder roles, representation grounding, leakage-controlled transfer, and human judgments. Together, they show strong structured readouts, nontrivial learned information, controlled anchor-held-out performance, and measurable human-evaluation differences, within a bounded simulator scope.
- Decoder-agnostic utility: 88.7 ± 0.7 Muscle EM and 93.3 ± 1.0 Direction EM are achieved by task-specific structured readouts, exceeding unified QA scores on identical tasks.The geometric structured readout also reaches 87.5 ± 1.1 macro Geometric Value Accuracy versus 74.0 ± 0.2 for unified QA.
- Representation analysis: 9.8 Muscle EM from directly thresholding the pretrained activation-regression head rises to substantially higher performance after readout training.The direct baseline uses no QA or readout training, indicating that readout training extracts additional information from the frozen representation.
- Anchor-held-out stress test: Near-anchor closure and maximin selection produce the held-out categories /o/, /U/, /l/, and /S,Z/ without using final-model scores.The selection excludes nearby anchors and labels or aliases appearing in prescriptive training text.
- Human evaluation: Factual Accuracy shows a model effect in prompt-level Friedman tests, with χ2(4) = 58.10 and p < 0.001.Three model-blind speech researchers rated Fluency and reference-based Factual Accuracy on 1–5 scales; agreement was ICC(2, k) = 0.816 and 0.831, respectively.
- Scope and limitations: 3DTongueQA supports simulator-grounded articulatory reasoning but does not constitute clinical validation or replace diagnosis or treatment.Its scope is constrained by one Badin anatomy, fixed topology, symmetric midsagittal representation, fixed lips, and uncoupled jaw motion.