Source-linked AI summary
ExploraTwin, a Non-Profit Research Platform for Digital Twin Simulations
Naveen Venkatanarayanan, Yuchen Qiu, Tianyi Peng, George Gui, Olivier Toubia
TL;DR
ExploraTwin addresses the practical friction of running digital twin simulations and supports survey and qualitative panel workflows, alongside a standardized format for persona banks. Across its survey-mode fidelity evaluation, 99.6% of 197,000 answer units were structurally valid on the first run, while the platform is offered to help researchers develop empirical evidence for their use cases.
Problem
Researchers need easier ways to test the validity of synthetic data and develop empirical evidence for their own use cases.
Method
ExploraTwin provides survey and panel pipelines, while CroissantTwin standardizes formats for adding digital-twin persona banks.
Results
99.6% of 197,000 asked question–answer units received an in-range, structurally valid response on the first run.
Takeaways & Limitations
ExploraTwin reduces the cost and friction of running digital twin simulations and helps researchers generate empirical evidence for their use cases.
Takeaways & Limitations
The platform does not provide any estimate of the validity of the data it simulates.
Abstract
from arXiv · showhide
Digital twin simulations show promise, but current empirical evidence suggests that the approach should be tested before being deployed in any particular context. To lower the friction for researchers and practitioners to test and deploy digital twin simulations, this brief commentary introduces ExploraTwin (https://exploratwin.org), an open-access, non-profit research platform for digital twin survey simulations. ExploraTwin supports two modes. In survey mode, researchers can upload a Qualtrics survey file or create a survey within the platform; select an available sample of digital twins; configure and run the simulation, and export analysis-ready data. In panel mode, researchers can assemble a small group of twins for open-ended conversations, document annotation, and moderated, focus-group-style voice discussions. We also developed CroissantTwin, a standardized data format for adding samples of digital twins to the platform. We demonstrate the survey mode workflow by using the platform to replicate 19 experiments on digital twins from the Twin-2K-500 dataset. ExploraTwin's survey execution fidelity is high: 99.6% of 197,000 answer units returned a structurally valid response on the first run.
THE EXPLORATWIN PIPELINE: SURVEY MODE
ExploraTwin’s survey mode converts researcher-created or uploaded surveys into LLM-readable simulations, lets researchers configure persona samples and execution, and validates and exports structured responses. It supports static and staged surveys while warning when uploaded features cannot be fully reproduced.
- Survey execution: Survey logic is preserved through one-call static execution or staged execution when later questions depend on earlier answers.In staged runs, ExploraTwin collects answers needed to resolve flow logic before presenting the remaining questions.
- Survey preparation: ExploraTwin accepts Qualtrics .qsf files or surveys created in an interactive builder, then converts them into a common parsed-template structure.The builder can start blank or generate a draft from a research brief; the QSF parser supports many Qualtrics question types and logic features.
- Scope and validation: The survey builder targets straightforward linear instruments, while complex logic such as randomization, branching, display logic, and piped text requires the QSF path.Unsupported uploaded features generate warnings before execution rather than being silently translated.
- Prompt construction: The simulation prompt combines behavioral instructions, a persona representation, an LLM-readable survey, and a parseable output schema.The survey component includes question text, options, stimuli, validation rules, and flow context.
- Response structure: The platform represents complex survey items as separate answer units and instructs the LLM to return one structured response for every unit.A matrix, multi-select, rank-order, or multi-statement slider can produce multiple answer units, though one model call may contain many questions.
- Configuration: Researchers configure the model, persona bank, representation kind, sample size, population filters, reporting options, and calling method before running a simulation.Batch calling takes longer but reduces token cost by roughly half.
THE EXPLORATWIN PIPELINE: PANEL MODE
Panel mode supports both independent qualitative elicitation and moderated group exchanges with digital twins. Researchers can select and filter personas, conduct text conversations or document reviews, and run voice discussions through TwinMeet.
- Panel-mode scope: Taken together, panel mode provides independent text and document-review interactions alongside moderated group exchanges, with conversations and associated materials available for analysis.The platform distinguishes text-based qualitative elicitation from shared-context voice discussion.
- Panel setup: Researchers select a persona bank, persona format, AI model, and filtered roster, with panels containing up to 10 personas.Available filters may include demographic attributes or custom traits from third-party CroissantTwin datasets.
- Text conversations: Text-based panel conversations operate as parallel interviews because each twin sees its own history but not other members’ responses.This prevents one twin’s response from directly anchoring another, while allowing follow-up questions to specific twins.
- Document review: Document review lets researchers upload images, PDFs, presentations, or text documents and receive annotations tied to passages, pages, cells, or visual locations.Annotations can be opened for direct follow-up exchanges with the twin that produced them.
- TwinMeet: TwinMeet provides moderated live voice conversations with up to five twins, using turn-taking and a shared transcript.The shared context allows twins to react to, build on, or disagree with earlier remarks, making the interaction closer to a moderated focus group.
CROISSANTTWIN: A STANDARDIZED DATA FORMAT FOR PERSONA BANKS
CroissantTwin standardizes how persona banks package personas, representations, filters, and responsible-use metadata. Conformant banks can be loaded into ExploraTwin without custom adapters, either publicly or privately at runtime.
- Motivation and standard: CroissantTwin addresses the costly need to write custom adapters for each persona-bank format by defining a unified packaging standard.It extends the MLCommons Croissant 1.1 dataset standard.
- Bank structure: A CroissantTwin bank contains data files and a manifest whose universal table and column labels allow automatic loading without custom code.The manifest also declares filtering and cataloging metadata.
- Data model: The data model distinguishes personas, representations, and representation kinds, allowing one persona to be rendered in multiple prompt formats.Researchers select the bank and representation kind when configuring a study.
- Responsible data use: CroissantTwin requires metadata describing persona basis, including observed individuals, human composites, synthetic personas, or mixed banks.Each persona row in a mixed-basis bank records its own basis.
- Distribution and access: ExploraTwin supports both a public persona-bank library and private runtime uploads that are discarded after execution.The library presents hosted banks in a common catalog, while runtime banks remain private to the active session.
- Supporting resources: The project lowers the cost of creating conformant banks through a public versioned specification and a coding-agent skill package that produces validated, upload-ready banks.The resources are intended to support conversion of individual-level datasets into CroissantTwin banks.
SURVEY-MODE FIDELITY: A 19-STUDY REPLICATION
ExploraTwin’s survey mode was evaluated by replicating 19 experiments with digital twins, using original Qualtrics files and standardized execution workflows. Across 197,000 asked answer units, first-run structural validity was high, while targeted repairs resolved nearly all execution failures.
- Replication design: 5,700 digital-twin response records replicated 19 experiments, with 300 respondents completing each study.The replications used Twin-2K-500 persona profiles and the full 500-question persona representation.
- Replication design: Original Qualtrics files were run through ExploraTwin using batch calling for static surveys and synchronous calling for surveys with dependent questions.Each study produced an export bundle identical to what a user would see on the website.
- Cost: The first-pass simulations cost $58.55 for 5,700 completed twins and 197,000 question-answer units, with $1.94 added for targeted repair.The total cost was $60.49, or about 1.1 cents per completed twin.
- First-run fidelity: 99.60% of 197,000 asked question–answer units received an in-range, nonblank response on the first run.The evaluation measured survey execution fidelity at the answer-unit level.
- First-run fidelity: 274 units (0.14%) were classified as survey execution failures, comprising 207 forced-response omissions and 67 out-of-range responses.The evaluation separately counted 521 optional blanks (0.26%), which were not classified as fidelity failures.
- Repair outcomes: 273 of 274 survey execution failures were resolved through rematching, targeted reruns, or a real-time retry.After repair, only one forced-response item out of 197,000 remained unanswered; optional blanks were mainly end-of-survey comment boxes.
GENERAL DISCUSSION
ExploraTwin reduces the cost and friction of digital twin simulation by providing a standardized workflow and extensible persona-bank format, while retaining important limitations around response validity, instrument support, and data validity assessment.
- ExploraTwin provides a standardized simulation pipeline covering survey construction or upload, prompt generation, model execution, validation, repair, and clean data export.
- The platform aims to help researchers experiment with digital twin simulations and develop empirical evidence for their own use cases.
- CroissantTwin standardizes the addition of synthetic-persona samples, supporting 11 persona banks and allowing new banks to be added or kept private.
- 0.14% of answer units required repair in the validation run, although post-hoc repair can re-ask questions after later survey information has been revealed.
- The platform accepts Qualtrics files and supports straightforward linear surveys, but unsupported custom JavaScript, dynamic displays, games, and interactive designs are flagged rather than executed.
- ExploraTwin does not estimate the validity of simulated data; instead, it makes testing synthetic-data validity easier in researchers’ own contexts.
POST-SIMULATION VALIDATION REPORT
The post-simulation validation report provides a compact, run-level assessment of survey simulations, covering configuration, fidelity, validity, randomization, descriptive results, and persona-bank coverage.
- Six validation components document run configuration, survey fidelity, response validity, randomization, descriptive results, and selected persona-bank coverage.
- Overall status classifies each study as ok, caution, or error based on problematic questions and failed calls.The status guides review while underlying flags remain visible.
PANEL MODE INTERFACES
Panel mode provides interfaces for text-based conversation, document review with location-based annotation, and moderated voice discussions through TwinMeet.
- Panel mode supports text-based conversations among digital twins.
- Panel mode supports document review and location-based annotation.
- TwinMeet provides a moderated voice interface for panel discussions.
SURVEY SUPPORT AND SIMULATION FIDELITY
ExploraTwin evaluates survey fidelity across question formats, settings, and flow features, preserving supported behavior while explicitly flagging unsupported or simplified features.
- Survey design layers: Survey fidelity is assessed across question format, question setting, and survey flow layers.These layers cover response tasks, local recording rules, and the sequence of questions participants see.
- Survey translation: Supported survey features are converted into an LLM-readable representation with structured questions, natural-language logic instructions, and visual stimuli when present.
- Support classifications: Features are classified as supported, partially supported, or not supported according to whether they are parsed, administered, constrained, and exported.
- Unsupported features: Unsupported features are detected at upload, warned about before the run, skipped during simulation, and recorded in the validation report.Purely visual formatting is listed as unreproduced without a warning because it changes appearance rather than respondent instructions.
- Boundary cases: Custom JavaScript is detected and reported but not executed because its browser-dependent behavior cannot be reliably recovered from the survey file.
- Boundary cases: Unsupported or inaccessible media, external services, browser-only interactions, and specialty question formats constrain what researchers can claim about a simulation.These limitations do not necessarily make a study unusable, but they are made explicit.
SIMULATION FIDELITY AND REPAIR
Across 19 replicated studies, ExploraTwin reports first-run response validity and repair outcomes while retaining original responses for auditability.
- Table A3 reports first-run response validity and repair outcomes across 19 replicated studies.
- Panel A distinguishes fidelity failures from blanks on optional questions, while Panel B reports responses resolved through rematching and targeted reruns.
- Original responses are retained so researchers can audit first-run performance.
- Optional blanks are counted as unanswered but reported separately from fidelity failures.F-SKIP denotes forced-question omissions, OPT optional blanks, and OOR out-of-range answers.
Brief Commentary: ExploraTwin, a Non-Profit Research Platform for Digital Twin
ExploraTwin combines structured response validation and repair with portable persona-bank standards, enabling more reproducible digital-twin simulations while preserving first-run data for inspection.
- Response validation and repair: Researchers can repair flagged cells through rule-based rematching or targeted rerunning, with repaired data stored separately from unchanged first-run responses.Rematching avoids a model call, while targeted rerunning re-asks unresolved questions and records the repair method.
- Response validation and repair: ExploraTwin validates every closed-ended response against the survey template using deterministic option matching, preserving unmatched responses as flagged rather than forcing them into the dataset.The matcher normalizes formatting and accepts only uniquely identified valid options; fuzzy matching is excluded from the initial run.
- Response validation and repair: Targeted reruns preserve prior valid answers and re-ask only unresolved cells, but repaired responses arise in a different information state from answers generated at the original survey position.Preserving first-run data supports inspection and sensitivity checks when sequential exposure matters.
- CroissantTwin and Build Skill: CroissantTwin defines a public, machine-readable, platform-independent structure for portable persona banks, allowing compatible tools to use them without custom code.The profile extends the MLCommons Croissant 1.1 standard and supports multiple representations of the same persona.
- CroissantTwin and Build Skill: CroissantTwin adds persona-specific metadata, including alternative representations, persona basis, approved filters, and disclosures about access and human-derived data.Its current implementation targets the versioned CT 0.1 working specification.
- CroissantTwin and Build Skill: The Build Skill converts an unfamiliar individual-level dataset and its documentation into a CroissantTwin bank, lowering the cost of creating new persona collections.The workflow is designed for coding agents and is self-contained.