Source-linked AI summary
Praxist: From Experimental Artifacts to Solution Lineages
Jin Li, Ahmed Murtadha, Zhiyu Wang, Qiwen Chen, William Chen, Yifei Wu, Guan Wang, Andy L. Siy, Jiayi Yang, Mengsha Huang, Wenhao Li, Yixuan Liu, Shuailin Pan, Mingli Yuan, Sen Song, Yuhao Sun
TL;DR
Autonomous R&D systems often retain evaluated experience without making its mechanisms, validation status, and recombination value explicit. PRAXIST converts reproducible artifacts and evaluator outcomes into typed, lineage-grounded evidence that directs later generations. On 75 MLE-bench tasks, it obtains 60 medals (80.0%) versus 55 (73.3%) for Claude Code on Claude Opus 4.8, while the recorded spend is approximately US$3,054 versus US$38,370; case studies extend the process to open-ended engineering problems.
Problem
Artifact-level search retains extensive evaluated experience but can prune components before their value emerges through combination, limiting what later attempts inherit.
Method
PRAXIST converts evaluated artifacts into typed findings, lane-structured frontier evidence, agendas, Gems, and an inspectable lineage that guides subsequent generations.
Results
60 medals (80.0%) across 75 MLE-bench tasks, including 49 gold, versus 55 medals (73.3%) and 34 gold for Claude Code on Claude Opus 4.8.
Takeaways & Limitations
PRAXIST couples evaluated artifact improvement with an inspectable solution lineage, while the four case studies report detailed discovery paths for open-ended engineering results.
Takeaways & Limitations
The full-horizon precision metric in the tokamak case is confounded by longer survival, requiring a common-horizon recomputation that discards late-horizon steps.
Abstract
from arXiv · showhide
Autonomous R\&D agents now write, run, and improve executable artifacts under automated evaluation---but largely as laboratory instruments: shown on curated benchmarks, with gains that are hard to trace to a cause and costs well above what sustained engineering practice absorbs. The limitation is structural. Most systems treat each attempt as nearly self-contained, so logs, memories, and search trees record what happened without establishing which design element produced an improvement, whether its evidence survived validation, or how it recombines with others. Long campaigns therefore keep re-learning the same lessons. We introduce Praxist, a lineage-centered generational system that converts reproducible artifacts and evaluator outcomes into a typed evidence graph of findings, lane-structured frontiers, and agendas. Separating local artifact construction from cohort-level evidence synthesis lets later attempts inherit validated mechanisms, unresolved claims, and useful constraints, and leaves results attached to an inspectable lineage. On the standardized 75-task MLE-bench suite, the finalized official-grader results give Praxist 60 medals (80.0\%), 49 of them gold, against 55 medals (73.3\%) and 34 gold for a Claude Code baseline on Claude Opus 4.8---at a recorded model spend of US\$3,054 versus US\$38,370, roughly a twelfth of the cost. Four case studies---quantitative trading, LiDAR-inertial-visual SLAM, tokamak magnetic control, and rocket landing---carry the same process into open-ended engineering problems, improving on each task-native baseline in headline accuracy, survival, or resource cost, with the discovery path on record. Stronger artifacts at an order of magnitude less spend, each backed by an auditable lineage, are, to our knowledge, first brought together here: the operating profile production research requires, not the one a benchmark demonstration establishes.
1 Introduction
Long-horizon autonomous R&D systems accumulate evaluated artifacts but retain too little actionable, recombinable evidence. PRAXIST addresses this by converting outcomes into typed evidence and lineage for subsequent research.
- Motivation: Artifact-level search can prune components before their utility emerges through combination, limiting what later attempts inherit.Campaigns generate extensive evaluated experience, but only a limited portion is retained in a form later attempts can build on.
- Motivation: Evidence inheritance retains evaluated outcomes as reusable elements whose roles and support determine how they influence later construction.This assembly-oriented primitive amortizes retained elements across later constructions rather than requiring search to anticipate every useful variant.
- PRAXIST: PRAXIST formulates selective retention and recombination of prior evidence as a systems requirement for long-horizon evaluator-grounded R&D.The formulation connects cumulative construction from reusable parts with future artifact development.
- PRAXIST: PRAXIST converts evaluated artifacts into typed, inheritable evidence, including findings, frontier lanes, agendas, and Gems, while maintaining an inspectable lineage graph.The system makes the resulting research state active during research and inspectable afterward.
- Evaluation: PRAXIST is evaluated across all 75 MLE-bench tasks and four open-ended case studies against local baselines.The evaluation scores outcomes, lineage process measures, and cost.
2 Method
PRAXIST runs autonomous R&D as a generational artifact-to-lineage process, selectively inheriting typed evidence through frontiers, agendas, and durable lessons. Local experimentation produces evaluated findings, while global synthesis decides what survives, records its lineage, and directs subsequent construction.
- Overview: PRAXIST materializes each attempt as a reproducible artifact, evaluates it externally, interprets the result into findings, and reports the final artifact with its production lineage.The lineage records artifacts, findings, decisions, agendas, memory, and typed relations between them.
- Evidence representation: Typed findings replace raw transcripts or scalar scores by assigning evidence operational roles such as reuse, validation, avoidance, diagnosis, preservation, or archiving.Findings also carry a type and evidence maturity, linking inheritance decisions to the thoroughness of evaluation.
- Generational cycle: Each generation reads inherited frontier evidence, an agenda, and Gems, runs a cohort of parallel peers, then emits a successor state through synthesis and inheritance.The generational state includes the frontier Fg, agenda Ag, Gems Gg, and accumulated lineage trace Lg.
- Global synthesis: Global synthesis sorts surviving evidence into confirmed, candidate, diagnostic, and validation lanes, preserving reliability, failure information, and validation priorities alongside measured performance.The Chair converts these decisions into per-direction agendas to continue, validate, stop, or explore.
- Design allocation: The Deep Innovation Gate assigns each peer a testable mechanism, intervention, parent lineage, evidence signature, validation hook, and forbidden changes before construction.Design allocation distributes contracts across distinct cells under mechanism, intervention-surface, and intent caps.
- Memory and failure handling: Negative and diagnostic findings remain inheritable constraints, while periodic memory compression distills recurring lessons into a bounded set of durable Gems.Gems may encode validated mechanisms, rejected assumptions, recurring failure modes, or procedural constraints; compression is optional and lane-balanced.
3 Experiments
PRAXIST is evaluated across MLE-bench and open-ended engineering case studies, combining outcome quality with cost and lineage evidence. It exceeds the cited benchmark baseline, achieves strong rocket and quantitative-finance results, and records how improvements accumulated.
- MLE-bench: 81.7% of PRAXIST’s medals are gold, compared with 61.8% for Claude Code, although both results are single locally measured runs rather than multi-seed estimates.The reported margins should therefore be interpreted as benchmark-wide outcomes under the stated reporting protocol.
- MLE-bench: On 19 selected Medium- and High-tier tasks, PRAXIST wins the better score on 12 and records 11 gold outcomes versus 5 for Claude Code.Both arms medal on 14 tasks; the selected-subset difference is concentrated in medal grade, while the full-suite raw-score comparison is closer.
- MLE-bench: US$3,054 versus US$38,370 in recorded 75-task spend separates PRAXIST from the Opus 4.8 sweep by roughly an order of magnitude.The cost comparison uses the same task suite, while medals remain determined by task scores and thresholds.
- Rocket: 100% landing success on 12,288/12,288 rocket trajectories improves over the 4.03% starting artifact and Weco’s reported 17.12%.The controller also reduces 95th-percentile sink speed, lateral speed, and tilt, while increasing gimbal total variation and roll-to-pitch/yaw coupling.
- Quantitative finance: 53% walk-forward CAGR versus 23% for the paired baseline gives the discovered quantitative policy a 2.3-fold growth-rate ratio across 28 windows.The policy is positive in 26 of 28 quarters, and its mean quarterly return remains positive under an additional 50 basis points per executed side.
4 Conclusion and Discussion
PRAXIST combines stronger benchmark and open-ended R&D outcomes with inspectable solution lineages. Its results suggest a research-collaborator profile spanning performance, reuse, and cost-aware discovery.
- Results: 60 medals (80.0%) across 75 MLE-bench tasks, including 49 gold, compared with 55 medals (73.3%) and 34 gold for Claude Code.The baseline uses Claude Opus 4.8.
- Open-ended case studies: Four case studies beat task-native baselines in headline accuracy, survival, or resource cost while recording their discovery paths.The cases cover trading, SLAM, tokamak magnetic control, and rocket landing.
- Results: 49 of 60 medals are gold (81.7%), indicating that most successful MLE-bench outcomes reach the highest medal tier.The paper reports this as a property distinguishing PRAXIST as a research collaborator.
- Lineages: Solution lineages expose the mechanisms, controls, and failures behind results, supporting reuse and extension by human scientists.The four case studies report these lineages in detail, while MLE-bench reports only graded outcomes per task.
- Scope and directions: PRAXIST targets model and algorithm development, controller synthesis, simulation-driven engineering design, and quantitative strategy research.The paper also proposes cross-campaign inheritance, slower or noisier evaluators, and deeper human–AI collaboration as future directions.
Code and Released Runs
The implementation, project page, and released runs are publicly available. The archive preserves generational records so reported trajectories can be inspected beyond aggregate scores.
- Availability: The implementation, project page, and runs behind the reported results are publicly available.The listed resources include the GitHub repository, project page, and released-runs archive.
- Archive contents: The released archive contains artifacts, evaluations, findings, frontier lanes, agendas, and lineage traces for each run.These records preserve the generational research state.
- Inspectability: Generation-by-generation records allow inspection of research trajectories rather than only aggregate scores.The archive supports inspection of the trajectories summarized in the paper and appendices.
- Task-family interfaces: Task-family realization components instantiate MLE-style tasks with competition context, submission artifacts, metric semantics, and grader outcomes.Other task families use their own artifact forms and evaluators.
A Method details
The appendix specifies PRAXIST’s methodological interfaces and state objects rather than its software mechanics. It defines how artifacts, findings, frontiers, agendas, Gems, and lineage records organize the research cycle.
- State representation: PRAXIST’s state transition is Sg = (Fg, Ag, Gg, Lg), with artifacts grounding attempts and lineage recording how the final artifact was produced.The appendix uses the same symbols and vocabulary as the main text.
- State objects: Findings interpret evaluated evidence, frontiers determine inheritance status, agendas direct later generations, and Gems preserve durable lessons.These objects define the roles represented in the generational state.
- Methodological scope: The appendix emphasizes methodological interfaces—what information is represented, which role consumes it, and what each role emits.Command invocations, storage layouts, and release-specific engineering details are left to implementation documentation or protocol descriptions.
- Generality: MLE-style tasks are a running example of a task-family realization, not the definition of PRAXIST.The method is presented independently of that single task family.
A.1 Task-Family Realizations
PRAXIST uses a task-agnostic research cycle instantiated by task-family-specific artifacts, evaluators, constraints, and evidence conventions. The common requirement is that evaluated artifacts produce traceable evidence.
- Realization interface: A task-family realization supplies task context, artifact form, evaluator semantics, role constraints, and evidence conventions.These components instantiate the common research-state objects across evaluator-grounded tasks.
- Task agnosticism: The same research-state objects can operate across different evaluator-grounded tasks because task-specific information is separated from the cycle.Table 8 lists the components a realization must supply.
- Examples: MLE-style realizations define competition description, data context, submission format, metric name and direction, and evaluator semantics.These instantiate the general realization interface for machine-learning engineering tasks.
- Examples: Proof search, software engineering, and simulation can instead use proof artifacts with checker outcomes, patches with test reports, or simulation bundles with score reports.The artifact and evaluator change by task family.
- Common requirement: Evaluated artifacts must yield traceable evidence across task families.This is the common requirement shared by the examples.
A.2 Artifacts, Evaluation Records, and Findings
PRAXIST turns evaluated artifacts into typed, maturity-aware findings and uses frontiers, agendas, and lineage to govern what later experiments inherit.
- Evidence records: Each attempt is represented by a reproducible artifact, an evaluation record, and findings that interpret the outcome as research claims.Findings carry maturity and inheritance recommendations, while conceptual evidence records can be instantiated differently across task families.
- Artifact status: Only committed artifacts provide primary sources for positive inheritance; partial and failed artifacts may still yield diagnostic or uncertain findings.Superseded artifacts are replaced as the current inheritable state.
- Evaluation: PRAXIST separates validity from score, treating an invalid high score as diagnostic rather than confirmed improvement.Evaluation records preserve protocol, metric, score, validity, evidence stage, lineage, and interpretive limitations.
- Finding types: Findings distinguish positive, negative, diagnostic, uncertain, and procedural roles, with evidence maturity kept distinct from outcome quality.These roles preserve reusable mechanisms, weakened assumptions, constraints, failure modes, and experimental requirements.
- Contracts and diversity: DIG makes each intended experiment explicit, while quantified diversity allocates contracts across distinct mechanism, intervention, and intent cells.The gate specifies the tested mechanism, intervention site, prior evidence, supporting or weakening evidence, and invalidating changes.
- Cohort synthesis: Peers publish local artifacts and findings; PI roles interpret them, and the Chair synthesizes an agenda with tests, claim boundaries, contracts, and archive decisions.The frontier assigns operational inheritance status, while Gems optionally compress durable lessons across longer horizons.
- Lineage and reporting: The final output is a reproducible artifact accompanied by lineage records showing supporting evidence, constraints, validation needs, Gems, and directing agendas.Reporting selects from the completed frontier and lineage outputs rather than invoking the selection rule inside the research loop.
A.8 Resource Scheduling and Mature-Evidence Debt
PRAXIST combines GPU admission and idle backfill with mature-evidence debt control, so scheduling addresses both hardware utilization and timely high-grade evidence.
- Resource admission: GPU admission considers declared utilization and peak memory, permits compatible co-location, and acquires multi-GPU experiments as gangs.Placement limits compute-load and memory fragmentation, with either dimension able to block admission.
- CPU contention: CPU demand affects modeled runtime through processor-sharing slowdown rather than per-experiment core allocation or rejection.The deployed scheduler additionally consults observed host pressure as a coarse launch gate.
- Maturity: Mature evidence requires full-stage execution with sufficient independent evaluation coverage; simulator thresholds are completed-work ratio ≥0.75 and coverage ratio ≥0.80.Scout-stage runs remain preliminary probes rather than mature results.
- Mature-evidence debt: The controller sets Q = max(1, ⌈C/4⌉), tracks completed mature results M_t, and launches up to K_t = min(C, 3D_t) mature-directed experiments.The threefold target provides bounded redundancy against failures and heavy-tailed durations, with recomputation after completion, failure, or resource release.
- Simulation protocol: The simulation spans 512 scenarios and 4,096,000 policy runs across heterogeneous peers, GPUs, memory, cores, durations, failures, costs, correlations, and horizons.The reported comparison uses 348 physically feasible scenarios.
- Simulation results: Equal utilization does not ensure equal evidence supply: Boolean and thin-token feedback trail the debt controller by 1.20 and 1.49 quota-success points.Under high CPU pressure the gap reaches 3.75 points, and under high GPU-memory pressure 4.43 points.
- Boundary and caveats: Production adds a host-pressure guard that reduces concurrency above roughly 92% CPU or 95% memory utilization and withholds launches under sustained pressure.CPU cores are not a declarative per-experiment admission dimension, and the study documents this as a production departure from the simulation.
A.11 MLE-Style Task Example
MLE-style tasks instantiate PRAXIST with submission-centered reproducible artifacts, task-provided evaluation, artifact-grounded findings, and inheritable frontiers.
- Task instantiation: The task context supplies the competition description, data, submission format, metric name, and metric direction for artifact construction.A peer builds a submission file and the supporting files required for reproduction.
- Evaluation and findings: The external evaluator returns task-grounded outcomes, while findings preserve raw evidence, limitations, and inheritance recommendations.This keeps interpretation attached to the evaluated artifact rather than to an ungrounded agent assessment.
- Generalization boundary: Other task families can change artifact forms and evaluator semantics while retaining traceable artifacts, typed findings, explicit inheritance, agenda control, selective memory, and lineage accumulation.The PRAXIST cycle is therefore preserved across task-family realizations.
A.12 Worked MLE-bench Research Trajectory
The Jigsaw trajectory records how evaluated parents, diagnostics, ensembles, and negative branches combine across generations while preserving the lineage to the selected artifact.
- Generation 0: 0.00037 improvement: generation 0 retained the matched BCE ablation after testing class-balanced focal loss.The task’s primary metric is column-wise ROC AUC, with higher values better.
- Generation 1: 0.98723: averaging the BCE parent with a second BERT seed and adding DistilRoBERTa produced the generation-1 ensemble branch.The trajectory reuses preserved prediction artifacts after a completed null-model diagnostic.
- Branch comparison: 0.00012 decrease: adding the weaker BiLSTM component reduced the score, preserving a negative branch in the lineage.Adding the independent RoBERTa-CLS parent instead raised the score by 0.00026 to 0.98749.
- Final ledger: 0.9880: the release ledger reports Jigsaw’s later highest-scoring integrity-clean attempt, distinct from the worked generation-1 trace.The worked trajectory illustrates structure rather than the finalized campaign entry.
B Experimental Setup Details
The appendix specifies how replication depth, execution conditions, and campaign cost support the reported results.
- Replication details state how many repetitions, hardware, time envelope, and campaign settings underpin each reported number.The intent is to make replication depth and headline-result cost directly readable rather than inferred from the narrative.
B.1 MLE-bench
The experimental setup defines shared GPU execution, study-level generation accounting, model-spend reporting, and a frozen rocket protocol with fixed initial states, actuation, and evaluation scope.
- Each task run used one NVIDIA H100 80GB GPU, with up to eight concurrent experiments, 24- or 36-hour caps, and a 3,600-second finalization grace period.Seventy tasks used the 24-hour cap and five used the 36-hour cap; these were admitted ceilings rather than measured runtimes.
- Table 29 counts committed generation boundaries for five studies, except Fusion, whose legacy trajectory records completed generation result sets.
- US$3,054 in model spend resulted from mixed DeepSeek V4 Pro and V4 Flash token schedules with cache-hit, cache-miss, and output pricing.The reported total combines CNY 20,694.84 across the two schedules.
- The Rocket protocol freezes the plant, initial state, evaluator integration, and actuation contract while allowing candidates to change only the controller.C05 uses a six-degree-of-freedom rigid body, RK4 integration, a 0.1 s step, and at most 900 steps.
- Complete private validation comprises 13,312 units: 12,288 landing trajectories across three banks plus 1,024 frozen roll-disturbance cases.This canonical scope is used for matched baseline comparisons, promotion, and post-run reporting.
B.2.1 First-contact controller discovery, matched evaluation, and full-bank audit
The Rocket study discovers a deterministic hybrid controller through lineage-guided guidance and allocation changes, then evaluates it under matched complete-protocol validation and a separate full-bank audit.
- First-contact controller discovery: The controller combines rolling zero-effort-miss guidance, a fuel-commit governor below 450 m, and committed descent and phase-release guards.The fuel governor activates when remaining main fuel falls below 80%, while the radius-conditioned branch changes release behavior for r0 < 450 m.
- First-contact controller discovery: The newest lineage contribution is a closed-form allocator solving independent box-constrained quadratic programs to split pitch and yaw torque between gimbals and grid fins.
- Matched evaluation: 12,288/12,288 complete-protocol successes belong to the committed generation-11 configuration, while the audited generation-12 controller is byte-identical in its controller implementation.
- Matched evaluation: 100% matched-protocol landing success replaced the baseline’s 4.0283%, a gain of 95.9717 percentage points on 12,288 fixed landing cases.The corresponding descriptive 95% Wilson lower bounds were 3.6948% and 99.9687%; neither is a post-selection population guarantee.
- Matched evaluation: The baseline already had 100% first-contact rate, so the improvement concerns contact quality rather than reaching contact.Baseline fuel depletion was 88.6393%, and baseline sink-speed P95 was 66.3928 m s−1.
- Matched evaluation: Gimbal total variation rose by 80.43% and roll-to-pitch/yaw coupling P95 by 87.55%, despite 100% roll-stability success on 1,024 cases.These shifts are consistent with moving pitch and yaw activity from grid fins onto the gimbal.
- Full-bank audit: The post-run full-bank audit evaluated 122,880 trajectories without using its output for selection, ranking, or promotion, and without a same-scale baseline.
- Full-bank audit: Exactly two audited trajectories failed, both only on the lateral-speed condition; every other audited trajectory passed the joint predicate.Both failures occurred in the [0, 450) m initial-radius bin, whose nominal and near-OOD rows each scored 3,686/3,687 = 99.972878%.
B.4 SLAM
The SLAM comparison uses accepted runs over fourteen sequences and reports resource and full-trajectory effects, while several protocol and coverage choices constrain interpretation.
- The comparison reports one accepted run per method–sequence pair across fourteen sequences, with no training seeds or run-to-run variance estimate.Several sequences were replayed or rerun, and acceptance was not governed by a pre-declared rule.
- The champion configuration uses PVTR_MODE=8, an observability-gated visual scheduler, and VMAP_DEDUP=1, a visual-map admission filter.
- APE uses ground-truth-associated poses whose coverage differs between arms, with COVSCHED’s associated-sample ratio ranging from 0.79× to 1.76× across sequences.Full-sequence comparisons should therefore be read with coverage.
- Short-window relative pose error is mixed across sequences, and component controls cover only one to three sequences rather than all fourteen.The rate-matched periodic-skip control also changes two respects at once by disabling the map-admission filter.
- Sampled campaign pods exposed 168 logical cores, roughly 25 GiB of memory, ROS Noetic, and no directly visible NVIDIA device, so timings are indicative rather than isolated-benchmark measurements.
- The campaign used DeepSeek V4 Pro, eight peers per generation, up to 200 generations, two promotions per generation, and five-hour generation windows.No compute budget, Gem configuration, evaluator seed list, or per-experiment GPU budget was declared in the effective specification.
B.5 Fusion
The Fusion case study evaluates synthesized Torch control code across repeated scenarios and seeds, while showing that survival results depend materially on the environment lifecycle used for evaluation.
- Evaluation protocol: 15 episodes and at most 1,500 survived simulator steps define the Fusion evaluation protocol across 5 scenarios, 3 seeds, and a 100-step horizon.The evaluated artifact is synthesized Torch control code, not a trained policy; meaningful iteration counts are generations, peer sessions, episodes, and simulator steps.
- Resource accounting: 8 GPU-hours per experiment is a scheduling reservation rather than a realized-compute measure because FreeGSNKE evaluation is CPU-heavy and process-parallel.Four H100 devices were visible on the host, but that hardware context does not establish workload occupancy.
- Campaign configuration: At most 12 generations with 5 peers and two promotions per generation produced the HybridJacobianPDV1 lineage in a local DeepSeek V4 Pro campaign.The campaign used a nominal 6-hour generation window, with each experiment declaring 8 GPU-hours and 20 GB of memory.
- Metrics: The Fusion table reports full-horizon and common-horizon WNRMSE p95, with lower values preferred and common-horizon scoring capped at the zero-feedback baseline’s survival length.Aggregate rows pool all scored steps rather than averaging across scenarios.
- Runner sensitivity: Under fresh reset with action clipping, the selected controller survives 1,264 steps at a 0.667 full-horizon completion rate, while a second fresh validation gives 1,298 steps and 0.733.The passage states that the default in-run evaluator uses a different reset lifecycle, so Fusion figures must be read with their runner.
C Full MLE-bench Per-Task Results
The full MLE-bench results compare PRAXIST with Claude Code + Opus 4.8 across all 75 competitions, using medal outcomes, metric-direction score comparisons, and integrity-adjudicated official attempts.
- Coverage: Tables 36 and 37 provide the complete per-task comparison between PRAXIST and the Claude Code + Opus 4.8 baseline across all 75 competitions.The tasks are grouped by category, with Table 36 covering image and text classification-related categories and Table 37 covering the remaining categories.
- Reading conventions: Medal shading distinguishes gold, silver, bronze, and no medal, while bold identifies the better score in the task metric’s direction.Bold assignment uses full-precision values, so displayed-equal scores may still differ.
- Organization: Tasks are sorted alphabetically within each category and numbered consecutively from 1 to 75 for cross-reference throughout the paper.The # column supplies the display-order index used for cross-reference.
- Integrity and selection: 60 tasks retain PRAXIST’s best official clean attempt under the corrected release ledger’s integrity adjudication.Selected submission payloads are verified by SHA-256 against the run journal.