Source-linked AI summary
A Metamorphic Artificial Age Score Decision-Support Prototype for Flight-Log-Based Drone Propeller Health Monitoring
Seyma Yaman Kayadibi
TL;DR
Drone propeller health monitoring needs interpretable post-flight decisions because faults can affect multiple flight-log channels rather than one diagnostic signal. This prototype transforms selected DronePropA logs into normalized indicators and AAS-based maintenance recommendations, distinguishing routine monitoring, maintenance review, and mandatory inspection across controlled healthy and defective cases.
Problem
Drone propeller health monitoring lacks interpretable decision-support outputs that explain which flight-log channels drive maintenance decisions beyond binary fault detection.
Method
The prototype extracts six indicators from selected DronePropA logs, normalizes them to a healthy baseline, and evaluates scoring policies using metamorphic adequacy and redundancy-adjusted AAS.
Results
The healthy case received routine monitoring, Severity 1 maintenance review, and Severity 2 and 3 mandatory inspection based on different dominant burden channels.
Takeaways & Limitations
Propeller fault effects may appear through different operational channels, supporting multi-indicator decision support for post-flight maintenance prioritization.
Takeaways & Limitations
The controlled retrospective subset contains one healthy baseline and three defective cases from the same fault group, speed profile, and trajectory, requiring broader validation before operational use.
Abstract
from arXiv · showhide
Drone propeller faults can create safety and reliability risks when their effects are distributed across multiple flight-log channels rather than appearing as a single diagnostic signal. This paper proposes a Metamorphic Artificial Age Score (AAS) decision-support prototype for flight-log-based drone propeller health monitoring. Using selected historical real flight logs from the 2024 DronePropA public dataset, the framework computes six health-related indicators from raw MATLAB matrices: trajectory tracking error, attitude instability, thrust-command burden, motor-command imbalance, ESC-command instability, and battery-level stress. These indicators are normalized relative to a healthy baseline and evaluated through candidate scoring policies, metamorphic adequacy relations, and a redundancy-adjusted AAS formulation. In this context, AAS is used as a structural policy-adequacy and burden measure rather than as a chronological age measure. A controlled retrospective evaluation was performed using one healthy baseline and three defective propeller cases under the same speed profile and trajectory. The healthy case was assigned to routine monitoring. The Severity 1 case was dominated by ESC-command instability and assigned to maintenance review. The Severity 2 case reached maximum motor-command and ESC-command burden, while the Severity 3 case reached maximum trajectory tracking error; both triggered mandatory inspection. The results show that propeller fault effects may appear through different operational channels, supporting the need for a multi-indicator decision-support layer for post-flight maintenance prioritization and autonomous-system oversight.
1 Introduction
The paper proposes an interpretable metamorphic AAS decision-support prototype that converts real DronePropA flight logs into six health indicators and evaluates scoring-policy adequacy. It addresses distributed fault effects by separating operational burden, maintenance recommendations, and structural policy assessment.
- Motivation: Propeller faults may distribute their effects across multiple flight-log channels instead of appearing as a single diagnostic signal.DronePropA provides historical real flight-log data for healthy and defective propeller conditions.
- Decision-support problem: Unlike binary classification, the decision-support problem requires explaining the recommendation, identifying the dominant indicator, and assessing whether evidence justifies inspection.The introduction frames maintenance needs as routine monitoring, maintenance review, or mandatory inspection.
- Contribution: The main contribution integrates DronePropA feature extraction, metamorphic adequacy testing, and redundancy-adjusted AAS modelling into one interpretable decision-support prototype.The proposed separation distinguishes observed operational burden, the propeller-health burden score, and the structural adequacy of the scoring policy.
- AAS interpretation: Artificial Age Score is adapted as a structural policy-adequacy measure that evaluates scoring-policy consistency under health-monitoring transformations rather than chronological age.The framework uses metamorphic relations, consistency loss, redundancy adjustment, and logarithmic penalty to assess policy behavior.
- Framework: The framework computes six normalized indicators: trajectory tracking error, attitude instability, thrust-command burden, motor-command imbalance, ESC-command instability, and battery-level stress.The indicators are derived from selected historical real DronePropA flight logs and raw MATLAB time-series matrices.
2 Methodology … 2.8 Metamorphic Adequacy Relations
The study presents a retrospective Metamorphic Artificial Age Score decision-support prototype that converts real DronePropA flight logs into six baseline-normalized health indicators, candidate scores, and metamorphic adequacy assessments. A controlled healthy-to-defective subset supports internally consistent comparison while outputs remain post-flight decision-support evidence rather than certified diagnosis or autonomous control.
- 2.1 Research Design: The prototype uses historical real DronePropA flight logs to transform flight behaviour into interpretable maintenance-support outputs rather than a conventional fault classifier.The methodology combines dataset-based feature extraction, scoring policies, metamorphic adequacy relations, and AAS evaluation.
- 2.2 Selected DronePropA Subset: The controlled subset contains one healthy baseline and three defective propeller cases sharing the same speed profile and trajectory type.This design limits variation attributable to trajectory demands or speed conditions during healthy-to-defective comparison.
- 2.3 DronePropA File Structure and Signal Extraction: Numerical values are extracted directly from MATLAB .mat files using scipy.io.loadmat, with commander and QDrone matrices used and stabilizer data excluded from the feature set.MATLAB row indices are converted to Python zero-based indexing during computation.
- 2.4 Flight-Log-Derived Input Vector: Each selected file becomes a six-dimensional input vector covering trajectory tracking error, attitude instability, thrust-command burden, motor-command imbalance, ESC-command instability, and battery-level stress.Commander data supplies position and reference-thrust signals, while QDrone data supplies attitude-rate, battery, motor-command, and ESC-command signals.
- 2.5 Raw Feature Extraction: The raw indicators combine flight-log signal characteristics including position error, attitude-rate variation, thrust variation, motor asymmetry, ESC variation, and battery voltage stress.The resulting raw feature vector is z = (z1, z2, z3, z4, z5, z6).
- 2.6 Baseline Normalization: Indicators are normalized against the healthy file, retaining only positive deviations as burden evidence in a vector ranging from 0 to 1.A value of 0 indicates no positive burden relative to baseline, while 1 indicates the maximum observed positive burden in the selected retrospective subset.
- 2.7 Candidate Scoring Policies: Three candidate scoring policies evaluate normalized burden: a linear weighted score, an operationally capped score, and a threshold-sensitive score.The illustrative weights are demonstration parameters requiring recalibration through expert judgement, sensitivity analysis, and broader empirical validation.
- 2.8 Metamorphic Adequacy Relations: Six metamorphic relations test whether candidate scores respond appropriately to uniform improvement, tracking-error escalation, attitude instability, motor imbalance, and battery-related transformations.For each policy and relation, zero violation indicates satisfaction, whereas a positive violation indicates failure.
2.9 Redundancy-Adjusted Artificial Age Score for Policy Adequacy · 2.10 Maintenance Recommendation and Confidence Measures
The prototype adapts redundancy-adjusted Artificial Age Score (AAS) from metamorphic testing to assess policy adequacy for drone propeller health monitoring. It combines burden and critical-indicator evidence into maintenance recommendations while separating policy confidence from decision confidence.
- 2.9 Redundancy-Adjusted Artificial Age Score for Policy Adequacy: AAS measures structural or behavioural policy burden through consistency loss, logarithmic penalty, and redundancy-aware aggregation rather than chronological age.The study applies this logic to evaluate scoring-policy adequacy under metamorphic relations.
- 2.9 Redundancy-Adjusted Artificial Age Score for Policy Adequacy: The formulation converts relation-level violation magnitudes into bounded consistency scores, applies a logarithmic penalty kernel, and aggregates violations with redundancy correction.This transfers a previously introduced metamorphic-testing transformation to DronePropA-based propeller health monitoring.
- 2.9 Redundancy-Adjusted Artificial Age Score for Policy Adequacy: Redundancy adjustment is estimated from feature-set overlap among active violated relations, preventing overlapping violations from being treated as fully independent evidence.A relation with no violation receives zero redundancy adjustment.
- 2.9 Redundancy-Adjusted Artificial Age Score for Policy Adequacy: The preferred policy is the candidate with the lowest redundancy-adjusted AAS within the tested policy set, not a globally optimal scoring policy.This represents the lowest redundancy-adjusted metamorphic inconsistency under the defined adequacy relations.
- 2.10 Maintenance Recommendation and Confidence Measures: The selected policy combines a propeller-health burden score with critical normalized indicators to produce Routine monitoring, Maintenance review / supervisory monitoring recommended, or Mandatory inspection required.Recommendations are generated from the scoring policy and critical indicators.
- 2.10 Maintenance Recommendation and Confidence Measures: Trajectory tracking error, motor-command imbalance, and ESC-command instability are critical indicators for escalation.Mandatory inspection follows high aggregate burden or a high normalized level in a critical indicator; moderate burden supports maintenance review.
- 2.10 Maintenance Recommendation and Confidence Measures: Policy confidence reflects the AAS margin between the best and second-best policies, whereas decision confidence reflects the strength of indicators supporting the maintenance recommendation.A strong recommendation can coexist with weak policy confidence when candidate policies are structurally close but one normalized indicator dominates.
2.11 Computational Implementation
The Python implementation transforms selected DronePropA MATLAB logs into normalized flight-log indicators, evaluates candidate scoring policies and metamorphic adequacy, and produces redundancy-adjusted AAS-based maintenance recommendations. It outputs structured intermediate and decision tables for a retrospective decision-support prototype, not certified diagnostic, maintenance-control, or autonomous flight-control software.
- Computational workflow: The Python workflow extracts six raw flight-log indicators, normalizes them to the healthy baseline, evaluates three candidate policies, applies six metamorphic relations, and computes redundancy-adjusted AAS values.The workflow uses commander and QDrone data from selected DronePropA .mat files and produces maintenance-support recommendations.
- Computational outputs: The implementation produces five structured outputs: raw feature, normalized input, policy-level AAS, policy-ranking, and DSS decision-summary tables.These outputs were used to construct the Results-section tables.
- Prototype scope: The implementation is limited to a retrospective decision-support prototype and is not certified diagnostic, maintenance-control, or autonomous flight-control software.Its structured outputs support the maintenance recommendations reported in the paper.
- Data representation: Raw feature vectors are stored in table Z, while table U contains six baseline-normalized burden indicators constrained to [0, 1]^6.The healthy baseline file is denoted F0.
- Policy evaluation: Candidate policies are ranked by ascending policy AAS after relation-level violations, consistency scores, penalties, redundancy adjustments, and policy-level outputs are computed.Stored policy outputs include violation count, total violation, R, AASCj, and yj.
3 Results
The results show that defective propeller effects were distributed across different flight-log channels, with SV1 dominated by ESC-command instability, SV2 by motor-command and ESC-command burden, and SV3 by trajectory tracking error. The prototype converted these patterns into maintenance recommendations, while policy confidence remained weak but decision confidence remained strong.
- Feature extraction: Six flight-log indicators were extracted: trajectory tracking error, attitude instability, thrust-command burden, motor-command imbalance, ESC-command instability, and battery-level stress.These indicators were derived from selected DronePropA MATLAB flight logs.
- Raw and normalized patterns: Defective cases did not increase uniformly across channels: SV1 and SV2 increased motor-command imbalance and ESC-command instability, while SV3 increased trajectory tracking error.Trajectory tracking error and thrust-command burden decreased in SV1 and SV2 relative to the healthy baseline.
- Policy evaluation: The healthy baseline produced zero AAS for all candidate policies, while SV2 and SV3 produced equal AAS values across C1, C2, and C3.For SV1, threshold-sensitive policy C3 produced the lowest AAS; the resulting policy-selection margin was weak.
- Decision outputs: SV1 received maintenance review, whereas SV2 and SV3 received mandatory inspection because ESC-command instability, motor-command imbalance, and trajectory tracking error reached the cited dominant burdens.SV2 selected C1 with policy AAS 0.016544 and dominant violation MR4; SV3 selected C1 with policy AAS 0.008272 and dominant violation MR2.
- Confidence interpretation: Decision confidence remained strong despite weak policy confidence because maintenance recommendations were supported by clear normalized indicators.The healthy case had no positive burden indicators, while each defective case had a distinct dominant indicator pattern.
4 Discussion
The prototype converts selected flight logs into interpretable maintenance-support outputs by identifying dominant propeller-health burden channels and evaluating scoring-policy adequacy through metamorphic relations. Results show that fault effects vary across operational indicators, while separating policy confidence from decision confidence strengthens interpretation of recommendations.
- Multi-indicator findings: Fault effects were distributed across channels: SV1 was dominated by ESC-command instability, SV2 by motor-command imbalance and ESC-command instability, and SV3 by trajectory tracking error.Increasing labelled severity did not produce a uniform increase across all extracted indicators.
- Multi-indicator findings: A multi-indicator approach is necessary because trajectory tracking error or motor-command imbalance alone would miss the strongest burden patterns in some defective cases.An aggregate score alone could also obscure which operational channel caused the maintenance recommendation.
- AAS interpretation: AAS evaluates the structural adequacy and consistency of candidate scoring policies rather than chronological drone or propeller age.The propeller-health score summarizes indicator burden, whereas policy-level AAS evaluates scoring-mechanism behaviour under metamorphic relations.
- Metamorphic adequacy: SV2’s dominant metamorphic violation involved motor-imbalance-not-masked, while SV3’s involved tracking-error-escalation.These relation-level outputs identify which structural expectation was most relevant to each case.
- Confidence interpretation: Policy confidence remained weak because candidate AAS margins were small, whereas decision confidence was strong because recommendations were supported by clear normalized indicators.All candidate policies produced equal AAS values in SV2 and SV3, while SV1 showed elevated ESC-command instability and SV2 and SV3 showed maximum normalized burdens in their dominant indicators.
- Methodological implications: Baseline normalization made differently scaled flight-log indicators comparable in [0, 1], while redundancy values were zero because active violations did not substantially overlap.Redundancy becomes informative only when overlapping metamorphic violations are active.
- Prototype value: The prototype transformed selected flight logs into normalized indicators, identified dominant burden channels, evaluated candidate policies through metamorphic relations, and generated interpretable maintenance recommendations.The healthy baseline produced zero normalized burden, while defective cases received maintenance review or mandatory inspection outputs.
5 Limitations and Future Work
The prototype is limited by retrospective evaluation on a narrow dataset subset, baseline normalization, closely ranked scoring policies, non-exhaustive adequacy relations, and currently uninformative redundancy adjustment. Future work should broaden validation, calibration, relation design, model comparison, and maintenance-oriented deployment.
- Limitations: The evaluation used one healthy baseline and three defective propeller cases, so broader validation is required before operational use.The cases came from the same fault group, speed profile, and trajectory.
- Limitations: Normalization used one healthy baseline and treated only positive deviations as burden evidence, requiring broader multi-flight and multi-condition baseline modeling.Future baselines should include multiple healthy flights, drones, trajectories, and speed conditions.
- Limitations: Policy confidence remained weak because AAS margins were small, indicating that C1, C2, and C3 were structurally close under the defined relations.Future work should test wider policy sets, adaptive weighting, nonlinear escalation, and data-informed threshold calibration.
- Limitations: The six metamorphic adequacy relations provide a transparent starting point but are not exhaustive, motivating additional drone-health relations.Existing relations include improvement, tracking-error escalation, motor-imbalance sensitivity, ESC-instability sensitivity, and battery-alone non-critical behaviour.
- Limitations: The mean redundancy value was zero because active violation patterns did not substantially overlap, so larger evaluations are needed to reveal informative redundancy.Redundancy becomes informative only when overlapping metamorphic violations are active.
- Future Work: Future evaluation should cover the full DronePropA dataset and compare AAS-DSS outputs with conventional fault-detection or classification models.Deployment-oriented work should also assess fleet maintenance ranking, dominant indicators, near-real-time monitoring, safety review, operational thresholds, and domain-specific calibration.
6 Conclusion
The paper presents a Metamorphic Artificial Age Score decision-support prototype that converts flight logs into six normalized propeller-health indicators and evaluates them through candidate policies and metamorphic adequacy relations. A controlled retrospective evaluation shows differentiated maintenance decisions, supporting multi-indicator post-flight monitoring rather than reliance on a single burden signal.
- Framework: The prototype derives six normalized indicators covering trajectory, attitude, thrust, motor commands, ESC commands, and battery stress from selected DronePropA flight logs.The indicators are evaluated using candidate scoring policies and metamorphic adequacy relations.
- Evaluation: One healthy baseline and three defective cases were evaluated under the same fault group, speed profile, and trajectory.The healthy baseline produced a zero normalized burden vector and was assigned to routine monitoring.
- Evaluation: Severity 1 was dominated by ESC-command instability and assigned to maintenance review.The study used controlled retrospective cases from the DronePropA dataset.
- Interpretation: Fault effects can emerge through different operational channels rather than a single monotonically increasing indicator, motivating a multi-indicator decision-support structure.The framework reports the dominant indicator, selected scoring policy, policy-level AAS, dominant metamorphic violation, policy confidence, and decision confidence.
- Contribution: AAS functions as a structural adequacy measure of candidate scoring-policy consistency, not a chronological aging metric.The integrated prototype combines DronePropA feature extraction, metamorphic adequacy testing, and redundancy-adjusted AAS modelling for interpretable maintenance prioritization.
Data Availability Statement
The study uses the publicly available DronePropA dataset, specifically a controlled subset of selected MATLAB .mat flight-log files.
- Data Availability Statement: The study used selected MATLAB .mat flight-log files from the publicly available DronePropA dataset.The original dataset is available through Mendeley Data, with an accompanying Data in Brief article.
Funding
The research received no external funding.
- The research received no external funding.
Ethics Statement
The study uses a publicly available drone flight-log dataset and involves no human participants, animals, social media data, or personally identifiable information.
- Ethics Statement: The study uses a publicly available drone flight-log dataset and involves no human participants, animals, social media data, or personally identifiable information.
Code Availability Statement
The Python implementation for feature extraction, normalization, metamorphic testing, AAS calculation, and DSS output generation is available from the author upon reasonable request.
- Code Availability Statement: The Python implementation is available from the author upon reasonable request.It generated the reported raw feature, normalized input, policy-level AAS, policy-ranking, and DSS decision outputs.