Source-linked AI summary
Semantic Signal-Assisted Inspection and Recovery Allocation in Reverse Logistics
Jiani He, Dingyan Shang, Yihua Xu, Shiqi Huang, Yan Lyu, Jize Li, Shangjing Tang
TL;DR
Reverse-logistics operators must inspect and route returned assets before condition is fully observed, but full inspection consumes scarce labor. The paper proposes SSADS, which converts return notes into condition and signal-quality scores guiding inspection and recovery allocation under shared capacity. Across synthetic benchmarks, it improves value over noisy full inspection while reducing inspection cost, although risk-blind no-inspection performs better under the purely economic objective and matched-cost targeting matters mainly in aircraft maintenance.
Problem
Reverse-logistics operators must decide inspection and recovery actions under condition uncertainty, while full inspection consumes scarce labor and existing models generally expect structured inputs instead of return notes.
Method
SSADS maps each return note to a condition factor and signal-quality score, then uses them to set inspection depth and allocate feasible recovery actions under shared labor capacity.
Results
Across three synthetic scenarios, SSADS–Keyword improves net recovery value over noisy full inspection while reducing inspection cost, whereas no-signal/no-inspection has higher simulated TRV under the purely economic objective; matched-cost targeting adds $53.9K per batch in S2.
Takeaways & Limitations
Narrative evidence can support inspection allocation before recovery decisions, with material semantic-targeting value in the aircraft scenario but negligible value in the IT and consumer configurations.
Takeaways & Limitations
RLDB uses stylized prices, yields, capacities, inspection noise, and fixed return-note templates, so it supports reproducible mechanism analysis rather than estimates of field effectiveness.
Abstract
from arXiv · showhide
Reverse-logistics operators often decide how to inspect and route returned assets before their condition is fully observed, while full inspection consumes scarce labor. Semantic Signal-Assisted Decision Support converts return notes into a condition factor and a signal-quality score that guide inspection depth and recovery allocation under shared labor capacity. We evaluate the framework in three synthetic benchmark scenarios spanning information technology decommissioning, aircraft maintenance, and consumer-electronics returns. Across 30 paired simulation seeds, the keyword implementation improves net recovery value relative to a structured-feature comparator with noisy full inspection while reducing inspection cost in all three scenarios. A risk-blind comparator that skips inspection altogether still records higher value under the benchmark's purely economic objective. At matched inspection cost, score-guided targeting adds 53.9 thousand United States dollars per batch in the aircraft scenario but has little economic effect in the other two configurations; phrase and large language model extractors provide further gains in the aircraft scenario. These results show how narrative evidence can support inspection allocation before recovery decisions are made.
I. Introduction
Reverse-logistics decisions must manage uncertain asset condition while conserving scarce inspection capacity. SSADS uses return-note evidence to allocate inspection and recovery decisions across varied operational scenarios.
- Motivation: The reverse-logistics challenge is recovering value when asset condition is uncertain and full inspection consumes time, labor, and sometimes specialized equipment.Existing optimization models generally require structured numerical inputs rather than free-text annotations.
- Framework: SSADS maps each return note to a condition factor and signal-quality score that adjust expected yield and inspection depth.Inspection may be skipped, quick, or full, while the optimizer ranks feasible dispositions under shared capacity.
- Framework: The framework treats inspection as an asset-level allocation decision rather than a uniform preprocessing step.The recovery optimizer combines note-derived information with inspection outcomes to allocate assets or scrap under capacity constraints.
- Benchmark: SSADS is evaluated across IT decommissioning, aircraft MRO, and consumer-electronics returns using descriptive technician notes, maintenance records, and customer descriptions.The scenarios differ in regulatory constraints, note noise, unit value, and the central role of triage.
- Contributions: The paper contributes inspection targeting from return notes, a modular recovery pipeline, and reproducible benchmark evaluation with shared baselines and paired seeds.The extractor ladder compares keyword rules, phrase matching, and large language model extraction under common downstream decision logic.
- Contributions: The paper evaluates how extractor accuracy translates into inspection targeting and recovery value while validating the allocator against the corresponding fixed-action optimum.The benchmark uses the Reverse Logistics Decision Benchmark across the three scenarios.
II. Background and Related Work
Prior reverse-logistics optimization typically relies on structured numerical inputs, while newer text-based and maintenance systems address adjacent tasks. SSADS instead connects narrative evidence to inspection depth and recovery allocation under yield uncertainty.
- Reverse-logistics optimization commonly requires structured inputs such as disassembly bills of materials, cost tables, and pre-specified yield rates.
- Recent language-based supply-chain systems primarily target forward-chain optimization or structure extraction rather than reverse-logistics disposition under yield uncertainty.
- Maintenance studies connect narrative records to work-order assignment, causal extraction, decision support, or insight generation, whereas SSADS maps them to inspection depth and recovery allocation.
- SSADS uses fixed signal-quality thresholds to allocate inspection, applying value-of-information reasoning without explicitly estimating information value.
III. Framework Design
SSADS selects inspection levels before recovery actions and combines expected-margin evaluation with shared labor constraints. Its decision engine uses semantic condition information to rank feasible recovery or scrap options.
- SSADS assigns each asset an inspection level q_m ∈ {0, 1, 2} before selecting a feasible recovery action or scrap.The levels represent skip, quick, and full inspection.
- Component-recovery value depends on component counts, recovery prices, and expected yields, while whole-unit refurbishment uses configured asset-level value multiplied by 𝜙(m).
- Partial-recovery teardown recovers 0.60 of component value at lower cost and time, but the margin-ranked allocator chooses between component recovery and whole-unit refurbishment.
- The decision engine uses 𝜙 to set the yield prior, 𝜎 to gate inspection depth, then ranks positive expected margins by value per processing minute under remaining shared labor capacity.
C. Semantic Extraction Layer
The semantic layer converts return notes into a condition factor and signal-quality score that feed a common inspection and recovery policy. Multiple extractor types can therefore be compared through the same downstream decision logic.
- Any extractor producing 𝜙 ∈ (0, 1] and 𝜎 ∈ [0, 1] can use the same inspection policy and decision engine.
- The benchmark’s keyword-and-pattern classifier falls back to 𝜙 = 1.0 when no signal triggers, while phrase and cached LLM extractors provide comparison points under identical downstream logic.
- The pipeline maps 𝜎 to inspection depth and 𝜙 to expected yield, then allocates shared labor after selected inspections are charged.
- The deterministic phrase matcher selects the most frequent declared condition, assigns fixed 𝜙 and 𝜎 = 0.90, and returns 𝜙 = 1.0 and 𝜎 = 0.20 when no phrase matches.
- The prompted LLM extractor independently scores condition and condition-information content using scenario-specific instructions and temperature 0.
- The analytical pre-inspection yield supports linear allocation without recourse, while realized yields are simulated from a Beta distribution with fixed concentration 20.
E. Adaptive Inspection Policy
SSADS uses signal quality to choose inspection depth, then allocates recovery work by expected margin per labor minute under capacity constraints.
- Adaptive inspection policy: Signal-quality thresholds select skip, quick, or full inspection before recovery allocation.S1 and S2 use thresholds τh=0.5 and τl=0.25; S3 uses τl=0.45.
- Adaptive inspection policy: Inspections update condition estimates with noisy observations weighted according to inspection depth.Quick and full inspections use standard deviations of 0.15 and 0.05, with observation weights of 50% and 90%.
- Capacity-constrained allocation: The allocator ranks positive-margin recovery candidates by expected margin per incremental labor minute while capacity remains.Each candidate is admitted only when the shared labor-minute budget permits it.
- Evaluation scenarios: The evaluation spans IT decommissioning, aircraft MRO, and consumer-electronics returns.Table I defines the benchmark scenario set used to evaluate the policy.
- Batch workflow: The batch workflow extracts signal scores, selects inspection, updates condition estimates, computes margins, and assigns feasible dispositions.The process also charges inspection and scrap-handling costs before ranking candidates.
G. Illustrative Asset Walkthrough
The walkthrough shows how note-derived condition and signal quality jointly determine inspection and recovery priority for individual assets.
- Healthy server: A healthy server note produces φ=0.925 and σ=0.925, so the policy skips inspection and prioritizes component recovery.This decision saves the $75 and 60-minute inspection required by full inspection.
- Damaged server: A damaged PSU/CPU note produces φ=0.400 and σ=0.875, so high signal quality still skips inspection while lower condition reduces allocation priority.The example shows that high-confidence text can indicate poor condition without necessarily implying scrap.
A. Scenarios, Data, and Baselines
The benchmark generates noisy narrative observations, compares routing and inspection baselines across three scenarios, and evaluates recovery value over paired seeds.
- Data generation: Notes are generated from latent condition with omission probability p=0.15 and severity-perturbation probability p=0.25.The resulting text has an intentional but imperfect link to true yield.
- Baselines: Seven alternatives use the same generated assets and seed, including random, FIFO, structured noisy full-inspection, and oracle-based approaches.The baselines differ in routing, inspection, or access to latent yield information.
- Evaluation design: Results report means over 30 paired seeds with identical asset populations, inspection perturbations, and realized component outcomes within each seed.The reported intervals and significance tests describe variation across simulated seeds, not field generalization.
- Full-inspection comparison: SSADS–Keyword exceeds structured noisy full inspection by 50.9%, 16.0%, and 1.8% in S1, S2, and S3, respectively.It also saves $34.3K, $28.8K, and $1.0K in inspection cost across those scenarios.
- Economic-only comparison: The no-signal/no-inspection policy has higher simulated TRV than SSADS–Keyword by 0.5%, 4.6%, and 14.3% in S1, S2, and S3.This comparison uses the benchmark's economic-only objective.
C. Condition-Factor Extraction Accuracy
The extraction analysis distinguishes note informativeness from calibrated probability and evaluates condition-factor accuracy through correlations with latent yield.
- Accuracy metric: Table III reports Pearson r between each extractor’s condition factor φ and the latent yield factor.The analysis evaluates extraction accuracy rather than probability calibration.
- Signal-quality interpretation: Signal-quality score σ measures note informativeness, so expected calibration error and Brier score are not applicable.Operational selectivity is instead assessed with high-score bad-skip rates and σ-error association.
D. Effect of Extractor Choice on Recovery Value
Extractor choice materially changes aircraft recovery value, while phrase matching and DeepSeek improve on keyword extraction most clearly in S2. Overestimation remains the central risk when high scores cause low-yield assets to bypass inspection.
- Phrase extractor: $117.1K: SSADS–Phrase rises over SSADS–Keyword in S2 and exceeds the no-signal / no-inspection comparator by 2.6%.SSADS–Phrase reaches $614K in S1, $1,720K in S2, and $80.9K in S3.
- LLM extractor: SSADS–DeepSeek reaches $604K in S1, $1,703K in S2, and $81.0K in S3, with only S2 exceeding no-signal / no-inspection by 1.6%.The correlation estimate uses 150 records, whereas TRV values use all 15,000/15,000/30,000 records.
- Risk analysis: 10.1% of high-score skips in S2 under the keyword reader have true yields below 0.30, precluding autonomous keyword use in aircraft MRO.DeepSeek reduces the corresponding rate to 0% in S2 among 2,675 skipped assets.
- Risk analysis: Higher signal-quality scores are generally more selective, with score correlations to absolute condition error of −0.52/−0.65/−0.36 across S1/S2/S3.In S2, 35 of 421 low-yield skips have negative simulated net value, although aggregate net value remains positive.
- Sensitivity analysis: At half labor capacity, SSADS–Keyword TRV falls 0.6% in S1 and 49.9% in S2, while the prespecified S3 plan is infeasible under the reduced budget.Changing omission and severity perturbation changes TRV by at most 0.19/2.75/0.42% in S1/S2/S3.
- Sensitivity analysis: Removing the negative vocabulary inverts condition-factor correlation in S1 and S2, but TRV changes only −0.3% and −6.9%, respectively.The S2 bad-skip rate remains 10.1%, motivating a conservative fallback prior for unmatched notes.
G. Matched-Cost Targeting Ablation
Matched-cost targeting isolates the value of assigning identical inspection counts to different assets. The economic benefit is concentrated in aircraft MRO, where allocation and extractor choice have the greatest downstream effect.
- Targeting result: +$53.9K: score-guided targeting increases mean TRV in S2 at matched inspection cost, while gains are negligible in S1 and S3.Mean TRV changes by $0 in S1 and +$18 in S3; the S2 increase has p<0.001.
- Targeting design: Matched-cost random assigns the same skip/quick/full counts as SSADS but shuffles those actions across assets.This comparison tests targeting separately from inspection-policy differences.
- Allocation ablation: Removing margin-per-minute ranking reduces TRV by 0.0/38.0/5.6% for S1/S2/S3, reinforcing that allocation is most consequential in aviation.The inspection policy and action set remain fixed in this ablation.
- Deployment: The framework can begin with inspection prioritization while retaining required checks and logging suggested depth, observed condition, and outcomes for local threshold estimation.In S2, the workflow prioritizes uncertain records and reserves labor for high-margin parts while retaining qualified inspection and authorized sign-off.
- Extractor choice: Extractor choice depends on downstream recovery performance as well as privacy, auditability, latency, and model-version stability.In S2, SSADS–DeepSeek reaches $1,703K, below SSADS–Phrase at $1,720K, while the common interface leaves the decision model unchanged.
B. Limitations
The evaluation is deliberately synthetic and economically narrow, so its results isolate the decision mechanism rather than establish field effectiveness or safety. Deployment therefore requires broader validation, mandatory safety controls, and human approval.
- Synthetic scope: RLDB uses stylized prices, yields, capacities, inspection noise, and fixed return-note templates to isolate the decision mechanism.The design supports reproducible analysis rather than an estimate of field effectiveness.
- Economic objective: TRV omits certification-error costs, latent safety failures, warranty exposure, and most downstream failures, potentially favoring risk-blind no-inspection policies.Unsafe S2 skips therefore cannot be interpreted as financially optimal under the benchmark objective.
- Reproducibility: The companion repository releases configurations, noisy-note generators, seven comparators, exact prompts, score caches, paired seeds, and reproduction scripts.The released reproduction uses Python 3.12.13 with pinned dependencies and requires no GPU.
- Supported conclusion: The keyword implementation improves net recovery value over noisy full inspection while reducing inspection cost, but matched-cost targeting is materially valuable only in aircraft.Phrase and DeepSeek extractors further improve the aircraft result.
- Deployment boundary: Prospective advisory deployment should evaluate operational notes while retaining mandatory safety inspections and human approval.The conclusion limits the proposed use to an advisory setting rather than autonomous operation.