Source-linked AI summary

Local Gains and Fixed-Assignment Set Losses in Shared Set Decoders

Ze Zhang, Yang Zhang

arXiv:2608.14717v1cs.CVcs.LG

TL;DR

Shared set decoders can improve an edited slot while worsening the jointly decoded prediction set. Using a matched active-control deletion across 710 selected units in two checkpoints, the paper finds local-positive but fixed-assignment-negative effects, with persistence differing by readout and checkpoint.

  • Problem

    The paper examines how a local query-relation edit propagates through a jointly decoded prediction set when assignment and selection determine final utility.

  • Method

    The study compares selected target deletions with matched active-control deletions across 710 selection-conditional paired units in two checkpoints and multiple readouts.

  • Results

    In both checkpoints, target deletion improves the local slot but reduces fixed-assignment set utility; rematched and native readouts cross zero for DETR but remain negative for DINO.

  • Takeaways & Limitations

    Selected relations show deletion sensitivity whose observed consequence depends on the readout and intervention operator.

  • Takeaways & Limitations

    The evidence is conditional on selected populations and a different-recipient composite control, so it does not establish arbitrary-relation prevalence, architecture-family effects, or detector-level degradation.

Abstract

from arXiv · show

A query-relation deletion can improve the edited slot while reducing the utility of the prediction set that contains it. We study this tension in two related ResNet-50 DETR-family checkpoints using recorded, selection-conditional evidence from 710 paired image-relation units per checkpoint. The primary comparison subtracts a matched active control, which deletes the same leader source at a different recorded recipient, from the selected target deletion. It is therefore a composite contrast rather than a same-recipient placebo. The target-minus-control contrast is locally positive and fixed-assignment negative in both checkpoints. The opposite-sign pattern occurs within 302/710 DETR units and 460/710 DINO units. After rematching, the corresponding counts are 285/710 and 433/710. Rematching and native selection absorb enough of the mean loss for DETR intervals to cross zero, whereas DINO intervals remain negative, so persistence across readouts differs by checkpoint. A fixed-map comparison between hard deletion and a mass-preserving edit also differs before rematching. That comparison is conditional on the outcome-blind map and does not establish same-dose transport. Local intervention success therefore does not determine the consequence for a jointly decoded set. The supported conclusion is selection-conditional deletion sensitivity whose persistence depends on the readout and intervention operator. We do not identify an intervention-invariant edge mechanism, detector-level degradation, population prevalence, or the value of a training-time regularizer.

1 Introduction

In shared set decoders, a query-relation deletion can improve the edited slot while reducing fixed prediction-set utility. The response depends on the comparator, intervention operator, and readout stage, with checkpoint-specific persistence after rematching and native selection.

  • Problem: A positive local response can coexist with lower fixed-assignment utility because shared queries are jointly decoded and final outputs depend on assignment and selection.The edited slot does not determine the utility of the jointly decoded prediction set in isolation.
  • Readouts: The same edit can produce different findings at the target slot, under fixed assignment, after Hungarian rematching, or after native output selection.Reassignment may recover utility without undoing the underlying state change, while native selection may reveal another response.
  • Interpretation: These results establish a local–set sign reversal and checkpoint-specific recovery, not an intervention-invariant mechanism, detector-level degradation, population prevalence, or training-time regularizer value.Hard deletion, attenuation, and mass-preserving replacement are distinct operators unless their mapping is established, and fixed-assignment loss is not detector average precision.
  • Design: The study uses a fixed target-minus-matched-active-control contrast and follows the response from local to fixed, rematched, and native readouts.It also compares hard deletion with a mass-preserving replacement under a separate fixed map.
  • Results: Across two checkpoints, the selected deletion contrast improves the local target slot but lowers fixed-assignment set utility, with persistence differing after rematching and native selection.DETR intervals cross zero after later readouts, whereas DINO intervals remain negative; the mapped hard-minus-mass difference appears before rematching.

2 Related Work

Prior work shows that intervention findings depend on both the intervention and evaluation metric, especially when jointly decoded set predictions make assignment part of measurement. This study applies that discipline to selected directed-relation deletions across assignment-sensitive readouts at fixed checkpoints, rather than introducing a new matching procedure.

  • Intervention Evaluation: Intervention and evaluation metric are part of the measured claim, not neutral implementation details.This follows lessons from attention-faithfulness and activation-intervention studies questioning whether successful corruptions or patches identify a component’s semantic role.
  • Object Detection: Object-detection research has studied query interaction, denoising, grouping, relation bias, routing, and matching stability as architecture or training problems.These studies motivate the importance of interactions among queries.
  • Study Positioning: This study instead asks how a selected directed-relation intervention appears across assignment-sensitive readouts at fixed checkpoints and what that response licenses scientifically.The comparison uses a selected deletion and an active control rather than treating the intervention response as architecture- or training-level evidence.
  • Assignment-Sensitive Measurement: Set prediction makes assignment part of measurement: original assignment preserves one correspondence, rematching permits reorganization, and native selection evaluates the final output rule.The contribution is an empirical account that keeps these stages separate while comparing a selected deletion with an active control and an indexed alternative.

3 Study Design and Estimands

The study analyzes 710 baseline-selection-conditional image–relation units per checkpoint, comparing a target deletion with a matched active control and tracking effects from the selected slot through fixed-assignment, rematched, and native set utilities. It separately evaluates hard deletion and mass-preserving edits, with checkpoint-specific simultaneous inference and sign-event definitions.

  • Study population: 710 complete paired units per checkpoint remain after feasibility screening from 1,500 discovery images and 752 jointly pair-eligible images.The units are selected by a baseline-only procedure and are not random query edges.
  • Intervention contrast: The target arm deletes qc ← qℓ, while the active-control arm deletes qh ← qℓ at a separately matched recipient chosen using baseline quantities.The two checkpoints use the same discovery image identities, but selected query relations are model-specific.
  • Intervention operators: Hard deletion sets the selected layer-3 pre-softmax attention logit to −∞ and redistributes its attention mass, whereas mass-preserving editing exchanges target and donor weights after softmax.The mass-preserving operator preserves the two-cell weight sum but not value content or later decoder state.
  • Estimands: Fixed-assignment, rematched, and native set effects are U2 − U0, U3 − U0, and U4 − U0, respectively, while the local endpoint is the selected recipient slot.The set endpoint is the maximum quality among queries assigned to the same target object, so these endpoints are not interchangeable outcomes.
  • Inference and decision rules: 10,000 deterministic hash-seeded paired-image bootstrap resamples support within-model max-studentized simultaneous 95% intervals, while sign events require local response > 10−4 and set response < −10−4.Event rates use two-sided Wilson 95% intervals, and all mean effects describe only the selected populations.

4 Results

Across selected units, deletion improves the target slot while reducing fixed-assignment set utility, with the persistence of that loss depending on checkpoint and readout. The mapped operator comparison and legacy analyses remain conditional and do not establish transport, mechanism, detector degradation, or regularizer value.

  • Primary deletion contrast: DETR’s local effect was +0.095621 [ +0.070160, +0.121082 ], versus a fixed-assignment effect of −0.005741 [−0.010889, −0.000592].DINO showed the same pattern: local +0.036633 [+0.030623, +0.042643] and fixed-assignment −0.008833 [−0.013130, −0.004536].
  • Interpretation: The local-positive/fixed-negative sign reversal shows that recipient-slot improvement does not summarize the consequence for the jointly decoded set.The negative result concerns this set-level readout, not detector average precision or training-time relation suppression.
  • Within-unit reversals: The opposite-sign pattern occurred within 302/710 DETR units and 460/710 DINO units, remaining after rematching in 285/710 and 433/710 units.These rates use the declared deadzone and are selection-conditional, not estimates of arbitrary-relation prevalence.
  • Readout persistence: DETR’s rematched and native means were −0.001575 and −0.001686, with intervals crossing zero, while DINO’s −0.007424 and −0.007429 remained negative.Rematching and native selection therefore absorbed enough of DETR’s mean loss to include zero, but not enough for DINO.
  • Mapped operator comparison: Under the fixed outcome-blind map, hard-minus-mass means were DETR +0.081169 local, −0.005091 spillover, and −0.004981 fixed-assignment; rematched and native intervals crossed zero.In DINO, the corresponding means were +0.034817, −0.007747, −0.007701, −0.006561, and −0.006571, all with intervals excluding zero.
  • Scope and limitations: The mapped comparison is conditional on a fixed outcome-blind map, is not same-dose, and does not establish operator transport, a unique mechanism, calibrated non-transport, or regularizer efficacy.The legacy analysis also did not meet its full conjunction, while the fixed-checkpoint design trained or evaluated no regularizer.

5 Discussion

The discussion shows that local gains can conflict with fixed-assignment set losses, while rematching changes the observed response differently across checkpoints. It also limits interpretation: operator pathways, transport, and training-time regularizer effects remain unresolved.

  • Readout disagreement: Local improvement at a selected recipient can coexist with a loss for the jointly decoded set under fixed correspondence.The opposite signs indicate that shared computation can redistribute responses through other slots and later states.
  • Readout disagreement: Rematching is a recovery stage that changes the visible intervention response rather than a neutral relabeling.DETR’s later-readout intervals cross zero, whereas DINO’s remain negative.
  • Intervention interpretation: The hard-versus-mass-preserving fixed-map difference appears before rematching but does not identify a unique causal pathway.Compatible explanations include renormalization, donor content, later interaction, destructive deletion, or another mechanism.
  • Intervention interpretation: Transport remains unresolved because nominal dose cannot be connected to realized dose, and the fresh assay differs in prerequisites, population, and baseline.The evidence neither establishes nor rules out transport in the tested settings.
  • Implications and limits: A fixed-checkpoint sensitivity can motivate a regularizer hypothesis but cannot evaluate training-time regularizer performance.Training would alter the learned model, optimization trajectory, and query-interaction distribution rather than reproduce a fixed-checkpoint deletion.
  • Implications and limits: Intervention reports should jointly specify computational location, comparator, selected population, realized-dose contract, and readout stage.Fixed-checkpoint diagnosis and training-time design require separate authorization and evaluation.

6 Limitations

The study’s findings are conditional on selected populations and a composite control, and they do not establish detector-level, architecture-wide, or population-wide conclusions. Additional limitations concern unmeasured dose, fixed mapping, changed assay context, mechanism, operator transport, and training regularizers.

  • Results are conditional on selected image–relation populations from two related checkpoints and do not estimate arbitrary-relation prevalence or an architecture-family effect.
  • The active control edits a different recipient, making the primary comparison composite; set utilities are not detector average precision and do not establish detector-level degradation.
  • Legacy outcomes lack co-recorded per-head realized dose and immediate message displacement, while the mapped hard-minus-mass comparison relies on a fixed outcome-blind map without propagated mapping uncertainty.
  • The fresh assay changes context and fails its hard-sensitivity prerequisite, so the evidence does not identify a unique mechanism, establish calibrated operator transport or population-wide non-transport, or evaluate a learned regularizer.

7 Conclusion

Across two related shared set-decoder checkpoints, deleting a selected query relation improves the local target slot but reduces fixed-assignment set utility relative to a matched active control. Within-unit reversals and checkpoint-specific persistence after rematching show that this disagreement is not merely an aggregate effect.

  • Deleting a selected query relation improves the local target slot while reducing fixed-assignment set utility relative to a matched active control.
  • Within-unit reversals show that the local-versus-set disagreement is not only an aggregate effect.
  • Rematched and native readouts reveal checkpoint-specific persistence of the disagreement across the two related checkpoints.
  • A conditional mapped hard- minus-mass difference appears before rematching, but available dose and fresh-context evidence do not establish its broader interpretation.

8 Data and Code Availability

Source code and an analysis-ready reproduction package are publicly available, including materials needed to reproduce the reported aggregate analyses. Image pixels and model weights are not redistributed.

  • 8 Data and Code Availability: The public repository provides source code and an analysis-ready reproduction package for reproducing the reported aggregate analyses.It includes per-image rows, analysis configurations, population and pair manifests, integrity records, expected aggregates, and CPU-only code.
  • 8 Data and Code Availability: The same frozen reproduction package is included as arXiv ancillary material.
  • 8 Data and Code Availability: Image pixels and model weights are not redistributed and must be obtained separately.

9 Broader Impact · A Supporting Analyses and Provenance

The section emphasizes precise interpretation of internal interventions to avoid overstating positive ablations or misreading fixed-assignment utility changes as detector-level degradation. It also records the study’s funding and competing-interest disclosures.

  • 9 Broader Impact: Precise interpretation of internal interventions can reduce overclaiming from positive ablations.
  • 9 Broader Impact: The main interpretive risk is reporting fixed-assignment utility changes as detector-level degradation.
  • 9 Broader Impact: Keeping the comparator, selected population, and readout stage attached to each result reduces interpretive errors.
  • 9 Broader Impact: The research received no specific grant from any public, commercial, or not-for-profit funding agency.
  • 9 Broader Impact: Wuhan United Imaging Surgical Co., Ltd. had no role in study design, data analysis, interpretation, or the submission decision.
  • 9 Broader Impact: The authors declare no competing interests.

A.1 Operator and context checks · A.2 Native endpoint provenance note · B Operator and Population Details

Operator checks do not establish calibrated transport or support a regularizer study, while native endpoint provenance requires using T0-defined estimands. The interventions operate on decoder attention weights under selection-conditional population contexts.

  • A.1 Operator and context checks: 1,420 intervention records pass identity, repeated-baseline, and zero-dose controls exactly, but no intermediate mass-preserving condition passes the prespecified joint criteria.The records lack per-head realized dose and immediate message displacement, so calibrated operator transport is not established.
  • A.1 Operator and context checks: The COCO assay records realized dose on 128 paired units per checkpoint, but the required hard-deletion sensitivity check fails in both checkpoints.Its pairwise operator intervals fall inside assay-specific bands.
  • A.1 Operator and context checks: None of the complete prespecified regularizer candidates passes on 256 images per checkpoint, so no training study is undertaken.This is a design-stopping result, not evidence about regularizer performance.
  • A.2 Native endpoint provenance note: H4-D and T0 are historical provenance aliases, not canonical endpoints: H4-D denotes the decomposition-specific record and T0 the common-estimand audit.The native readout reported in H4-D is a derived decomposition-specific summary that differs from the T0 native endpoint.
  • A.2 Native endpoint provenance note: Primary quantitative claims use T0-defined estimands; numerical values from H4-D and T0 are not combined, compared, or substituted.The endpoints differ in estimand construction and analytical purpose.
  • B Operator and Population Details: The hard-deletion hook sets the selected target logit to −∞ in the decoder layer-3 pre-softmax attention tensor before softmax.Because the denominator changes, deleted mass can redistribute over remaining sources.
  • B Operator and Population Details: The mass-preserving arm edits post-softmax target/donor weights at the same layer while preserving their two-cell mass.Target and active-control arms use different recipients.
  • B Operator and Population Details: The denominator is selection-conditional because discovery and supporting assays differ in population, dataset, ground-truth formation, ranking, and admission contexts.Both use the same target-pair rule; the intervention-arm mapping is summarized in Table 2.

C Readout and Reproducibility Contract

The section distinguishes fixed-assignment, rematched, and native readouts while documenting the reconstruction contract. Operator checks fail the declared conjunction or hard-deletion sensitivity requirement without establishing population non-transport or regularizer performance.

  • Readout and reproducibility contract: Fixed-assignment, rematched, and native readouts are distinct endpoints, and native summaries are not combined or substituted.The supplementary manifest records hashes, receipt schemas, statistical configurations, and implementation files for aggregate reconstruction.
  • Operator and context checks: Figure 4 summarizes the legacy dose conjunction and the fresh hard-sensitivity check.No intermediate dose passes the complete frozen conjunction.
  • Operator and context checks: The legacy mass-preserving conditions fail different conjunction components, while the fresh assay fails its required hard-deletion sensitivity check.These results do not establish population non-transport or regularizer performance.
Loading 2608.14717v1…