Source-linked AI summary
Whose record is this? Diagnosing and authorizing record use in personalized multimodal models
Xinyu Mao, Junsi Li, Chenyang Liu, Haoji Zhang, Ming Sun
TL;DR
Visual personalization can retrieve a true record yet bind it to the wrong subject. The paper formalizes authorization as presence, edge validity, and answer support; constructs a matched diagnostic suite; and finds that typed authorization sharply reduces unsafe release, with support verification dominating the gain. The conclusions apply to the evaluated contracts rather than natural prevalence, consent, or visual identity.
Problem
Personalized multimodal models may use a relevant record for the wrong visual subject, while existing systems primarily optimize relevance or answer quality.
Method
The paper defines record authorization as P ∧E ∧S and constructs RecordAuth-Diag, a 3,690-case matched suite that changes one image–record edge while holding supplied content fixed.
Results
Typed authorization reduces unauthorized exposure from 43.63% to 3.06% on RecordAuth-Diag, while DAVIS unsafe release falls from 6.79% to 0.89% at comparable release rates.
Takeaways & Limitations
The observed improvement is dominated by support verification within an E ∧S decision, while appearance-based edge evidence remains conditional on presence.
Takeaways & Limitations
The study establishes inducibility under frozen contracts rather than natural prevalence or realized harm, and its relevance replay can produce an optimistic lower bound.
Abstract
from arXiv · showhide
Contextualized visual personalization can retrieve a true record yet apply it to the wrong visual subject. We formalize when a record may condition an answer as \emph{record authorization}: subject presence ($P$), record-edge validity ($E$), and answer support ($S$) must all hold. We call violations visual memory misbinding (VMM). We construct RecordAuth-Diag, a 3,690-case matched diagnostic suite that changes one image--record edge while holding the query, question, record text, and image multiset fixed. Card removal and nonce relabeling attribute these failures to supplied records. Raw-bank failures span Qwen-, Phi-, and Gemma-family interfaces: Gemma-3-4B-IT reaches 63.69\% local unauthorized use at 25.75\% clean recall. CoViP remains at 26.02\%, versus 22.49\% for its Qwen backbone at similar clean recall. Typed pre-generation authorization reduces Qwen card exposure on RecordAuth-Diag from 43.63\% to 3.06\%, while positive recall changes from 86.26\% to 60.90\%. Full $P\wedge E\wedge S$ validation uses 560 localized DAVIS cases: top-1 relevance and typed authorization have comparable release (28.93\% and 28.39\%) but 6.79\% and 0.89\% unsafe release, respectively. Of the 33 additional unsafe cases removed, 27 are support, 4 edge, 2 clean, and 0 boundary cases. Thus the observed increment is an $E\wedge S$ decision dominated by support, not an edge check alone. Appearance supplies $E$ evidence only conditional on $P$; authenticated subject tokens instantiate the missing presence witness as a sufficiency control. The claims concern the evaluated contracts, not natural prevalence, consent, or visual identity
1 INTRODUCTION
The paper frames visual personalization as an authorization problem: a relevant record is safe to use only when subject presence, record-edge validity, and answer support all hold. RecordAuth-Diag isolates edge-binding failures and evaluates typed authorization as a repair.
- Figure 1 changes only one image–record edge while fixing the query, question, record text, and image multiset, isolating binding errors.
- Record authorization requires three predicates: queried-subject presence (P), valid record attachment (E), and answer support (S).
- The study first measures unauthorized use after edge corruption, then tests which predicates systems can decide before generation, and finally examines the missing observation when the subject is absent.
- 43.63% to 3.06%: card-level edge authorization cuts unauthorized exposure on RecordAuth-Diag.On 560 DAVIS cases, release remains comparable at 28.93% for top-1 relevance and 28.39% for typed authorization, while unsafe release is 6.79% versus 0.89%.
- 27 of 33 additional unsafe cases removed by typed authorization are support failures, compared with 4 edge, 2 clean, and 0 boundary cases.The reported increment is therefore an E ∧S decision dominated by support rather than an edge check alone.
- The paper contributes the P ∧E ∧S contract, typed pre-generation authorization, and RecordAuth-Diag’s 3,690 matched cases.The suite changes one image–record edge while fixing the query, question, record text, and image multiset.
2 RELATED WORK
Prior personalized-vision work primarily measures whether models can use personal concepts or context, whereas this paper asks whether a particular retrieved record is authorized for the current subject. It connects that question to provenance, misbinding, and selective multimodal grounding while distinguishing record ownership from source authority.
- Existing contextualized visual personalization systems evaluate recognition and use of user-specific concepts or accumulated context, not whether a retrieved record is authorized for the current subject.MMPB scales capability evaluation to 10,000 image–query pairs and 111 concepts, while CoViP formalizes contextualized visual personalization and personalized captioning.
- RecordAuth-Diag changes which visual subject owns a record while keeping the query, question, record text, and image multiset fixed.This extends capability benchmarks along an authorization axis rather than measuring capability alone.
- Record ownership differs from source authority: the former asks whether information applies to a subject, while the latter asks whether it came from a trusted source.Record ownership additionally requires visual subject assignment.
- Related misbinding work studies evidence attached to the wrong speaker, time, modality, or generated subject, whereas this paper studies ownership of an external record across subjects.
- Selective prediction and visual grounding evaluate abstention, risk–coverage, or claim support, while record authorization makes a pre-generation decision for each supplied record.
3 RECORD AUTHORIZATION: DECISION AND IDENTIFIABILITY
The paper defines release as a typed conjunction of presence, edge validity, and answer support, with NULL when no unique supported card can be authorized. Its analysis separates what the method decides from the external observation needed to establish presence.
- For a candidate answer citing card m_jk, P indicates subject presence, E indicates group attachment, and S indicates support for the requested answer field.
- If no unique card can be cited or any predicate fails, the operational action is NULL and no personal card reaches the generator.Target absence changes P, reassignment changes E, and stale, missing, or unsupported fields change S.
- The method decides edge validity and answer support, while presence remains an external observation in the visual experiments.Ground-truth masks and authenticated tokens are supplied to analysis rather than learned as presence estimates.
- Behavioral VMM measures non-abstaining answers matching answer-bearing cards outside the authorized set, while card removal and counterfactual relabeling test supplied-record attribution.Clean accuracy, positive recall, coverage, and NULL specificity prevent always-NULL behavior from appearing safe.
- Relative edge evidence is conditional on subject presence and visual separability, and the stated total-variation inequality is an identifiability tool rather than a numerical certificate for evaluated policies.
- Authenticated subject-equivalent tokens can witness P, E, and S jointly when all tokens match and support holds, without requiring appearance-based inference.The result is a sufficiency control under the explicit token contract.
4 TYPED RECORD AUTHORIZATION
Typed record authorization traces a cited record and gates release with separate P, E, and S witnesses, using appearance and metadata only conditionally on presence. The evaluated implementation supports pre-generation or post-generation authorization and returns NULL for failed or ambiguous traces.
- Typed authorization can act before generation by serializing only records whose edge and requested field pass, or after generation by authorizing the cited record.Both placements require a traced record, valid edge E, and answer support S.
- An answer-conditioned trace selects one cited record, while separate P/E/S witnesses gate release; failed or ambiguous traces return NULL.Appearance-based E remains conditional on P, and authenticated provenance supplies the missing presence credential.
- Nonce-bearing values map answers to event IDs, and support requires an active, image-backed, latest usable record containing the requested field.Duplicate values, aliases, and non-unique paraphrases return ⊥.
- The edge margin Δ_g compares coherence under the claimed group with every counterfactual assignment and is valid only under subject presence and visual separability.
- The dual-path implementation freezes guard and raw-rescue thresholds, releasing supported guard candidates first and supported raw candidates only under the stricter threshold.The offline matched control applies the frozen group/card allow-list to both candidates without a new authorization call.
- ATRA uses separate full-frame, localized, and authenticated-ID observations, with learned edge/group features intersected with a deterministic anchor–triad rule.
5 EXPERIMENTAL DESIGN
The experimental design combines matched diagnosis, localized validation, typed authorization, and controlled confirmation to evaluate record authorization and its safety–utility trade-offs. It also reports computational cost and limits confirmation generalization through a track-disjoint but not video-disjoint split.
- Diagnostic suite: 3,690 matched cases across 87 identities and five identity-clustered folds evaluate multiple multimodal interfaces, with Idefics3 excluded by a preregistered capability gate.The suite evaluates Qwen3-VL-8B, CoViP, Phi-3.5-Vision, RAP, and Gemma-3-4B-IT; SigLIP2 and DINOv2 remain frozen.
- Localized validation: Full P ∧E ∧S validation uses 156 masked DAVIS tracks, split into 100 development tracks and 56 frozen confirmation tracks.RecordAuth-Diag diagnoses large-scale behavior and card authorization but lacks instance masks, while DAVIS supplies localized query and record evidence.
- Diagnostic controls: The matched HGB uses 24 candidate features and leave-two-fold-out training, but remains diagnostic rather than confirmatory because it was specified after DAVIS-56 inspection.Guard/raw thresholds are 0.832/0.984 and optimize DAVIS-100 clean accuracy subject to aggregate and condition-wise risk constraints.
- Metrics: The evaluation reports unauthorized memory-answer rate alongside clean accuracy, positive recall, coverage, NULL specificity, fixed-false-release frontiers, and separate aggregate and worst-condition risk.Unless otherwise stated, rates use all 560 cases; clean uses 56, and worst-condition risk is the largest ten-condition rate.
- Compute contract: Typed authorization uses 560 MLLM calls and 0.285M input tokens, making it 7.1× cheaper in input tokens than the full-bank arm at matched calls.The full-bank arm serializes every record, using 2.032M input tokens and 539.0 seconds versus 300.6 seconds for the one-path guard.
- Limitations: The 56 confirmation tracks are track-disjoint but not video-disjoint, and a post-hoc 19-track video-disjoint sensitivity analysis estimates a 10.53-point clean gain.The sensitivity interval is [0.00, 26.32], with one unsafe case for each policy.
6 RESULTS
The results diagnose supplied-record use, test typed card-level authorization, compare it with relevance pruning, and identify the boundary imposed by missing subject presence.
- 6.1 DOES VMM USE A SUPPLIED RECORD?: All five raw interfaces show substantial unauthorized use under local edge corruptions, with whole-group swaps treated separately because ownership is not identifiable without invariant side information.The reported conditional rate is measured after a correct clean answer.
- 6.1 DOES VMM USE A SUPPLIED RECORD?: Card removal preserves no exact aligned unsafe answer, while nonce relabeling exposes copied record content, supporting supplied-record dependence for measured failures.Nonce copying ranges from 38.10% to 92.78%.
- 6.2 TYPED RECORD AUTHORIZATION REMOVES MOST UNAUTHORIZED RELEASE: 43.63/44.09% to 3.06/4.61%: typed edge authorization lowers Qwen/CoViP card exposure, while positive recall falls from 86.26/86.91% to 60.90/66.63%.Group routing and card authorization are separate decisions; parse failures return NULL.
- 6.2 TYPED RECORD AUTHORIZATION REMOVES MOST UNAUTHORIZED RELEASE: 0 unsafe releases occur in the decidable class and five at the boundary under typed authorization.The boundary includes cases where the queried subject is absent or the whole-bank swap is unidentifiable; the reported whole-bank reduction relies on unpermuted anchors.
- 6.2 TYPED RECORD AUTHORIZATION REMOVES MOST UNAUTHORIZED RELEASE: 149 unsafe releases (26.61%) fall to 5 (0.89%) across 560 cases with typed authorization, while correct releases rise from 129 to 154.Input tokens also fall from 2.032M to 0.285M.
- 6.2 TYPED RECORD AUTHORIZATION REMOVES MOST UNAUTHORIZED RELEASE: 28.93% vs. 28.39% release: top-1 relevance and typed authorization are comparable, but unsafe release is 6.79% versus 0.89%.Of 33 additional unsafe cases removed beyond top-1 relevance, 27 are support, 4 edge validity, 2 clean, and 0 boundary cases.
- 6.4 WHAT OBSERVATION IS MISSING?: Authenticated tokens remove the target-absence residue and reach the candidate-union ceiling with 0 of 560 observed false releases.The tokens supply the missing presence witness; supplied evidence decides edge validity and support, whereas presence requires provenance.
- 6.4 WHAT OBSERVATION IS MISSING?: Target-absence false answers fall to 1.61% on free-form questions, while attribute-absence false answers remain 46.30–53.70%.The unvalidated judge indicates a predicate shift rather than accuracy, and MyVLM remains a boundary audit.
7 LIMITATIONS AND DISCUSSION
The study establishes inducibility under frozen evaluation contracts rather than natural prevalence or realized harm, and its replay design imposes several scope constraints.
- The study establishes inducibility under frozen contracts, not prevalence or realized harm.
8 CONCLUSION
The paper frames unauthorized visual-memory use as a record-authorization problem, diagnoses it with matched interventions, and evaluates typed authorization and observation boundaries. Results show substantial reductions in unauthorized release, while authenticated provenance remains necessary for presence and several evaluation limits remain.
- RecordAuth-Diag isolates authorization from recall and observes raw-bank misbinding across Qwen-, Phi-, and Gemma-family interfaces.
- Authenticated tokens make typed false acceptance zero under the stated sufficiency assumptions, but DAVIS track identities are supplied evaluation metadata rather than visually inferred credentials.
- The evaluation uses a track-disjoint but not video-disjoint split, so the confirmation set does not establish video-disjoint generalization.
- Both raw generators fall below the registered 61% capability floor, limiting interpretation of the counterfactual-policy comparison.
I TRAINING AND SELECTION PROTOCOL
The protocol fixes identity folds, validation-only threshold selection, and deterministic training procedures before evaluating relational authorization across held-out conditions and encoders.
- Training and calibration: Five identity folds and seeds 1, 7, 21, 42, 89 use AdamW with validation-based calibration and early stopping.Training uses learning rate 10^-3, weight decay 10^-4, batch size 64, patience 20, and at most 200 epochs.
- Training and calibration: ATRA thresholds use seven empirical validation quantiles plus fixed boundary values for group probability, group margin, and observed-edge probability.Flat-LogReg uses five quantiles with the same boundary conventions and an additional NULL-probability grid.
- Evaluation design: The final ensemble aggregates each seed across five held-out folds rather than selecting a best seed.The ensemble improves over the seed mean on both encoders, while best-seed results are excluded from headline tables.
- Held-out generalization: On held-out mechanisms, ATRA transfers most clearly in safety, while accuracy changes vary by encoder.With SigLIP2, ATRA changes accuracy/unsafe exposure from 79.77/9.76% to 74.07/8.67%; with DINOv2, from 66.49/25.38% to 60.61/12.60%.
- Attribution checks: The diagnostic controls attribute record use through card removal, nonce relabeling, and fixed visual retrieval keys.Nonce and shuffle alter record text while holding visual keys and retrieval scores fixed; removal exposes no card.
Q STRUCTURED SAME-MLLM VERIFIERS AND REFERENCE ROBUSTNESS
Structured same-MLLM and independent RAP verifiers test whether explicit group and card authorization improves safety under matched cases and reference audits.
- Structured verifiers: The staged verifier first selects one group, then returns an exact KEEP list of cards or NULL, with parse failures mapped to NULL.The prompts are frozen and evaluated without RecordAuth-Diag labels, demonstrations, training, or threshold selection.
- Evaluation criteria: Matched end-to-end tables evaluate identical cases and frozen decoders using stress unsafe exposure, positive recall, and NULL specificity.Stress unsafe measures unauthorized memory use on non-clean interventions, while positive recall prevents an abstention-only safety result.
- Structured verifiers: The one-pass control receives the same full-bank evidence as the group prompt but emits group and card decisions in a single call.It is information-matched rather than compute-matched, because the joint output is longer despite fewer input tokens.
- Independent pipeline: The independent RAP transfer uses the same cases while allowing the frozen guard to filter cards before native CLIP top-2 retrieval.RAP receives one diagnostic card per record using full-frame retrieval rather than detector crops.
- Reference robustness: Reference audits use held-out identities, temporal views, and split-overlap checks to test robustness beyond the primary diagnostic construction.The audit retains serialized split information and checks forbidden temporal-view overlap across DAVIS assets.
R OMNI-PERSONA VISUAL PROTOCOL AND COMPLETE ROUTING RESULTS
The Omni-Persona evaluation tests judge-free visual routing under a schema shift, while complete routing results separate identity selection from attribute support and preserve implementation audits.
- Visual protocol: Omni-Persona contains 231 visual items and 860 referenced assets, with four persona contexts interleaved per item.Its mapping reuses one context reference as both enrollment anchor and one-card group, removing anchor–card diversity.
- Visual protocol: The integrated verifier presents four context images and the query image, then requires exactly one context index or NULL before answering.Raw and strict arms use released context order and deterministic decoding with frozen SigLIP2 and ATRA thresholds.
- Routing results: Four-order unanimous voting lowers unsafe routing to 0.87/1.30% for Qwen/CoViP but lowers positive recall to 81.66/89.35%.This voting result is exploratory and excluded from confirmatory claims.
- Routing results: Frozen SigLIP2 similarity reaches 71.43% exact routing and 10.82% unsafe routing on the Omni evaluation.The result is a gate decision and does not verify whether the selected biography supports the requested attribute.
- Implementation audit: The composer audit finds byte-identical covered predictions and fixed abstentions for rejected predictions across bypass arms.Thus the Table 34 difference is attributed to operational NULL handling rather than regeneration or prompt editing.
- External transfer: All generator-level custom POPE and MMVP contamination cases produce abstentions for both raw and guarded Qwen/Phi paths.These retained null results show neither repair gain nor positive personalization on those contamination sets.
U ADDITIONAL MEASUREMENT AND COMPARISON LIMITATIONS
Additional measurements establish scope boundaries around training-corpus overlap, semantic attribution, filtering effects, and the architecture dependence of learned aggregation.
- Scope limitations: The complete training corpora of CoViP and RAP cannot be independently audited, so foundation-model image overlap remains possible.Frozen base Qwen and Phi isolate CoViP-specific post-training effects but do not eliminate generic pretraining exposure.
- Attribution limitations: Generator attribution is conservatively based on suffix-preserving equality after a synthetic-location intervention rather than semantic validity in free-form narratives.Human agreement and semantic-parser validity are not claimed until the specified blinded annotation packet is collected.
- Filtering interpretation: Filtering changes whether the released generator is invoked, while NULL bypass returns a fixed abstention when routing rejects.Table 34 therefore separates filtered generation from operational bypass behavior.
- Architecture comparison: Flat-MLP matches ATRA on SigLIP2 but is significantly worse on both accuracy and unsafe exposure with DINOv2.The result makes learned hierarchical aggregation an encoder-dependent robustness choice rather than a universal necessity.
V REPRODUCIBILITY CONTRACT
The reproducibility contract freezes model, intervention, analysis, and evidence registries, with fail-closed verification requirements and explicit reproducibility artifacts. It also reports generator-transfer and hardware-reproducibility results in dedicated tables.
- Registry requirements: 3,690 unique cases per core arm produce 44,280 aligned rows, while RAP arms contain 3,680 cases each and 18,400 rows after the registered parent exclusion.The registries also require frozen 10,000-replicate analyses and a current ten-test regression log.
- Registry requirements: 52 original inference endpoints, 21 analyses, and 14 static evidence summaries must pass fail-closed verification before generating the immutable V3 package.A separate post-review registry verifies 12 core model–intervention endpoints plus five RAP arms.
- Frozen software and models: The frozen package pins model revisions, metadata hashes, and RAP component revisions in a registry.The listed frozen revisions include Qwen3-VL-8B-Instruct, CoViP-Qwen3-VL-8B-GSPO, Phi-3.5-Vision, Gemma-3-4B-IT, SigLIP2-base-patch16-224, and DINOv2-base.
- Reproducibility results: Table 36 reports generator-level transfer on independent all-negative banks using raw-to-guard percentages and parent-cluster 95% intervals.Accuracy is task-defined NULL correctness, and abstention is shown explicitly so refusal is not counted as repair.
- Reproducibility results: Table 37 reports deterministic A800-to-A100 reproducibility on the registered 300-case subset.
- Additive audit registries: The additive Round-2 and Round-3 registries freeze prompts, parsers, fold decision sets, aggregate audits, model outputs, gate contrasts, and paired end-to-end analyses.The Round-3 registry also records a failed CoViP attempt and a clean retry covering the frozen fold exactly.