Source-linked AI summary
PhotoBench: Beyond Visual Matching Towards Personalized Intent-Driven Photo Retrieval
Tianyi Xu, Rong Shan, Junjie Wu, Jiadeng Huang, Teng Wang, Jiachen Zhu, Wenteng Chen, Minxin Tu, Quantao Dou, Zhaoxiang Wang, Changwang Zhang, Weinan Zhang, Jun Wang, Jianghao Lin
TL;DR
Personal photo retrieval requires reasoning over temporal, social, metadata, and visual context that existing isolated-image benchmarks do not capture. PhotoBench builds an authentic-album benchmark with multi-source profiling and trajectory-grounded intent queries, revealing modality gaps in unified embeddings and source-fusion failures in agentic systems. The paper therefore points toward agentic retrieval systems designed for multi-source constraint satisfaction.
Problem
Existing retrieval benchmarks use context-isolated images and shallow descriptive queries, omitting the personalized multi-source reasoning required by authentic photo-album requests.
Method
PhotoBench profiles each image with visual semantics, spatial-temporal metadata, social identity, and temporal events, then synthesizes intent-driven queries rooted in users’ life trajectories.
Results
Evaluations expose a modality gap in unified embedding models and a source fusion paradox in agentic systems performing complex personalized photo retrieval.
Takeaways & Limitations
Personal multimodal retrieval requires moving beyond unified embeddings toward robust, lightweight agentic systems for multi-source fusion and constraint satisfaction.
Takeaways & Limitations
Public release requires privacy screening that removes or masks highly private or identifying content from the collected albums.
Abstract
from arXiv · showhide
Personal photo albums are not merely collections of static images but living, ecological archives defined by temporal continuity, social entanglement, and rich metadata, which makes the personalized photo retrieval non-trivial. However, existing retrieval benchmarks rely heavily on context-isolated web snapshots, failing to capture the multi-source reasoning required to resolve authentic, intent-driven user queries. To bridge this gap, we introduce PhotoBench, the first benchmark constructed from authentic, personal albums. It is designed to shift the paradigm from visual matching to personalized multi-source intent-driven reasoning. Based on a rigorous multi-source profiling framework, which integrates visual semantics, spatial-temporal metadata, social identity, and temporal events for each image, we synthesize complex intent-driven queries rooted in users' life trajectories. Extensive evaluation on PhotoBench exposes two critical limitations: the modality gap, where unified embedding models collapse on non-visual constraints, and the source fusion paradox, where agentic systems perform poor tool orchestration. These findings indicate that the next frontier in personal multimodal retrieval lies beyond unified embeddings, necessitating robust agentic reasoning systems capable of precise constraint satisfaction and multi-source fusion. Our PhotoBench is available.
1 Introduction
PhotoBench addresses the mismatch between ecologically rich personal albums and benchmarks built from isolated images and shallow visual queries. It profiles multi-source personal context, synthesizes intent-driven queries, and reveals failures in both unified embeddings and agentic retrieval systems.
- Motivation: Personal albums combine temporal continuity, social relationships, metadata, and visual content, so authentic queries require multi-source reasoning rather than visual matching alone.Examples include events, social relationships, and spatial-temporal constraints.
- Limitations of Existing Benchmarks: Existing benchmarks lack ecological fidelity because web-scraped images omit the temporal continuity and rich metadata needed for temporal- and social-based reasoning.MSCOCO and Flickr30k are cited as context-isolated examples.
- Limitations of Existing Benchmarks: Existing datasets also provide shallow user intent through descriptive captions that fail to represent multi-source entanglement and non-visual constraints.INQUIRE and VisualNews are cited as examples.
- Findings: Evaluations expose a modality gap in unified embeddings and a source fusion paradox in agentic systems as query constraints and complexity increase.Unified embeddings collapse on precise non-visual constraints, while tool-equipped agents degrade non-linearly with complexity.
- Implications: The findings motivate lightweight agentic reasoning systems that can traverse modality gaps and perform precise multi-source fusion and constraint satisfaction.The paper presents this as a direction for personal multimodal retrieval in photo-album scenarios.
- PhotoBench: PhotoBench is introduced as a benchmark from authentic, metadata-rich personal albums, using multi-source profiling to evaluate reasoning beyond visual matching.Each image is modeled through visual semantics, spatial-temporal metadata, social identity, and temporal events.
- PhotoBench: Intent-driven query synthesis generates narrative queries from users’ life trajectories, followed by exhaustive ground-truth mining and zero-ground-truth queries for reliability evaluation.The benchmark targets complex personalized retrieval and rejection of plausible but nonexistent images.
2 Related Work
Related work progresses from visual-caption matching toward semantic, compositional, metadata-aware, and reasoning-based retrieval. PhotoBench is positioned within this shift while focusing on personalized multi-source retrieval.
- Multimodal Retrieval Benchmarks: Early multimodal retrieval benchmarks such as MSCOCO and Flickr30k evaluate simple matching between visual content and captions.Later datasets extend complexity through compositional reasoning and large-scale retrieval.
- Multimodal Retrieval Benchmarks: Recent benchmarks target compositional reasoning, large-scale retrieval, metadata, and conversational search across different task settings.Examples include Winoground, INQUIRE, Visual News, and VisDial.
- Multimodal Representation: Cross-modal representation methods build latent spaces from models such as CLIP and SigLIP, with later work improving discriminative power through hard negatives and batch mining.Generative multimodal backbones are also being used to train embedding models.
- Multimodal Representation: Existing latent semantic spaces can fail to decompose personal queries involving complex metadata or OCR results.This motivates retrieval approaches that handle information beyond unified visual-text representations.
- Agentic and Reasoning-based Retrieval: Reasoning-based retrieval systems increasingly call external tools, using agentic, multi-agent, or reasoning-agent frameworks for complex search.The cited directions include ReAct, Reflexion, MMSearch, AutoCIR, XR, MRA-CIR, MM-R1, and MMSearch-R1.
3 Dataset Construction
PhotoBench constructs an ecologically grounded benchmark from continuous personal albums by profiling multiple information sources and synthesizing trajectory-aware intent queries. It combines automated candidate mining, human verification, and zero-ground-truth cases to support dense evaluation and rejection testing.
- Dataset Construction: PhotoBench captures ecological validity through authentic, temporally continuous albums and intent-driven queries grounded in personalized multi-source contexts.The construction explicitly contrasts personal albums with isolated web-scale snapshots.
- Dataset Construction: The dataset construction has two stages: album collection with multi-source profiling, followed by intent-driven query synthesis from users’ life trajectories.Profiling integrates visual semantics, spatio-temporal metadata, social cues, and hierarchical event structures.
- Album Collection: Authentic albums are collected from a diverse demographic pool while preserving high-fidelity timestamps, GPS coordinates, device headers, and other original metadata.The collection protocol prioritizes metadata integrity and temporal continuity.
- Album Collection: Privacy screening removes or masks highly private material, while additional manual pruning and aesthetic curation are intentionally avoided.Participants flag sensitive content and experts conduct further review for public release.
- Multi-Source Profiling: Each image profile integrates visual features, semantic spatio-temporal metadata, social identity, and temporal events into a structured multi-source representation.Social identity uses face clustering and expert-assigned roles; temporal events are hierarchically clustered and summarized.
- Intent-Driven Query Synthesis: Intent synthesis infers a user intention from an anchor image, its profile, and preceding event summaries before composing natural queries from multiple information sources.Queries must remain consistent with the inferred intention and require source intersections to resolve ambiguity.
- Ground-Truth Mining: Ground-truth mining combines visual, semantic, and agentic multi-tool retrieval, with K=50 candidates manually reviewed to identify valid positives and ambiguous queries.This produces dense ground truth beyond sparse web-scale labels.
- Zero-Ground-Truth Evaluation: Counterfactual synthesis creates zero-ground-truth queries to test whether systems reject plausible but nonexistent images.The scenarios include people queried at events they did not attend.
4 Dataset Statistics
PhotoBench combines authentic personal-album images with dense metadata and a source-aware taxonomy to evaluate multi-source retrieval. Its queries exhibit long-tailed ground-truth counts and substantial compositional complexity.
- Dataset Composition: PhotoBench contains 3,582 images from three authentic personal albums and 1,188 bilingual Chinese-English queries.The benchmark provides dense, exhaustive ground truth within a continuous life-logging context.
- Dataset Characteristics: 83.4% of images retain valid high-precision GPS and timestamp metadata across a 2018–2025 temporal range.The albums span metropolitan areas in China and international points of interest across East and Southeast Asia.
- Source-Aware Taxonomy: The source-aware taxonomy classifies queries by necessary and sufficient information sources for precise failure attribution.Its atomic dimensions are Vision, Metadata, and Face, with compositional categories for multi-source intents.
- Query Distribution: Composite categories such as S_VM and S_VMF constitute a significant proportion of queries, emphasizing heterogeneous-signal fusion beyond visual matching.The taxonomy is strict and non-overlapping, so composite queries are not double-counted under their atomic sources.
- Query Distribution: Figure 2 reports query distributions by ground-truth count and source-aware taxonomy.The ground-truth distribution is long-tailed, spanning needle-in-a-haystack retrieval and broader event-level recall.
5 Experiment
The experiments compare unified embedding models with hybrid, tool-augmented retrieval systems on personalized intent-driven queries. Results show that explicit tool orchestration improves multi-source retrieval, but increasing constraint complexity exposes persistent fusion and reliability failures.
- Experimental Setup: The evaluation compares unified embedding models with hybrid retrieval systems that use multi-step, tool-augmented reasoning.Tool-based agents invoke vector search, metadata filtering, face search, and set composition tools.
- Main Results: Caption-based text embedding pipelines consistently underperform multimodal embedding models despite adding structured metadata.The paper attributes this gap to irreversible semantic loss when dense visual signals are converted into discrete textual intermediates.
- Main Results: Tool-based agentic systems significantly outperform unified embedding models by explicitly orchestrating heterogeneous signals for constrained personal-album retrieval.The reported gains support treating personal retrieval as a multi-source constrained problem rather than visual matching alone.
- Mobile Gallery Comparison: Agentic systems achieve higher set-level F1 than evaluated mobile systems on normal queries, but zero-ground-truth queries expose a trade-off between reasoning capability and reliability.Commercial systems generally use lightweight retrieval methods, while agents provide a higher performance ceiling for entangled queries.
- Source-Aware Analysis: Unified embeddings perform strongly on visual queries but collapse on explicit metadata or identity queries, where agents outperform them by vast margins.Embeddings remain competitive on some compositional queries because distinctive visual anchors can correlate with non-visual constraints.
- Tool Ablation: The metadata filter raises S_M F1 from 2.6 to 54.7, while the face engine raises S_F F1 from 8.7 to 69.0.For S_VMF, however, the full tool suite scores 32.5 F1 versus 35.1 with the visual tool alone, demonstrating source-fusion degradation.
- Complexity Scaling: Both agentic and commercial systems suffer significant performance decay when moving from single-source to dual-source queries.At the triple-source level, Phones B, C, and E instead improve by +15% to +30% over dual-source queries through a visual-anchor effect.
- Complexity Scaling: Commercial-system rebounds on triple-source queries reflect reliance on visual anchors rather than successful fusion of metadata and identity constraints.The systems can salvage recall by prioritizing visual similarity when non-visual logic fails.
6 Conclusion and Future Direction
PhotoBench shifts personal photo retrieval toward multi-source, intent-driven reasoning by modeling authentic album context and users’ life trajectories. Its findings point beyond stronger unified embeddings toward robust agentic systems for precise constraint satisfaction and heterogeneous-signal fusion.
- Conclusion: PhotoBench shifts mobile photo-retrieval evaluation from visual matching to multi-source, intent-driven reasoning.The benchmark reconstructs visual, spatial-temporal, social, and event information from authentic personal albums.
- Future Direction: The paper concludes that personal multimodal retrieval requires robust agentic reasoning systems capable of precise constraint satisfaction, proactive abstention, and reliable heterogeneous-signal fusion.This conclusion follows the benchmark’s exposure of modality-gap and source-fusion limitations in existing systems.
- Benchmark Construction: The construction pipeline synthesizes event narratives incrementally from clustered images, refining storylines through temporal connectors and contradiction correction.The final summaries are continuous, evidence-based, non-redundant, and globally coherent.
- Benchmark Construction: Capture-intention inference maps behavioral trajectories, photo roles, and motivations into a single declarative account of why an image was taken.Supported motivations include memory preservation, information retrieval, artistic expression, and social sharing.
- Benchmark Construction: Intent-driven query generation produces diverse colloquial event-centric, object-centric, and freeform queries using nicknames and metadata naturally.Queries are constrained to fewer than 15 characters and omit filler phrases such as “photo of”.
A.4 Zero-GT queries Synthesis Method
PhotoBench synthesizes zero-ground-truth queries that are plausible but unfulfillable, then verifies that no relevant images exist. It uses semantic deviations, contradictory metadata, and altered relational roles to test rejection behavior.
- Zero-GT Synthesis: Zero-GT queries are generated to test whether systems reject contextually plausible requests without visual matches.The synthesis methods target realistic but absent variants and conflicting constraints.
- Zero-GT Synthesis: Metadata perturbation introduces conflicting timestamps or altered relational roles, while semantic variation changes key query elements without creating a dataset match.These modifications preserve realistic descriptions while removing corresponding ground-truth images.
- Verification: Human experts verify that zero-GT queries have no relevant ground-truth images in the dataset.The verification process is shared with the benchmark’s broader query-validation procedure.
- Zero-GT Synthesis: Semantic deviation and constraint injection create realistic absent entities or scenes, hyper-specific details, and conflicting metadata.The constraints require colloquial realism and exhaustive VL-model cross-checking to ensure zero matches.
B.1 Metrics for Query Linguistic Analysis
The section defines linguistic and rejection metrics used to analyze PhotoBench queries and system behavior. It emphasizes that PhotoBench queries are shorter and structurally flatter while retaining diverse vocabulary.
- Query length and noun density: Average Query Length measures the mean number of tokens per query, while Noun Density measures the ratio of nouns and proper nouns to total tokens.Noun Density captures the concentration of informational entities.
- Syntactic depth: Syntactic Depth measures the maximum root-to-leaf path length in a dependency parse tree, using the maximum across sentences in multi-sentence queries.Lower depth indicates flatter, more fragmented search-style structure.
- Lexical diversity: MTLD evaluates vocabulary richness through a factor-based, bidirectional calculation that is robust to total text length.The score is based on total tokens divided by the average number of forward and backward factors.
- Lexical diversity: A higher MTLD score signifies broader and more specialized vocabulary.This metric distinguishes lexical richness from simple text length effects.
- Rejection ability: Reject-Precision, Reject-Recall, and Reject-F1 evaluate whether systems correctly identify queries with no relevant photo.Reject-F1 balances over-rejection and under-rejection through the harmonic mean of precision and recall.
- Overall linguistic profile: PhotoBench queries are structurally streamlined, with lower length and syntactic depth, yet remain lexically diverse and search-oriented.The linguistic comparison contrasts search-style queries with narrative-style benchmark language.
C.1 Comparsion with classic benchmarks
PhotoBench differs from classic retrieval benchmarks in both query language and ecological coverage. Its concise, search-style queries retain specialized vocabulary and are grounded in temporally continuous, socially varied personal albums with cognitive dimensions.
- Linguistic comparison: PhotoBench queries are shorter and syntactically shallower than COCO2014 and Flickr30k captions, favoring concise noun-phrase-centric search inputs.The comparison frames this as a shift from descriptive narration toward search-style retrieval.
- Linguistic comparison: PhotoBench nevertheless achieves greater lexical diversity than standard benchmarks, indicating specialized vocabulary for personalized contexts.Concise structure does not imply semantic simplicity.
- Album diversity: Its albums represent event-centric, person-centric, and balanced user archetypes with differing metadata coverage, portrait ratios, and social density.Album 1 has 97.7% metadata coverage, Album 2 has a 43.5% portrait ratio, and Album 3 is balanced.
- Temporal coverage: The benchmark spans 2018 to 2025, with staggered album activity producing relatively uniform aggregate temporal coverage.Interleaved album distributions reduce temporal bias and support long-range temporal evaluation.
- Source-aware taxonomy: Queries are annotated across Location, Time, Person, Object, and Concept dimensions as either Fact-based or Cognitive.A single query can contain multiple dimensions and multiple cognitive labels.
- Source-aware taxonomy: Most queries involve one to two cognitive dimensions, with Cognitive requirements especially prominent in Time and Concept.A cognitive-label count of zero denotes a purely Fact-based query or one without specific dimension tags.
D.3 The Counterintuitive Result
The counterintuitive result is that embedding models perform better on Cognitive than Fact queries, while agents face greater difficulty with cognitive multi-source queries. The analysis therefore shifts attention from reasoning depth to information-source accessibility.
- Counterintuitive result: Embedding models perform better on Cognitive queries than Fact queries, producing a −28.52 gap.The result contradicts the expectation that cognitive queries would be the primary difficulty.
- Modality gap: The lower Fact performance reflects a Modality Gap because many Fact queries depend on metadata such as dates or locations that visual embeddings cannot access.Cognitive queries can contain visually recognizable concepts such as “cozy moments.”
- Source fusion paradox: Agents find Cognitive queries slightly more challenging because resolving implicit dimensions requires complex tool-calling and information pruning.These operations increase the likelihood of error propagation.
- Diagnostic axis: Information-source accessibility, represented by VMF, is presented as a more actionable diagnostic axis than the Fact/Cognitive distinction.The paper prioritizes VMF in its main analysis based on the empirical data.
E Additional Experiments on English Version of PhotoBench
Additional experiments on the English version of PhotoBench produce results consistent with the primary findings. The authors present this consistency as evidence of benchmark generalizability and language independence.
- English-version experiments: The English PhotoBench experiments replicate the primary experimental findings.Table 10 provides the English version of Table 4 and reports Recall@10 performance by source-aware query type.
- English-version experiments: The authors state that these consistent results demonstrate benchmark generalizability and confirm that the observations are language-independent.The conclusion is based on the English-version experiments.
F Experimental Protocol of Commercial System
The evaluation combines standardized commercial-system testing, intent-focused case studies, and a specialized agentic framework for personal photo retrieval. Results illustrate the value of multi-source orchestration, while also defining a scope boundary relative to commercial systems’ specialized visual optimization.
- Experimental Protocol: Commercial systems were factory-reset, loaded with album images, indexed for 24 hours, queried through native gallery interfaces, and independently evaluated by two annotators.Returned results were recorded up to 100 per query.
- Intent-Driven Cases: Embedding retrieval returns visually similar but temporally or contextually incorrect results, whereas agentic retrieval filters by events, identities, time ranges, or locations to return the correct results or an explained empty set.The contrasting cases include business-trip receipts, New Year’s Eve dinner with parents, and nonexistent beach photos from the specified period.
- Agentic Retrieval Framework: The specialized framework combines a 4B VLM for planning, evaluation, and captioning with a 2B VLM embedding model and three-phase routing through rules, metadata or hybrid retrieval, and agentic synthesis.The framework is designed for the semantic complexity of gallery-based queries and personal photo management.
- Agentic Retrieval Framework: 63.3% F1 on normal queries and 52.0% Rej-F1 are reported as the framework’s strongest overall and rejection results, respectively.Recall is reported as 70.2%, slightly below the best agentic baseline, with the passage attributing this primarily to the smaller 2B retrieval model.
- Evaluation Scope: The benchmark does not aim to exhaustively evaluate every fine-grained visual category, and commercial-system results may diverge because evaluation dimensions and data distributions differ.Commercial systems are described as optimized for specialized visual scenarios and hardcoded metadata rules.