Information Retrieval
Papers filed under cs.IR on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
1,561 to 1,620 of 1,626
Auxiliary uncertainty signals for LLM-assisted systematic review screening: a benchmark across eight Cohen drug-class reviews
Arya Rahgozar, Pouria Mortezaagha
cs.CLcs.DLcs.IRarXiv:2608.14551v12026The Commercial Tax: Rent-vs-Own Blind Spots in Multi-Hop Retrieval Benchmarks
Luis M. Sanchez, Kosrow Dehnad
cs.IRcs.CLarXiv:2608.16096v12026Static Pruning Across Sparse Retrieval Regimes: What Transfers, What Breaks, and What Still Helps
Zirui Song, Yuye Zhu, Yang Yang
cs.IRcs.AIarXiv:2608.16309v12026Decoupled Temporal Encoding for Generative Recommendation
Pengfei Jia, Jingjian Wang, Jingmao Li +2
cs.IRcs.AIarXiv:2608.16274v12026MVEB: Massive Video Embedding Benchmark
Adnan El Assadi, Roman Solomatin, Isaac Chung +13
cs.CVcs.IRcs.LGarXiv:2606.14958v12026Beyond Monolingual Deep Research: Evaluating Agents and Retrievers with Cross-Lingual BrowseComp-Plus
Yuheng Lu, Qingcheng Zeng, Heli Qi +6
cs.CLcs.IRarXiv:2606.15345v22026Sparse Coverage: Semantic Center Representations for Patent Prior-Art Retrieval
You Zuo, Kim Gerdes, Éric de la Clergerie +1
cs.IRcs.AIarXiv:2608.16918v12026pico-type: A 1.5M-Parameter Byte-Level Multi-Head Content Classifier
Gautam Kishore
cs.LGcs.AIcs.CLarXiv:2608.14658v12026NeuRoute: Logit-Guided Neural Routing for Billion-Scale Vector Search with Sub-Hour Index Construction
Xingqiao Wang, Zi Wang, Xiaowei Xu
cs.DBcs.IRcs.LGarXiv:2608.15438v12026SAGA: Structure-Attended Generative Action Embedding Model that encodes Multi-Surface User Action Sequences
Tsz Fung Pang, Po Jen Chen, Nimish Ronghe +2
cs.LGcs.IRarXiv:2608.15429v12026OGX: An Open-Source, Vendor-Neutral Generative AI Application Server
Francisco Javier Arceo, Sébastien Han, Matthew Farrellee +8
cs.AIcs.IRarXiv:2608.14580v12026MCompassRAG: Topic Metadata as a Semantic Compass for Paragraph-Level Retrieval
Amirhossein Abaskohi, Raymond Li, Gaetano Cimino +3
cs.CLcs.IRarXiv:2606.18508v12026SproutRAG: Attention-Guided Tree Search with Progressive Embeddings for Long-Document RAG
Amirhossein Abaskohi, Issam H. Laradji, Peter West +1
cs.CLcs.IRarXiv:2606.18381v12026ChartWalker: Benchmarking the Cross-Chart RAG Task with Hierarchical Knowledge Graphs
Ning Tang, Chenghan Xie, Hanyang Yuan +6
cs.IRarXiv:2606.23997v12026PrivacyAlign: Contextual Privacy Alignment for LLM Agents
Manveer Singh Tamber, Abhay Puri, Marc-Etienne Brunet +3
cs.CLcs.AIcs.IRarXiv:2606.21710v12026UniDot: A Unified Network for Sequence Modeling and Feature Interaction in Large-scale Recommendation
Rongcheng Lin, Yan Sun, Jamey Zhang +4
cs.IRcs.AIarXiv:2608.16797v12026Skill2Query: Exploiting Skill Structure to Generate Pseudo-Queries for Agent Skill Retrieval
Lihui Ding, Zihan Guo, Bingwei Lu +5
cs.CLcs.IRarXiv:2608.16071v12026From Student Risk Prediction to SC2R: Semantics-Constrained Counterfactual Recourse for Educational Decision Support
Ngoc Luyen Le, Marie-Hélène Abel, Bertrand Laforge
cs.IRcs.AIarXiv:2608.17618v12026DEPT: Document Embedding Preservation Tuning for Unified Query Expansion and Retrieval
Jingyuan Wang, Richong Zhang, Zhijie Nie +2
cs.IRcs.AIarXiv:2608.17632v12026Chain of Evidence: Pixel-Level Visual Attribution for Iterative Retrieval-Augmented Generation
Peiyang Liu, Ziqiang Cui, Xi Wang +2
cs.CVcs.AIcs.CLarXiv:2605.01284v22026ORBIT: Preserving Foundational Language Capabilities in GenRetrieval via Origin-Regulated Merging
Neha Verma, Nikhil Mehta, Shao-Chuan Wang +7
cs.CLcs.IRcs.LGarXiv:2605.12419v12026SciAtlas: A Large-Scale Knowledge Graph for Automated Scientific Research
Shuofei Qiao, Yunxiang Wei, Jiazheng Fan +8
cs.AIcs.CLcs.IRarXiv:2605.22878v12026When Tool-Backed Skill Retrieval Fails: Source-Style Collapse in Executable Capability Retrieval
Yiqi Liu, Joseph James, Yang Wang +2
cs.LGcs.IRarXiv:2608.16502v12026Xetrieval: Mechanistically Explaining Dense Retrieval
Zhixin Cai, Jun Bai, Yang Liu +7
cs.AIcs.IRarXiv:2605.29507v12026Exploring Autonomous Agentic Data Engineering for Model Specialization
Yujie Luo, Xiangyuan Ru, Jingsheng Zheng +10
cs.CLcs.AIcs.IRarXiv:2605.30407v22026Q-RAG: Long Context Multi-step Retrieval via Value-based Embedder Training
Artyom Sorokin, Nazar Buzun, Alexander Anokhin +7
cs.LGcs.IRarXiv:2511.07328v22025Coverage Is Not Containment: A Fundamental Limit of Admission-Time Defenses Against Coordinated Poisoning of Vector Retrieval
Prashant Kumar Pathak, Tarun Kumar Sharma
cs.CRcs.CLcs.IRarXiv:2608.16044v12026Grounding Healthcare LLMs in a Causal Knowledge Graph: Framework, Metrics, and a Cardiovascular Pilot
Ummara Mumtaz, Aimen Noor, Awais Ahmed
cs.AIcs.CLcs.IRarXiv:2608.15382v12026Ask to Be Sure: Informative Interactions for Confident Multi-Turn LLM Recommendation
Cedar Site Bai, Duanshun Li, Zhenyu Liao +6
cs.IRcs.AIcs.CLarXiv:2608.15949v12026Decomposing Staleness in Recommender Systems: A Dual-Filter Framework for Supersession and Decay
Di Bai, Feng Han, Zhenwei Tang +3
cs.IRcs.AIarXiv:2608.15780v12026ToolSense: A Diagnostic Framework for Auditing Parametric Tool Knowledge in LLMs
Ashutosh Hathidara, Sai Shruthi Sistla, Sebastian Schreiber +1
cs.AIcs.IRcs.LGarXiv:2606.12451v12026POI Recommendation with LLM-Augmented Multi-Graph Learning and Contrastive Alignment
Burak Tamer, Wolfram Höpken, Zehui Wang
cs.IRcs.LGarXiv:2608.16407v12026Cost Scales with Change, Not Corpus Size: Incrementally Maintaining an Evolving Semantic Substrate
Yusuke Takahashi, Kyle Wild, Asako Uraki
cs.AIcs.DBcs.IRarXiv:2608.16621v12026GEO-Flag: Detecting and Measuring GEO-Optimized Web Content
Junjie Chu, Ye Leng, Mingjie Li +3
cs.LGcs.CRcs.IRarXiv:2608.16824v12026Large language model-assisted discovery of cohorts from scientific literature
Moritz Sturm, Lisa M. Berg, Inken Berg +6
cs.IRcs.CLarXiv:2608.15909v12026Memory is Reconstructed, Not Retrieved: Graph Memory for LLM Agents
Shuo Ji, Yibo Li, Bryan Hooi
cs.AIcs.IRarXiv:2606.06036v12026GrepSeek: Training Search Agents for Direct Corpus Interaction
Alireza Salemi, Chang Zeng, Atharva Nijasure +4
cs.CLcs.AIcs.IRarXiv:2605.29307v12026ENTLORE: A Graph-Grounded Benchmark for Latent Organizational Reasoning in Enterprise Question Answering
Akrin Zheng, Alexander Wu, Alaia Liu
cs.IRcs.AIcs.CLarXiv:2608.10679v22026MemEye: A Visual-Centric Evaluation Framework for Multimodal Agent Memory
Minghao Guo, Qingyue Jiao, Zeru Shi +14
cs.CVcs.CLcs.IRarXiv:2605.15128v12026ConceptFormer: Learning Adaptive Latent Concepts for Query-Document Alignment in Visual Document Retrieval
Peng Chunyi, Xu Zhipeng, Yan Yukun +9
cs.CVcs.IRarXiv:2608.15698v12026Do AI chatbots find what experts would? Effects of model, user role, and sample size on study retrieval for medical questions
Qingfang Liu, Qiao Jin, Joe D. Menke +2
cs.IRcs.AIcs.CLarXiv:2608.13786v12026GRASP: GRanularity-Aware Search Policy for Agentic RAG
Varun Gandhi, Jaewook Lee, Shantanu Todmal +4
cs.AIcs.IRarXiv:2607.10463v12026Beyond Semantic Similarity: Rethinking Retrieval for Agentic Search via Direct Corpus Interaction
Zhuofeng Li, Haoxiang Zhang, Cong Wei +16
cs.IRcs.AIarXiv:2605.05242v12026Summaries:한국어Multi-Turn Agentic Scientific Literature Search via Workflow Induction
Jisen Li, Bingxuan Li, Nanyi Jiang +10
cs.CLcs.IRarXiv:2607.00597v22026Quantifying and Expanding the Theoretical Capacity of Late-Interaction Retrieval Models
Julian Killingback, Varad Ingale, Hamed Zamani +1
cs.IRarXiv:2607.05803v12026Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval
Suhyeong Park, Junha Jung, Jungwoo Park +1
cs.IRcs.AIcs.CLarXiv:2607.04605v22026SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration
Yuyao Zhang, Junjie Gao, Zhengxian Wu +11
cs.AIcs.IRarXiv:2607.15257v12026RecGPT-V3 Technical Report
Bowen Zheng, Chao Yi, Dian Chen +26
cs.IRarXiv:2607.15591v22026RL-Index: Reinforcement Learning for Retrieval Index Reasoning
Yongjia Lei, Nedim Lipka, Zhisheng Qi +7
cs.IRcs.AIcs.LGarXiv:2606.16316v22026TheoremGraph: Bridging Formal and Informal Mathematics
Simon Kurgan, Evan Wang, Eric Leonen +6
cs.IRcs.AImath.HOarXiv:2606.25363v12026Taste-aware music retrieval from audio embeddings
Matteo Spanio, Antonio Rodà
cs.SDcs.IRcs.LGarXiv:2607.03296v12026Evidence Attribution in Visual Document Understanding without Coordinates or Region Labels
Zhuchenyang Liu, Yao Zhang, Yu Xiao
cs.CVcs.CLcs.IRarXiv:2607.24651v12026A Frozen 12B Beats Frontier Models on Verified Work: 100% Accuracy, 0 Tokens, Bit-Exact, Forever
Sietse Schelpe
cs.CLcs.AIcs.IRarXiv:2607.23806v12026LAMAR: An Open Language-Aware Multilingual Alignment Reranker
Seongtae Hong, Youngjoon Jang, Jungseob Lee +2
cs.IRarXiv:2607.22042v22026AutoIndex: Learning Representation Programs for Retrieval
Sam O'Nuallain, Nithya Rajkumar, Ramya Narayanasamy +3
cs.IRcs.AIcs.CLarXiv:2607.18603v12026EMBL AI Librarian: Life-Sciences Knowledge Layer for AI Agents
Luigi Sigillo, Matteo Silvestri, Francesco Tabaro +9
cs.CLcs.AIcs.IRarXiv:2607.28229v12026UEmbed: Unified Sparse and Dense Multimodal Embeddings
Tingyu Song, Mingxin Li, Yanzhao Zhang +5
cs.CVcs.AIcs.CLarXiv:2608.02583v12026AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis
Bing Yan, Gregory Wolfe, Stefano Martiniani +1
cs.CLcs.AIcs.IRarXiv:2607.28618v12026RecHarness: A Bandit-Routed Agentic Harness for Self-Evolving Recommender Systems
Haoran Ling, Yuecheng Li, Zeyu Song +5
cs.IRcs.AIcs.CLarXiv:2607.29241v12026Search, Inspect, Fetch: Exploiting Structure-Aware Boolean Retrieval for Deep-Research Agents
Shuai Wang, Haodong Chen, Yu Yin +3
cs.IRcs.AIcs.CLarXiv:2608.02751v22026