Information Retrieval

Papers filed under cs.IR on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

1,561 to 1,620 of 1,626

  1. Auxiliary uncertainty signals for LLM-assisted systematic review screening: a benchmark across eight Cohen drug-class reviews

    Arya Rahgozar, Pouria Mortezaagha

    cs.CLcs.DLcs.IRarXiv:2608.14551v12026
  2. The Commercial Tax: Rent-vs-Own Blind Spots in Multi-Hop Retrieval Benchmarks

    Luis M. Sanchez, Kosrow Dehnad

    cs.IRcs.CLarXiv:2608.16096v12026
  3. Static Pruning Across Sparse Retrieval Regimes: What Transfers, What Breaks, and What Still Helps

    Zirui Song, Yuye Zhu, Yang Yang

    cs.IRcs.AIarXiv:2608.16309v12026
  4. Decoupled Temporal Encoding for Generative Recommendation

    Pengfei Jia, Jingjian Wang, Jingmao Li +2

    cs.IRcs.AIarXiv:2608.16274v12026
  5. MVEB: Massive Video Embedding Benchmark

    Adnan El Assadi, Roman Solomatin, Isaac Chung +13

    cs.CVcs.IRcs.LGarXiv:2606.14958v12026
  6. Beyond Monolingual Deep Research: Evaluating Agents and Retrievers with Cross-Lingual BrowseComp-Plus

    Yuheng Lu, Qingcheng Zeng, Heli Qi +6

    cs.CLcs.IRarXiv:2606.15345v22026
  7. Sparse Coverage: Semantic Center Representations for Patent Prior-Art Retrieval

    You Zuo, Kim Gerdes, Éric de la Clergerie +1

    cs.IRcs.AIarXiv:2608.16918v12026
  8. pico-type: A 1.5M-Parameter Byte-Level Multi-Head Content Classifier

    Gautam Kishore

    cs.LGcs.AIcs.CLarXiv:2608.14658v12026
  9. NeuRoute: Logit-Guided Neural Routing for Billion-Scale Vector Search with Sub-Hour Index Construction

    Xingqiao Wang, Zi Wang, Xiaowei Xu

    cs.DBcs.IRcs.LGarXiv:2608.15438v12026
  10. SAGA: Structure-Attended Generative Action Embedding Model that encodes Multi-Surface User Action Sequences

    Tsz Fung Pang, Po Jen Chen, Nimish Ronghe +2

    cs.LGcs.IRarXiv:2608.15429v12026
  11. OGX: An Open-Source, Vendor-Neutral Generative AI Application Server

    Francisco Javier Arceo, Sébastien Han, Matthew Farrellee +8

    cs.AIcs.IRarXiv:2608.14580v12026
  12. MCompassRAG: Topic Metadata as a Semantic Compass for Paragraph-Level Retrieval

    Amirhossein Abaskohi, Raymond Li, Gaetano Cimino +3

    cs.CLcs.IRarXiv:2606.18508v12026
  13. SproutRAG: Attention-Guided Tree Search with Progressive Embeddings for Long-Document RAG

    Amirhossein Abaskohi, Issam H. Laradji, Peter West +1

    cs.CLcs.IRarXiv:2606.18381v12026
  14. ChartWalker: Benchmarking the Cross-Chart RAG Task with Hierarchical Knowledge Graphs

    Ning Tang, Chenghan Xie, Hanyang Yuan +6

    cs.IRarXiv:2606.23997v12026
  15. PrivacyAlign: Contextual Privacy Alignment for LLM Agents

    Manveer Singh Tamber, Abhay Puri, Marc-Etienne Brunet +3

    cs.CLcs.AIcs.IRarXiv:2606.21710v12026
  16. UniDot: A Unified Network for Sequence Modeling and Feature Interaction in Large-scale Recommendation

    Rongcheng Lin, Yan Sun, Jamey Zhang +4

    cs.IRcs.AIarXiv:2608.16797v12026
  17. Skill2Query: Exploiting Skill Structure to Generate Pseudo-Queries for Agent Skill Retrieval

    Lihui Ding, Zihan Guo, Bingwei Lu +5

    cs.CLcs.IRarXiv:2608.16071v12026
  18. From Student Risk Prediction to SC2R: Semantics-Constrained Counterfactual Recourse for Educational Decision Support

    Ngoc Luyen Le, Marie-Hélène Abel, Bertrand Laforge

    cs.IRcs.AIarXiv:2608.17618v12026
  19. DEPT: Document Embedding Preservation Tuning for Unified Query Expansion and Retrieval

    Jingyuan Wang, Richong Zhang, Zhijie Nie +2

    cs.IRcs.AIarXiv:2608.17632v12026
  20. Chain of Evidence: Pixel-Level Visual Attribution for Iterative Retrieval-Augmented Generation

    Peiyang Liu, Ziqiang Cui, Xi Wang +2

    cs.CVcs.AIcs.CLarXiv:2605.01284v22026
  21. ORBIT: Preserving Foundational Language Capabilities in GenRetrieval via Origin-Regulated Merging

    Neha Verma, Nikhil Mehta, Shao-Chuan Wang +7

    cs.CLcs.IRcs.LGarXiv:2605.12419v12026
  22. SciAtlas: A Large-Scale Knowledge Graph for Automated Scientific Research

    Shuofei Qiao, Yunxiang Wei, Jiazheng Fan +8

    cs.AIcs.CLcs.IRarXiv:2605.22878v12026
  23. When Tool-Backed Skill Retrieval Fails: Source-Style Collapse in Executable Capability Retrieval

    Yiqi Liu, Joseph James, Yang Wang +2

    cs.LGcs.IRarXiv:2608.16502v12026
  24. Xetrieval: Mechanistically Explaining Dense Retrieval

    Zhixin Cai, Jun Bai, Yang Liu +7

    cs.AIcs.IRarXiv:2605.29507v12026
  25. Exploring Autonomous Agentic Data Engineering for Model Specialization

    Yujie Luo, Xiangyuan Ru, Jingsheng Zheng +10

    cs.CLcs.AIcs.IRarXiv:2605.30407v22026
  26. Q-RAG: Long Context Multi-step Retrieval via Value-based Embedder Training

    Artyom Sorokin, Nazar Buzun, Alexander Anokhin +7

    cs.LGcs.IRarXiv:2511.07328v22025
  27. Coverage Is Not Containment: A Fundamental Limit of Admission-Time Defenses Against Coordinated Poisoning of Vector Retrieval

    Prashant Kumar Pathak, Tarun Kumar Sharma

    cs.CRcs.CLcs.IRarXiv:2608.16044v12026
  28. Grounding Healthcare LLMs in a Causal Knowledge Graph: Framework, Metrics, and a Cardiovascular Pilot

    Ummara Mumtaz, Aimen Noor, Awais Ahmed

    cs.AIcs.CLcs.IRarXiv:2608.15382v12026
  29. Ask to Be Sure: Informative Interactions for Confident Multi-Turn LLM Recommendation

    Cedar Site Bai, Duanshun Li, Zhenyu Liao +6

    cs.IRcs.AIcs.CLarXiv:2608.15949v12026
  30. Decomposing Staleness in Recommender Systems: A Dual-Filter Framework for Supersession and Decay

    Di Bai, Feng Han, Zhenwei Tang +3

    cs.IRcs.AIarXiv:2608.15780v12026
  31. ToolSense: A Diagnostic Framework for Auditing Parametric Tool Knowledge in LLMs

    Ashutosh Hathidara, Sai Shruthi Sistla, Sebastian Schreiber +1

    cs.AIcs.IRcs.LGarXiv:2606.12451v12026
  32. POI Recommendation with LLM-Augmented Multi-Graph Learning and Contrastive Alignment

    Burak Tamer, Wolfram Höpken, Zehui Wang

    cs.IRcs.LGarXiv:2608.16407v12026
  33. Cost Scales with Change, Not Corpus Size: Incrementally Maintaining an Evolving Semantic Substrate

    Yusuke Takahashi, Kyle Wild, Asako Uraki

    cs.AIcs.DBcs.IRarXiv:2608.16621v12026
  34. GEO-Flag: Detecting and Measuring GEO-Optimized Web Content

    Junjie Chu, Ye Leng, Mingjie Li +3

    cs.LGcs.CRcs.IRarXiv:2608.16824v12026
  35. Large language model-assisted discovery of cohorts from scientific literature

    Moritz Sturm, Lisa M. Berg, Inken Berg +6

    cs.IRcs.CLarXiv:2608.15909v12026
  36. Memory is Reconstructed, Not Retrieved: Graph Memory for LLM Agents

    Shuo Ji, Yibo Li, Bryan Hooi

    cs.AIcs.IRarXiv:2606.06036v12026
  37. GrepSeek: Training Search Agents for Direct Corpus Interaction

    Alireza Salemi, Chang Zeng, Atharva Nijasure +4

    cs.CLcs.AIcs.IRarXiv:2605.29307v12026
  38. ENTLORE: A Graph-Grounded Benchmark for Latent Organizational Reasoning in Enterprise Question Answering

    Akrin Zheng, Alexander Wu, Alaia Liu

    cs.IRcs.AIcs.CLarXiv:2608.10679v22026
  39. MemEye: A Visual-Centric Evaluation Framework for Multimodal Agent Memory

    Minghao Guo, Qingyue Jiao, Zeru Shi +14

    cs.CVcs.CLcs.IRarXiv:2605.15128v12026
  40. ConceptFormer: Learning Adaptive Latent Concepts for Query-Document Alignment in Visual Document Retrieval

    Peng Chunyi, Xu Zhipeng, Yan Yukun +9

    cs.CVcs.IRarXiv:2608.15698v12026
  41. Do AI chatbots find what experts would? Effects of model, user role, and sample size on study retrieval for medical questions

    Qingfang Liu, Qiao Jin, Joe D. Menke +2

    cs.IRcs.AIcs.CLarXiv:2608.13786v12026
  42. GRASP: GRanularity-Aware Search Policy for Agentic RAG

    Varun Gandhi, Jaewook Lee, Shantanu Todmal +4

    cs.AIcs.IRarXiv:2607.10463v12026
  43. Beyond Semantic Similarity: Rethinking Retrieval for Agentic Search via Direct Corpus Interaction

    Zhuofeng Li, Haoxiang Zhang, Cong Wei +16

    cs.IRcs.AIarXiv:2605.05242v12026
    Summaries:한국어
  44. Multi-Turn Agentic Scientific Literature Search via Workflow Induction

    Jisen Li, Bingxuan Li, Nanyi Jiang +10

    cs.CLcs.IRarXiv:2607.00597v22026
  45. Quantifying and Expanding the Theoretical Capacity of Late-Interaction Retrieval Models

    Julian Killingback, Varad Ingale, Hamed Zamani +1

    cs.IRarXiv:2607.05803v12026
  46. Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval

    Suhyeong Park, Junha Jung, Jungwoo Park +1

    cs.IRcs.AIcs.CLarXiv:2607.04605v22026
  47. SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration

    Yuyao Zhang, Junjie Gao, Zhengxian Wu +11

    cs.AIcs.IRarXiv:2607.15257v12026
  48. RecGPT-V3 Technical Report

    Bowen Zheng, Chao Yi, Dian Chen +26

    cs.IRarXiv:2607.15591v22026
  49. RL-Index: Reinforcement Learning for Retrieval Index Reasoning

    Yongjia Lei, Nedim Lipka, Zhisheng Qi +7

    cs.IRcs.AIcs.LGarXiv:2606.16316v22026
  50. TheoremGraph: Bridging Formal and Informal Mathematics

    Simon Kurgan, Evan Wang, Eric Leonen +6

    cs.IRcs.AImath.HOarXiv:2606.25363v12026
  51. Taste-aware music retrieval from audio embeddings

    Matteo Spanio, Antonio Rodà

    cs.SDcs.IRcs.LGarXiv:2607.03296v12026
  52. Evidence Attribution in Visual Document Understanding without Coordinates or Region Labels

    Zhuchenyang Liu, Yao Zhang, Yu Xiao

    cs.CVcs.CLcs.IRarXiv:2607.24651v12026
  53. A Frozen 12B Beats Frontier Models on Verified Work: 100% Accuracy, 0 Tokens, Bit-Exact, Forever

    Sietse Schelpe

    cs.CLcs.AIcs.IRarXiv:2607.23806v12026
  54. LAMAR: An Open Language-Aware Multilingual Alignment Reranker

    Seongtae Hong, Youngjoon Jang, Jungseob Lee +2

    cs.IRarXiv:2607.22042v22026
  55. AutoIndex: Learning Representation Programs for Retrieval

    Sam O'Nuallain, Nithya Rajkumar, Ramya Narayanasamy +3

    cs.IRcs.AIcs.CLarXiv:2607.18603v12026
  56. EMBL AI Librarian: Life-Sciences Knowledge Layer for AI Agents

    Luigi Sigillo, Matteo Silvestri, Francesco Tabaro +9

    cs.CLcs.AIcs.IRarXiv:2607.28229v12026
  57. UEmbed: Unified Sparse and Dense Multimodal Embeddings

    Tingyu Song, Mingxin Li, Yanzhao Zhang +5

    cs.CVcs.AIcs.CLarXiv:2608.02583v12026
  58. AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis

    Bing Yan, Gregory Wolfe, Stefano Martiniani +1

    cs.CLcs.AIcs.IRarXiv:2607.28618v12026
  59. RecHarness: A Bandit-Routed Agentic Harness for Self-Evolving Recommender Systems

    Haoran Ling, Yuecheng Li, Zeyu Song +5

    cs.IRcs.AIcs.CLarXiv:2607.29241v12026
  60. Search, Inspect, Fetch: Exploiting Structure-Aware Boolean Retrieval for Deep-Research Agents

    Shuai Wang, Haodong Chen, Yu Yin +3

    cs.IRcs.AIcs.CLarXiv:2608.02751v22026