Information Retrieval

Papers filed under cs.IR on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

1 to 60 of 1,626

  1. HABERTOR: An Efficient and Effective Deep Hatespeech Detector

    Thanh Tran, Yifan Hu, Changwei Hu +4

    cs.CLcs.AIcs.IRarXiv:2010.08865v12020
  2. Finding Generalizable Evidence by Learning to Convince Q&A Models

    Ethan Perez, Siddharth Karamcheti, Rob Fergus +3

    cs.CLcs.AIcs.IRarXiv:1909.05863v12019
  3. Keyword search is all you need: Achieving RAG-Level Performance without vector databases using agentic tool use

    Shreyas Subramanian, Adewale Akinfaderin, Yanyan Zhang +4

    cs.IRcs.AIarXiv:2602.23368v12025
  4. Enabling Large Language Models to Generate Text with Citations

    Tianyu Gao, Howard Yen, Jiatong Yu +1

    cs.CLcs.IRcs.LGarXiv:2305.14627v22023
  5. Offline Evaluation Measures of Fairness in Recommender Systems

    Theresia Veronika Rampisela

    cs.IRarXiv:2604.25032v12026
  6. NovBench: Evaluating Large Language Models on Academic Paper Novelty Assessment

    Wenqing Wu, Yi Zhao, Yuzhuo Wang +4

    cs.CLcs.AIcs.DLarXiv:2604.11543v12026
  7. MGDiff: Multi-Interest Sequence Recommendation with Masking GNN-Guided Diffusion

    Wenjing Xiao, Hao Ding

    cs.IRarXiv:2609.01619v12026
  8. VideoSET: Video Summary Evaluation through Text

    Serena Yeung, Alireza Fathi, Li Fei-Fei

    cs.CVcs.CLcs.IRarXiv:1406.5824v12014
  9. Equal Ranking Quality, Different Decisions: Training Order-Consistent LLM Scorers

    Markus Frohmann, Mahdiyar Alavi, Elizabeth Lingg +1

    cs.CLcs.IRcs.LGarXiv:2608.26762v12026
  10. What is Wrong with Topic Modeling? (and How to Fix it Using Search-based Software Engineering)

    Amritanshu Agrawal, Wei Fu, Tim Menzies

    cs.SEcs.AIcs.CLarXiv:1608.08176v42016
  11. OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations

    Linke Ouyang, Yuan Qu, Hongbin Zhou +17

    cs.CVcs.AIcs.IRarXiv:2412.07626v22024
  12. Measuring Semantic Similarity by Latent Relational Analysis

    Peter D. Turney

    cs.LGcs.CLcs.IRarXiv:cs/0508053v12005
  13. Learn Before Represent: Bridging Generative and Contrastive Learning for Domain-Specific LLM Embeddings

    Xiaoyu Liang, Yuchen Peng, Jiale Luo +3

    cs.IRcs.AIarXiv:2601.11124v12026
  14. MagicLens: Self-Supervised Image Retrieval with Open-Ended Instructions

    Kai Zhang, Yi Luan, Hexiang Hu +5

    cs.CVcs.AIcs.CLarXiv:2403.19651v22024
  15. Imagine All The Relevance: Scenario-Profiled Indexing with Knowledge Expansion for Dense Retrieval

    Sangam Lee, Ryang Heo, SeongKu Kang +1

    cs.IRarXiv:2503.23033v22025
  16. Fair and Diverse DPP-based Data Summarization

    L. Elisa Celis, Vijay Keswani, Damian Straszak +3

    cs.LGcs.CYcs.IRarXiv:1802.04023v12018
  17. Addressing the Item Cold-start Problem by Attribute-driven Active Learning

    Yu Zhu, Jinhao Lin, Shibi He +4

    cs.IRcs.LGstat.MLarXiv:1805.09023v12018
  18. Understanding AI Provider Recommendations in Local Service Markets

    Hazem Ibrahim, Yasir Zaki

    cs.CYcs.CLcs.IRarXiv:2609.18341v12026
  19. Quanta: A Self-Contained Python Library for Hybrid Retrieval over Quantised Embeddings, Lexical Indexes, and Knowledge Graphs

    Ioannis E. Livieris

    cs.IRcs.AIarXiv:2609.18248v12026
  20. LIGE-GR: A Smooth Leap from Ranking to Generative Recommendation in the LLM Era

    Venkat Srinivas, Chenzhang He, Sam Woodmansee +59

    cs.LGcs.IRarXiv:2609.18148v12026
  21. Beyond Static RAG: An Adaptive, Tri-Metric Routing Framework for Efficient Long-Context Inference on Commodity GPUs

    Saipraveen Vabbilisetty, Ajay Kumar Boddepalli, Deep Narayan Mishra +3

    cs.LGcs.IRarXiv:2609.17564v12026
  22. GenRec: Large Language Model for Generative Recommendation

    Jianchao Ji, Zelong Li, Shuyuan Xu +4

    cs.IRcs.AIcs.CLarXiv:2307.00457v22023
  23. Can We Do Interpretable NLI with Graphs Based on Atomic Propositions?

    Younes Boufouss, Luc Pommeret, Thomas Gerald +2

    cs.AIcs.IRarXiv:2609.16814v22026
  24. PinDCO: Whole-Page Aware Dynamic Creative Optimization at Scale

    Yu Hao, Yuchun Li, Peimeng Sui +5

    cs.IRcs.LGarXiv:2609.11943v12026
  25. Who Are We Recommending To? Recommender Systems in the Agentic Web

    Himan Abdollahpouri, Kyle Kretschman, Sai Ravindranath +2

    cs.IRarXiv:2609.11945v12026
  26. Position: Recommender Systems Should Move Beyond Platform-Centric Ranking toward Personal Agent-Mediated Recommendation

    Haohan Yuan, Peng He, Dan Zhang +2

    cs.IRarXiv:2609.11942v12026
  27. MemRetriever: Learning to Search, Reflect, and Retrieve from Long-Term Memory

    Ruiyang Jiang, Chunyu Li, Zhiyu Li

    cs.IRarXiv:2609.11951v12026
  28. The Death of Schema Linking? Text-to-SQL in the Age of Well-Reasoned Language Models

    Karime Maamari, Fadhil Abubaker, Daniel Jaroslawicz +1

    cs.CLcs.AIcs.IRarXiv:2408.07702v22024
  29. InitGen: Candidate Generation for Interaction Initiation in Intelligent Assistants

    Ruize Shi, Jinhua Chen, Hong Huang +5

    cs.IRarXiv:2609.11953v12026
  30. Let the LLMs Talk: Simulating Human-to-Human Conversational QA via Zero-Shot LLM-to-LLM Interactions

    Zahra Abbasiantaeb, Yifei Yuan, Evangelos Kanoulas +1

    cs.CLcs.AIcs.IRarXiv:2312.02913v12023
  31. Contrastive Triple Extraction with Generative Transformer

    Hongbin Ye, Ningyu Zhang, Shumin Deng +4

    cs.CLcs.AIcs.DBarXiv:2009.06207v82020
  32. Semantics-Aware Denoising: A PLM-Guided Sample Reweighting Strategy for Robust Recommendation

    Xikai Yang, Yang Wang, Yilin Li +1

    cs.IRarXiv:2602.15359v12026
  33. Jointly Cross- and Self-Modal Graph Attention Network for Query-Based Moment Localization

    Daizong Liu, Xiaoye Qu, Xiao-Yang Liu +3

    cs.CVcs.IRarXiv:2008.01403v22020
  34. Beyond Quacking: Deep Integration of Language Models and RAG into DuckDB

    Anas Dorbani, Sunny Yasser, Jimmy Lin +1

    cs.DBcs.AIcs.IRarXiv:2504.01157v12025
  35. CORE: Simple and Effective Session-based Recommendation within Consistent Representation Space

    Yupeng Hou, Binbin Hu, Zhiqiang Zhang +1

    cs.IRcs.AIarXiv:2204.11067v12022
  36. Prompting for Multimodal Hateful Meme Classification

    Rui Cao, Roy Ka-Wei Lee, Wen-Haw Chong +1

    cs.CLcs.IRcs.MMarXiv:2302.04156v12023
  37. Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation

    Koutian Wu, Junjie Zhou, Ergan Shang +5

    cs.AIcs.IRarXiv:2609.11115v22026
  38. Response Ranking with Deep Matching Networks and External Knowledge in Information-seeking Conversation Systems

    Liu Yang, Minghui Qiu, Chen Qu +5

    cs.IRcs.CLarXiv:1805.00188v32018
  39. RankMixer: Scaling Up Ranking Models in Industrial Recommenders

    Jie Zhu, Zhifang Fan, Xiaoxie Zhu +18

    cs.IRarXiv:2507.15551v32025
  40. A Comprehensive Survey and Experimental Comparison of Graph-Based Approximate Nearest Neighbor Search

    Mengzhao Wang, Xiaoliang Xu, Qiang Yue +1

    cs.IRcs.DBarXiv:2101.12631v22021
  41. A Comparison of Word Embeddings for the Biomedical Natural Language Processing

    Yanshan Wang, Sijia Liu, Naveed Afzal +5

    cs.IRarXiv:1802.00400v32018
  42. Digital Nudging with Recommender Systems: Survey and Future Directions

    Mathias Jesse, Dietmar Jannach

    cs.HCcs.IRarXiv:2011.03413v22020
  43. Clicks can be Cheating: Counterfactual Recommendation for Mitigating Clickbait Issue

    Wenjie Wang, Fuli Feng, Xiangnan He +2

    cs.IRarXiv:2009.09945v42020
  44. Leveraging Low-Level Symbolic Competences for Unsupervised Grounding in Hallucination Detection

    Renato Vukovic, Hsien-chin Lin, Carel van Niekerk +5

    cs.CLcs.AIcs.IRarXiv:2609.05025v12026
  45. Efficient and robust approximate nearest neighbor search using Hierarchical Navigable Small World graphs

    Yu. A. Malkov, D. A. Yashunin

    cs.DScs.CVcs.IRarXiv:1603.09320v42016
  46. MedCPT: Contrastive Pre-trained Transformers with Large-scale PubMed Search Logs for Zero-shot Biomedical Information Retrieval

    Qiao Jin, Won Kim, Qingyu Chen +4

    cs.IRcs.AIcs.CLarXiv:2307.00589v22023
  47. STREAM: An Objective-Driven and Uncertainty-Aware Framework for Industrial Energy Data Acquisition

    Zhipeng Ma, Bo Nørregaard Jørgensen, Zheng Grace Ma

    cs.IRarXiv:2608.26754v12026
  48. Twitter Sentiment Analysis: Lexicon Method, Machine Learning Method and Their Combination

    Olga Kolchyna, Tharsis T. P. Souza, Philip Treleaven +1

    cs.CLcs.IRcs.LGarXiv:1507.00955v32015
  49. An Event is Worth One Token: Event Tokenization for Industrial-scale LLM Recommendation

    Fan Xia, Zhaoheng Zheng, Iman Setayesh +12

    cs.IRarXiv:2608.25546v12026
  50. A Dataset for Movie Description

    Anna Rohrbach, Marcus Rohrbach, Niket Tandon +1

    cs.CVcs.CLcs.IRarXiv:1501.02530v12015
  51. DGRec: Graph Neural Network for Recommendation with Diversified Embedding Generation

    Liangwei Yang, Shengjie Wang, Yunzhe Tao +4

    cs.IRarXiv:2211.10486v22022
  52. Natural Language Processing of Clinical Notes on Chronic Diseases: Systematic Review

    Seyedmostafa Sheikhalishahi, Riccardo Miotto, Joel T Dudley +3

    cs.CYcs.AIcs.CLarXiv:1908.05780v12019
  53. Video In Sentences Out

    Andrei Barbu, Alexander Bridge, Zachary Burchill +15

    cs.CVcs.CLcs.IRarXiv:1408.6418v12014
  54. Redakto - The Incognito Tab for LLMs

    Saurav Kumar Saha, Tom Röhr, Felix Bießmann

    cs.AIcs.CLcs.CRarXiv:2608.18260v12026
  55. Sparse Subspace Clustering: Algorithm, Theory, and Applications

    Ehsan Elhamifar, Rene Vidal

    cs.CVcs.IRcs.ITarXiv:1203.1005v32012
  56. IndicMedDialog: A Parallel Multi-Turn Medical Dialogue Dataset for Accessible Healthcare in Indic Languages

    Shubham Kumar Nigam, Suparnojit Sarkar, Piyush Patel

    cs.CLcs.AIcs.IRarXiv:2605.13292v12026
  57. From Frequency to Meaning: Vector Space Models of Semantics

    Peter D. Turney, Patrick Pantel

    cs.CLcs.IRcs.LGarXiv:1003.1141v12010
  58. $\mathrm{ECI}_{\mathrm{sem}}$: Semantic Residual Effective Contrastive Information for Evaluating Hard Negatives

    Aarush Sinha, Rahul Seetharaman, Aman Bansal

    cs.IRcs.AIarXiv:2603.20990v42026
  59. Generating Sentiment-Preserving Fake Online Reviews Using Neural Language Models and Their Human- and Machine-based Detection

    David Ifeoluwa Adelani, Haotian Mai, Fuming Fang +3

    cs.CLcs.CRcs.IRarXiv:1907.09177v22019
  60. Autoregressive Entity Retrieval

    Nicola De Cao, Gautier Izacard, Sebastian Riedel +1

    cs.CLcs.IRcs.LGarXiv:2010.00904v32020