Computation and Language

Papers filed under cs.CL on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

2,581 to 2,640 of 11,233

  1. RLHF Deciphered: A Critical Analysis of Reinforcement Learning from Human Feedback for LLMs

    Shreyas Chaudhari, Pranjal Aggarwal, Vishvak Murahari +5

    cs.LGcs.AIcs.CLarXiv:2404.08555v22024
  2. Evaluating Step-by-step Reasoning Traces: A Survey

    Jinu Lee, Julia Hockenmaier

    cs.CLarXiv:2502.12289v32025
  3. AutoAgent: A Fully-Automated and Zero-Code Framework for LLM Agents

    Jiabin Tang, Tianyu Fan, Chao Huang

    cs.AIcs.CLarXiv:2502.05957v32025
  4. A Comparative Survey of Recent Natural Language Interfaces for Databases

    Katrin Affolter, Kurt Stockinger, Abraham Bernstein

    cs.DBcs.CLcs.LGarXiv:1906.08990v12019
  5. Reduced, Reused and Recycled: The Life of a Dataset in Machine Learning Research

    Bernard Koch, Emily Denton, Alex Hanna +1

    cs.LGcs.CLcs.CVarXiv:2112.01716v12021
  6. Self-Alignment with Instruction Backtranslation

    Xian Li, Ping Yu, Chunting Zhou +5

    cs.CLarXiv:2308.06259v32023
  7. A Calibrated Reflection Approach for Enhancing Confidence Estimation in LLMs

    Umesh Bodhwani, Yuan Ling, Shujing Dong +3

    cs.CLarXiv:2609.04539v12026
  8. Deception Abilities Emerged in Large Language Models

    Thilo Hagendorff

    cs.CLcs.AIcs.LGarXiv:2307.16513v22023
  9. Accelerating LLM Inference with Staged Speculative Decoding

    Benjamin Spector, Chris Re

    cs.AIcs.CLarXiv:2308.04623v12023
  10. Mind the Value-Action Gap: Do LLMs Act in Alignment with Their Values?

    Hua Shen, Nicholas Clark, Tanushree Mitra

    cs.HCcs.AIcs.CLarXiv:2501.15463v42025
  11. DataSciBench: An LLM Agent Benchmark for Data Science

    Dan Zhang, Sining Zhoubian, Min Cai +7

    cs.CLcs.AIcs.LGarXiv:2502.13897v12025
  12. SkillCraft: Can LLM Agents Learn to Use Tools Skillfully?

    Shiqi Chen, Jingze Gai, Ruochen Zhou +13

    cs.CLcs.SEarXiv:2603.00718v22026
  13. Overview of the PsyDefDetect Shared Task at BioNLP 2026: Detecting Levels of Psychological Defense Mechanisms in Supportive Conversations

    Hongbin Na, Zimu Wang, Zhaoming Chen +8

    cs.CLarXiv:2605.24907v12026
  14. Refuse without Refusal: A Structural Analysis of Safety-Tuning Responses for Reducing False Refusals in Language Models

    Minji Kim, Hyounghun Kim

    cs.CLcs.AIarXiv:2609.04714v12026
  15. Why We Care About Understanding: Competence through Predictive Compression

    Matthieu Queloz, Pierre Beckmann

    cs.AIcs.CLarXiv:2609.04962v12026
  16. Embedding Text in Hyperbolic Spaces

    Bhuwan Dhingra, Christopher J. Shallue, Mohammad Norouzi +2

    cs.CLcs.LGarXiv:1806.04313v12018
  17. GraSP: Graph-Structured Skill Compositions for LLM Agents

    Tianle Xia, Lingxiang Hu, Yiding Sun +5

    cs.CLarXiv:2604.17870v12026
  18. Inducing Relational Knowledge from BERT

    Zied Bouraoui, Jose Camacho-Collados, Steven Schockaert

    cs.CLcs.AIarXiv:1911.12753v12019
  19. A Survey of Multimodal Retrieval-Augmented Generation

    Lang Mei, Siyu Mo, Zhihan Yang +1

    cs.IRcs.AIcs.CLarXiv:2504.08748v12025
  20. UXAgent: An LLM Agent-Based Usability Testing Framework for Web Design

    Yuxuan Lu, Bingsheng Yao, Hansu Gu +7

    cs.HCcs.CLarXiv:2502.12561v32025
  21. Which Attention Heads Matter for In-Context Learning?

    Kayo Yin, Jacob Steinhardt

    cs.LGcs.AIcs.CLarXiv:2502.14010v12025
  22. Knowledge Distillation and Dataset Distillation of Large Language Models: Emerging Trends, Challenges, and Future Directions

    Luyang Fang, Xiaowei Yu, Jiazhang Cai +23

    cs.CLcs.LGstat.MLarXiv:2504.14772v22025
  23. Exploring the Responses of Large Language Models to Beginner Programmers' Help Requests

    Arto Hellas, Juho Leinonen, Sami Sarsa +3

    cs.CYcs.AIcs.CLarXiv:2306.05715v12023
  24. Scaling Text-Rich Image Understanding via Code-Guided Synthetic Multimodal Data Generation

    Yue Yang, Ajay Patel, Matt Deitke +8

    cs.CVcs.CLarXiv:2502.14846v22025
  25. Do Massively Pretrained Language Models Make Better Storytellers?

    Abigail See, Aneesh Pappu, Rohun Saxena +2

    cs.CLcs.AIcs.LGarXiv:1909.10705v12019
  26. Efficient Reasoning with Hidden Thinking

    Xuan Shen, Yizhou Wang, Yufa Zhou +4

    cs.CLcs.AIcs.LGarXiv:2501.19201v22025
  27. Fusion of Detected Objects in Text for Visual Question Answering

    Chris Alberti, Jeffrey Ling, Michael Collins +1

    cs.CLcs.CVcs.LGarXiv:1908.05054v22019
  28. JSUT corpus: free large-scale Japanese speech corpus for end-to-end speech synthesis

    Ryosuke Sonobe, Shinnosuke Takamichi, Hiroshi Saruwatari

    cs.CLarXiv:1711.00354v12017
  29. Who's in Charge? Disempowerment Patterns in Real-World LLM Usage

    Mrinank Sharma, Miles McCain, Raymond Douglas +1

    cs.CYcs.AIcs.CLarXiv:2601.19062v12026
  30. A Survey on (M)LLM-Based GUI Agents

    Fei Tang, Haolei Xu, Hang Zhang +12

    cs.HCcs.AIcs.CLarXiv:2504.13865v22025
  31. An overview of model uncertainty and variability in LLM-based sentiment analysis. Challenges, mitigation strategies and the role of explainability

    David Herrera-Poyatos, Carlos Peláez-González, Cristina Zuheros +4

    cs.CLcs.AIarXiv:2504.04462v12025
  32. SoulChat: Improving LLMs' Empathy, Listening, and Comfort Abilities through Fine-tuning with Multi-turn Empathy Conversations

    Yirong Chen, Xiaofen Xing, Jingkai Lin +4

    cs.CLarXiv:2311.00273v12023
  33. XMerge: Cross-Axis Selection and Reconstructive Layer Merging for LLM Depth Compression

    Jundong Hu, Shekar Ramachandran

    cs.LGcs.CLarXiv:2609.02083v12026
  34. SOLID: A Large-Scale Semi-Supervised Dataset for Offensive Language Identification

    Sara Rosenthal, Pepa Atanasova, Georgi Karadzhov +2

    cs.CLarXiv:2004.14454v22020
  35. How do Large Language Models Handle Multilingualism?

    Yiran Zhao, Wenxuan Zhang, Guizhen Chen +2

    cs.CLcs.AIarXiv:2402.18815v32024
  36. Achieving Verified Robustness to Symbol Substitutions via Interval Bound Propagation

    Po-Sen Huang, Robert Stanforth, Johannes Welbl +5

    cs.CLcs.CRcs.LGarXiv:1909.01492v22019
  37. PlugMem: A Task-Agnostic Plugin Memory Module for LLM Agents

    Ke Yang, Zixi Chen, Xuan He +6

    cs.CLcs.AIcs.IRarXiv:2603.03296v12026
  38. That is a Known Lie: Detecting Previously Fact-Checked Claims

    Shaden Shaar, Giovanni Da San Martino, Nikolay Babulkov +1

    cs.CLcs.IRcs.LGarXiv:2005.06058v12020
  39. WebLINX: Real-World Website Navigation with Multi-Turn Dialogue

    Xing Han Lù, Zdeněk Kasner, Siva Reddy

    cs.CLcs.CVcs.LGarXiv:2402.05930v22024
  40. Benchmarking Prompt Sensitivity in Large Language Models

    Amirhossein Razavi, Mina Soltangheis, Negar Arabzadeh +3

    cs.CLcs.AIcs.IRarXiv:2502.06065v12025
  41. PLAID: An Efficient Engine for Late Interaction Retrieval

    Keshav Santhanam, Omar Khattab, Christopher Potts +1

    cs.IRcs.CLarXiv:2205.09707v12022
  42. Large Memory Layers with Product Keys

    Guillaume Lample, Alexandre Sablayrolles, Marc'Aurelio Ranzato +2

    cs.CLcs.LGarXiv:1907.05242v22019
  43. Values in the Wild: Discovering and Analyzing Values in Real-World Language Model Interactions

    Saffron Huang, Esin Durmus, Miles McCain +7

    cs.CLcs.AIcs.CYarXiv:2504.15236v12025
  44. The Dynamics of Continuous Mixture Collapse in Language Models

    Ali Backour

    cs.LGcs.CLarXiv:2609.02049v12026
  45. WinoQueer-NL: Assessing Bias in Dutch Language Models toward LGBTQ+ Identities

    Jiska Beuk, Gerasimos Spanakis

    cs.CLarXiv:2609.02651v12026
  46. Cross-Modal Retrieval in the Cooking Context: Learning Semantic Text-Image Embeddings

    Micael Carvalho, Rémi Cadène, David Picard +3

    cs.CLcs.CVcs.IRarXiv:1804.11146v12018
  47. LongCat-Flash Technical Report

    Meituan LongCat Team, Bayan, Bei Li +179

    cs.CLcs.AIcs.DCarXiv:2509.01322v22025
  48. Retrieval-Augmented Generation: A Comprehensive Survey of Architectures, Enhancements, and Robustness Frontiers

    Chaitanya Sharma

    cs.IRcs.CLarXiv:2506.00054v12025
  49. SETS: Leveraging Self-Verification and Self-Correction for Improved Test-Time Scaling

    Jiefeng Chen, Jie Ren, Xinyun Chen +4

    cs.AIcs.CLarXiv:2501.19306v52025
  50. PlotMachines: Outline-Conditioned Generation with Dynamic Plot State Tracking

    Hannah Rashkin, Asli Celikyilmaz, Yejin Choi +1

    cs.CLarXiv:2004.14967v22020
  51. EmoStance: Response-Side Affective-Orientation Control for Empathetic Response Generation via Emoji Weak Supervision

    Ziyuan Jin, Yuxuan Ge, Zheng Tian

    cs.AIcs.CLarXiv:2609.02133v12026
  52. Vision-Language Models for Edge Networks: A Comprehensive Survey

    Ahmed Sharshar, Latif U. Khan, Waseem Ullah +1

    cs.CVcs.AIcs.CLarXiv:2502.07855v22025
  53. SelfDoc: Self-Supervised Document Representation Learning

    Peizhao Li, Jiuxiang Gu, Jason Kuen +5

    cs.CVcs.CLarXiv:2106.03331v12021
  54. Towards Trustworthy Retrieval Augmented Generation for Large Language Models: A Survey

    Bo Ni, Zheyuan Liu, Leyao Wang +17

    cs.CLcs.AIarXiv:2502.06872v12025
  55. Content-Based Citation Recommendation

    Chandra Bhagavatula, Sergey Feldman, Russell Power +1

    cs.CLcs.DLcs.IRarXiv:1802.08301v12018
  56. Think Deep, Not Just Long: Measuring LLM Reasoning Effort via Deep-Thinking Tokens

    Wei-Lin Chen, Liqian Peng, Tian Tan +5

    cs.CLarXiv:2602.13517v22026
  57. DIET: Lightweight Language Understanding for Dialogue Systems

    Tanja Bunk, Daksh Varshneya, Vladimir Vlasov +1

    cs.CLarXiv:2004.09936v32020
  58. GME: Improving Universal Multimodal Retrieval by Multimodal LLMs

    Xin Zhang, Yanzhao Zhang, Wen Xie +7

    cs.CLcs.IRarXiv:2412.16855v22024
  59. Prompting4Debugging: Red-Teaming Text-to-Image Diffusion Models by Finding Problematic Prompts

    Zhi-Yi Chin, Chieh-Ming Jiang, Ching-Chun Huang +2

    cs.CLcs.CVarXiv:2309.06135v32023
  60. Large language models for automated scholarly paper review: A survey

    Zhenzhen Zhuang, Jiandong Chen, Hongfeng Xu +2

    cs.AIcs.CLcs.DLarXiv:2501.10326v22025