Computation and Language

Papers filed under cs.CL on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

2,941 to 3,000 of 11,403

  1. Ask in Any Modality: A Comprehensive Survey on Multimodal Retrieval-Augmented Generation

    Mohammad Mahdi Abootorabi, Amirhosein Zobeiri, Mahdi Dehghani +5

    cs.CLcs.AIcs.IRarXiv:2502.08826v32025
  2. ResearchBench: Benchmarking LLMs in Scientific Discovery via Inspiration-Based Task Decomposition

    Yujie Liu, Zonglin Yang, Tong Xie +7

    cs.CLcs.AIcs.CEarXiv:2503.21248v32025
  3. Disentangling Statistical Preemption from Entrenchment in Language Models' Avoidance of Overgeneralization

    Yixuan Wang, Freda Shi, Kanishka Misra

    cs.CLarXiv:2609.01794v12026
  4. InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model

    Yuhang Zang, Xiaoyi Dong, Pan Zhang +10

    cs.CVcs.CLarXiv:2501.12368v22025
  5. Stochastic Language Generation in Dialogue using Recurrent Neural Networks with Convolutional Sentence Reranking

    Tsung-Hsien Wen, Milica Gasic, Dongho Kim +4

    cs.CLarXiv:1508.01755v12015
  6. Aligning Cross-Lingual Entities with Multi-Aspect Information

    Hsiu-Wei Yang, Yanyan Zou, Peng Shi +3

    cs.CLarXiv:1910.06575v12019
  7. AVERT: Audio-Verified Adjudication for Spoken Dialogue State Tracking

    Chunggi Lee, Hanspeter Pfister

    cs.CLarXiv:2609.01828v12026
  8. Interpretable Symptom Vectors for Depression in a Large Language Model

    Fangyi Zhu, Ajay Subramanian, Allison Constant +3

    cs.CLcs.AIcs.LGarXiv:2609.01832v12026
  9. Forgetting Transformer: Softmax Attention with a Forget Gate

    Zhixuan Lin, Evgenii Nikishin, Xu Owen He +1

    cs.LGcs.AIcs.CLarXiv:2503.02130v22025
  10. The Obvious Invisible Threat: LLM-Powered GUI Agents' Vulnerability to Fine-Print Injections

    Chaoran Chen, Zhiping Zhang, Bingcan Guo +8

    cs.HCcs.CLcs.CRarXiv:2504.11281v12025
  11. TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling

    Yizhi Li, Qingshui Gu, Zhoufutu Wen +14

    cs.LGcs.CLarXiv:2508.17445v12025
  12. SensorLM: Learning the Language of Wearable Sensors

    Yuwei Zhang, Kumar Ayush, Siyuan Qiao +17

    cs.LGcs.AIcs.CLarXiv:2506.09108v12025
  13. RealSafe-R1: Safety-Aligned DeepSeek-R1 without Compromising Reasoning Capability

    Yichi Zhang, Zihao Zeng, Dongbai Li +3

    cs.AIcs.CLarXiv:2504.10081v12025
  14. Self-contradictory Hallucinations of Large Language Models: Evaluation, Detection and Mitigation

    Niels Mündler, Jingxuan He, Slobodan Jenko +1

    cs.CLcs.AIcs.LGarXiv:2305.15852v32023
  15. Rethinking Mixture-of-Agents: Is Mixing Different Large Language Models Beneficial?

    Wenzhe Li, Yong Lin, Mengzhou Xia +1

    cs.CLcs.LGarXiv:2502.00674v12025
  16. MemInsight: Autonomous Memory Augmentation for LLM Agents

    Rana Salama, Jason Cai, Michelle Yuan +4

    cs.CLarXiv:2503.21760v22025
  17. When Truth Is Overridden: Uncovering the Internal Origins of Sycophancy in Large Language Models

    Keyu Wang, Jin Li, Shu Yang +2

    cs.CLarXiv:2508.02087v42025
  18. MathTutorBench: A Benchmark for Measuring Open-ended Pedagogical Capabilities of LLM Tutors

    Jakub Macina, Nico Daheim, Ido Hakimi +3

    cs.CLcs.AIcs.LGarXiv:2502.18940v22025
  19. Missing Premise exacerbates Overthinking: Are Reasoning Models losing Critical Thinking Skill?

    Chenrui Fan, Ming Li, Lichao Sun +1

    cs.AIcs.CLcs.LGarXiv:2504.06514v22025
  20. PRISM: Self-Pruning Intrinsic Selection Method for Training-Free Multimodal Data Selection

    Jinhe Bi, Aniri, Zengjie Jin +11

    cs.CVcs.AIcs.CLarXiv:2502.12119v52025
  21. Part I: Tricks or Traps? A Deep Dive into RL for LLM Reasoning

    Zihe Liu, Jiashun Liu, Yancheng He +13

    cs.LGcs.CLarXiv:2508.08221v32025
  22. AI4Research: A Survey of Artificial Intelligence for Scientific Research

    Qiguang Chen, Mingda Yang, Libo Qin +13

    cs.CLcs.AIarXiv:2507.01903v22025
  23. How Does BERT Answer Questions? A Layer-Wise Analysis of Transformer Representations

    Betty van Aken, Benjamin Winter, Alexander Löser +1

    cs.CLcs.IRarXiv:1909.04925v12019
  24. ACEBench: Who Wins the Match Point in Tool Usage?

    Chen Chen, Xinlong Hao, Weiwen Liu +13

    cs.CLarXiv:2501.12851v82025
  25. SuperBPE: Space Travel for Language Models

    Alisa Liu, Jonathan Hayase, Valentin Hofmann +3

    cs.CLcs.LGarXiv:2503.13423v32025
  26. Jailbreaking Multimodal Large Language Models via Shuffle Inconsistency

    Shiji Zhao, Ranjie Duan, Fengxiang Wang +7

    cs.CRcs.AIcs.CLarXiv:2501.04931v22025
  27. Waypoint Models for Instruction-guided Navigation in Continuous Environments

    Jacob Krantz, Aaron Gokaslan, Dhruv Batra +2

    cs.CVcs.CLcs.ROarXiv:2110.02207v12021
  28. Contamination Inflates Scores but Rarely Reorders Large Language Model Leaderboards

    Xingyao Xiao, Yihong Cheng

    cs.CLstat.APstat.MEarXiv:2609.02899v12026
  29. Calibrated Language Models Must Hallucinate

    Adam Tauman Kalai, Santosh S. Vempala

    cs.CLcs.AIarXiv:2311.14648v32023
  30. RePolicy: Reinforcement Learning for Safety-Policy Invocation in Agent Safeguards

    Houcheng Jiang, Boxuan Zhang, Qiyong Zhong +3

    cs.AIcs.CLarXiv:2608.24275v12026
  31. How to Train a Critic Stably and Efficiently

    Penghui Qi, Xiangxin Zhou, Wee Sun Lee

    cs.LGcs.AIcs.CLarXiv:2608.23566v12026
  32. The politics of postmortem privacy

    Mauricio Figueroa

    cs.CYcs.CLcs.SIarXiv:2608.16905v12026
  33. CAST: Game Solvers as Turn-Level Teachers for LLM Agents

    Yu Wang, Yi-Kai Zhang, Wentao Shi +8

    cs.CLcs.AIarXiv:2607.25308v12026
  34. Graph-Native Reinforcement Learning Enables Traceable Scientific Hypothesis Generation through Conceptual Recombination

    Subhadeep Pal, Shashwat Sourav, Tirthankar Ghosal +1

    cs.AIcond-mat.mtrl-scics.CLarXiv:2607.00924v12026
    Summaries:한국어
  35. Visual Framing for News Stance Detection via Image Generation

    Dahyun Lee, Jiyoung Han, Kunwoo Park

    cs.CLcs.AIcs.CVarXiv:2609.00685v12026
  36. Hints Help But Do They Teach? Evaluating Skills Transfer in Code Generation

    Will Badr

    cs.SEcs.AIcs.CLarXiv:2609.01106v12026
  37. MMBench-GUI: Hierarchical Multi-Platform Evaluation Framework for GUI Agents

    Xuehui Wang, Zhenyu Wu, JingJing Xie +25

    cs.CVcs.CLarXiv:2507.19478v12025
  38. OS-ATLAS: A Foundation Action Model for Generalist GUI Agents

    Zhiyong Wu, Zhenyu Wu, Fangzhi Xu +8

    cs.CLcs.CVcs.HCarXiv:2410.23218v12024
  39. Look Inward to Explore Outward: Learning Temperature Policy from LLM Internal States via Hierarchical RL

    Yixiao Zhou, Yang Li, Dongzhou Cheng +2

    cs.LGcs.AIcs.CLarXiv:2602.13035v12026
  40. Seed-Coder: Let the Code Model Curate Data for Itself

    ByteDance Seed, Yuyu Zhang, Jing Su +24

    cs.CLcs.SEarXiv:2506.03524v22025
  41. AI chatbots versus human healthcare professionals: a systematic review and meta-analysis of empathy in patient care

    Alastair Howcroft, Amber Bennett-Weston, Ahmad Khan +3

    cs.HCcs.AIcs.CLarXiv:2602.05628v12026
  42. CodeI/O: Condensing Reasoning Patterns via Code Input-Output Prediction

    Junlong Li, Daya Guo, Dejian Yang +3

    cs.CLcs.AIarXiv:2502.07316v42025
  43. OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models

    Han Zhu, Lingxuan Ye, Wei Kang +7

    cs.CLeess.ASarXiv:2604.00688v32026
  44. Understanding and Improving Information Transfer in Multi-Task Learning

    Sen Wu, Hongyang R. Zhang, Christopher Ré

    cs.LGcs.CLarXiv:2005.00944v12020
  45. Large Language Model Routing with Benchmark Datasets

    Tal Shnitzer, Anthony Ou, Mírian Silva +5

    cs.CLcs.LGarXiv:2309.15789v12023
  46. Latent On-Policy Self-Distillation

    Guibin Zhang, Jiayang Lyu, Ran Sun +4

    cs.LGcs.CLarXiv:2608.13040v12026
  47. Can Sensitive Information Be Deleted From LLMs? Objectives for Defending Against Extraction Attacks

    Vaidehi Patil, Peter Hase, Mohit Bansal

    cs.CLcs.AIcs.LGarXiv:2309.17410v12023
  48. Covo-Audio Technical Report

    Wenfu Wang, Chenxing Li, Liqiang Zhang +23

    cs.SDcs.CLeess.ASarXiv:2602.09823v22026
  49. Specializing Large Language Models to Simulate Survey Response Distributions for Global Populations

    Yong Cao, Haijiang Liu, Arnav Arora +3

    cs.CLarXiv:2502.07068v22025
  50. A Systematic Literature Review of Retrieval-Augmented Generation: Techniques, Metrics, and Challenges

    Andrew Brown, Muhammad Roman, Barry Devereux

    cs.DLcs.AIcs.CLarXiv:2508.06401v32025
  51. Eagle and Finch: RWKV with Matrix-Valued States and Dynamic Recurrence

    Bo Peng, Daniel Goldstein, Quentin Anthony +27

    cs.CLcs.AIarXiv:2404.05892v42024
  52. Agent S: An Open Agentic Framework that Uses Computers Like a Human

    Saaket Agashe, Jiuzhou Han, Shuyu Gan +3

    cs.AIcs.CLcs.CVarXiv:2410.08164v12024
  53. Can LLM feedback enhance review quality? A randomized study of 20K reviews at ICLR 2025

    Nitya Thakkar, Mert Yuksekgonul, Jake Silberg +6

    cs.AIcs.CLcs.HCarXiv:2504.09737v12025
  54. Command A: An Enterprise-Ready Large Language Model

    Team Cohere, :, Aakanksha +227

    cs.CLcs.AIcs.LGarXiv:2504.00698v22025
  55. DistiLLM-2: A Contrastive Approach Boosts the Distillation of LLMs

    Jongwoo Ko, Tianyi Chen, Sungnyun Kim +4

    cs.CLcs.AIcs.LGarXiv:2503.07067v22025
  56. Offensive Language and Hate Speech Detection for Danish

    Gudbjartur Ingi Sigurbergsson, Leon Derczynski

    cs.CLarXiv:1908.04531v22019
  57. Sequence-to-Sequence Generation for Spoken Dialogue via Deep Syntax Trees and Strings

    Ondřej Dušek, Filip Jurčíček

    cs.CLarXiv:1606.05491v12016
  58. LiveCodeBench Pro: How Do Olympiad Medalists Judge LLMs in Competitive Programming?

    Zihan Zheng, Zerui Cheng, Zeyu Shen +16

    cs.SEcs.AIcs.CLarXiv:2506.11928v12025
  59. Improving Data Efficiency for LLM Reinforcement Fine-tuning Through Difficulty-targeted Online Data Selection and Rollout Replay

    Yifan Sun, Jingyan Shen, Yibin Wang +4

    cs.LGcs.AIcs.CLarXiv:2506.05316v42025
  60. Tactical Rewind: Self-Correction via Backtracking in Vision-and-Language Navigation

    Liyiming Ke, Xiujun Li, Yonatan Bisk +6

    cs.CLcs.CVcs.LGarXiv:1903.02547v22019