Computation and Language
Papers filed under cs.CL on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
2,341 to 2,400 of 11,259
How Language Models Choose Sides: Internal Representations of Instruction Hierarchy
Enrique Balp-Straffon, Chih-Hao Hsu, Rushiraj Gadhvi +3
cs.AIcs.CLcs.LGarXiv:2608.28648v12026Tweet2Vec: Character-Based Distributed Representations for Social Media
Bhuwan Dhingra, Zhong Zhou, Dylan Fitzpatrick +2
cs.LGcs.CLarXiv:1605.03481v22016Coding Agents are Effective Long-Context Processors
Weili Cao, Xunjian Yin, Bhuwan Dhingra +1
cs.CLcs.AIarXiv:2603.20432v12026A parallel corpus of Python functions and documentation strings for automated code documentation and code generation
Antonio Valerio Miceli Barone, Rico Sennrich
cs.CLcs.AIarXiv:1707.02275v12017The Halt Vector: Internalizing a Causal Steering Intervention for Efficient Reasoning
Dylan Jayabahu, Tinuade Adeleke
cs.LGcs.AIcs.CLarXiv:2608.28859v12026WebXSkill: Skill Learning for Autonomous Web Agents
Zhaoyang Wang, Qianhui Wu, Xuchao Zhang +12
cs.AIcs.CLarXiv:2604.13318v22026Improving Named Entity Recognition by External Context Retrieving and Cooperative Learning
Xinyu Wang, Yong Jiang, Nguyen Bach +4
cs.CLcs.AIcs.LGarXiv:2105.03654v32021DAWN: Dependency-Aware Fast Inference for Diffusion LLMs
Lizhuo Luo, Zhuoran Shi, Jiajun Luo +4
cs.CLarXiv:2602.06953v12026Beyond the Context Window: A Cost-Performance Analysis of Fact-Based Memory vs. Long-Context LLMs for Persistent Agents
Natchanon Pollertlam, Witchayut Kornsuwannawit
cs.CLarXiv:2603.04814v12026MemoryCD: Benchmarking Long-Context User Memory of LLM Agents for Lifelong Cross-Domain Personalization
Weizhi Zhang, Xiaokai Wei, Wei-Chieh Huang +4
cs.CLarXiv:2603.25973v12026OctoBench: Benchmarking Scaffold-Aware Instruction Following in Repository-Grounded Agentic Coding
Deming Ding, Shichun Liu, Enhui Yang +12
cs.CLcs.AIarXiv:2601.10343v22026CommonsenseQA 2.0: Exposing the Limits of AI through Gamification
Alon Talmor, Ori Yoran, Ronan Le Bras +4
cs.CLcs.AIcs.LGarXiv:2201.05320v12022Graphix-T5: Mixing Pre-Trained Transformers with Graph-Aware Layers for Text-to-SQL Parsing
Jinyang Li, Binyuan Hui, Reynold Cheng +7
cs.CLcs.DBarXiv:2301.07507v12023Noise Contrastive Estimation and Negative Sampling for Conditional Models: Consistency and Statistical Efficiency
Zhuang Ma, Michael Collins
cs.CLcs.LGstat.MEarXiv:1809.01812v12018Cross Language Image Matching for Weakly Supervised Semantic Segmentation
Jinheng Xie, Xianxu Hou, Kai Ye +1
cs.CVcs.CLarXiv:2203.02668v22022Scaling Beyond Masked Diffusion Language Models
Subham Sekhar Sahoo, Jean-Marie Lemercier, Zhihan Yang +4
cs.LGcs.CLarXiv:2602.15014v12026SE-Search: Self-Evolving Search Agent via Memory and Dense Reward
Jian Li, Yizhang Jin, Dongqi Liu +9
cs.CLarXiv:2603.03293v12026Rethinking the Value of Multi-Agent Workflow: A Strong Single Agent Baseline
Jiawei Xu, Arief Koesdwiady, Sisong Bei +8
cs.MAcs.CLcs.LGarXiv:2601.12307v12026Why Reasoning Fails to Plan: A Planning-Centric Analysis of Long-Horizon Decision Making in LLM Agents
Zehong Wang, Fang Wu, Hongru Wang +8
cs.AIcs.CLcs.LGarXiv:2601.22311v12026Nüwa: Mending the Spatial Integrity Torn by VLM Token Pruning
Yihong Huang, Fei Ma, Yihua Shao +4
cs.CVcs.AIcs.CLarXiv:2602.02951v12026Prompting and Evaluating Large Language Models for Proactive Dialogues: Clarification, Target-guided, and Non-collaboration
Yang Deng, Lizi Liao, Liang Chen +3
cs.CLarXiv:2305.13626v22023MAS-Orchestra: Understanding and Improving Multi-Agent Reasoning Through Holistic Orchestration and Controlled Benchmarks
Zixuan Ke, Yifei Ming, Austin Xu +7
cs.AIcs.CLcs.MAarXiv:2601.14652v52026How much do language models copy from their training data? Evaluating linguistic novelty in text generation using RAVEN
R. Thomas McCoy, Paul Smolensky, Tal Linzen +2
cs.CLarXiv:2111.09509v12021Topic Matching in the Wild: Benchmark and Lessons from Real-World ASR Transcripts
Saman Rahbar, Xiliang Zhu, Irvin Cardoza +1
cs.CLcs.AIcs.LGarXiv:2609.00330v12026How Identity and Opinion Shape Political Sycophancy in LLMs
Li-Ni Fu, Chang-Chih Meng, Chien-Hua Chen +2
cs.AIcs.CLcs.CYarXiv:2608.29198v12026AgentHallu: Benchmarking Automated Hallucination Attribution of LLM-based Agents
Xuannan Liu, Xiao Yang, Zekun Li +2
cs.CLarXiv:2601.06818v12026End-to-End Attention based Text-Dependent Speaker Verification
Shi-Xiong Zhang, Zhuo Chen, Yong Zhao +2
cs.CLstat.MLarXiv:1701.00562v12017MemBuilder: Reinforcing LLMs for Long-Term Memory Construction via Attributed Dense Rewards
Zhiyu Shen, Ziming Wu, Fuming Lai +2
cs.CLarXiv:2601.05488v42026Attention Sensitivity Is Not Enough: Dissociating Attention-Level and Behavioural In-Context Learning under Fine-Tuning
Jinyuan Zhang, Peng He, He Hu +2
cs.LGcs.AIcs.CLarXiv:2609.00064v12026BenchGuard: Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks
Xinming Tu, Tianze Wang, Yingzhou +4
cs.CLcs.AIcs.SEarXiv:2604.24955v12026Learning to Judge: LLMs Designing and Applying Evaluation Rubrics
Clemencia Siro, Pourya Aliannejadi, Mohammad Aliannejadi
cs.CLcs.LGarXiv:2602.08672v12026Sustainable Hybrid Document-Routed Retrieval for Financial RAG: Resolving the Robustness-Precision Trade-off
Zhiyuan Cheng, Longying Lai, Yue Liu
cs.CLcs.AIcs.IRarXiv:2603.26815v32026Medically Aware GPT-3 as a Data Generator for Medical Dialogue Summarization
Bharath Chintagunta, Namit Katariya, Xavier Amatriain +1
cs.CLcs.AIcs.LGarXiv:2110.07356v12021Breaking the News: First Impressions Matter on Online News
Julio Reis, Fabrıcio Benevenuto, Pedro O. S. Vaz de Melo +3
cs.CYcs.CLarXiv:1503.07921v22015IH-Challenge: A Training Dataset to Improve Instruction Hierarchy on Frontier LLMs
Chuan Guo, Juan Felipe Ceron Uribe, Sicheng Zhu +10
cs.AIcs.CLcs.CRarXiv:2603.10521v12026HorizonBench: Long-Horizon Personalization with Evolving Preferences
Shuyue Stella Li, Bhargavi Paranjape, Kerem Oktar +9
cs.CLcs.AIarXiv:2604.17283v12026GEMBA-MQM: Detecting Translation Quality Error Spans with GPT-4
Tom Kocmi, Christian Federmann
cs.CLarXiv:2310.13988v12023Theory of Mind for Multi-Agent Collaboration via Large Language Models
Huao Li, Yu Quan Chong, Simon Stepputtis +4
cs.CLcs.AIarXiv:2310.10701v32023FinCon: A Synthesized LLM Multi-Agent System with Conceptual Verbal Reinforcement for Enhanced Financial Decision Making
Yangyang Yu, Zhiyuan Yao, Haohang Li +14
cs.CLarXiv:2407.06567v32024GPT-RE: In-context Learning for Relation Extraction using Large Language Models
Zhen Wan, Fei Cheng, Zhuoyuan Mao +4
cs.CLarXiv:2305.02105v32023Watch Out for Your Agents! Investigating Backdoor Threats to LLM-Based Agents
Wenkai Yang, Xiaohan Bi, Yankai Lin +3
cs.CRcs.AIcs.CLarXiv:2402.11208v22024Embodied-Reasoner: Synergizing Visual Search, Reasoning, and Action for Embodied Interactive Tasks
Wenqi Zhang, Mengna Wang, Gangao Liu +10
cs.CLcs.CVarXiv:2503.21696v22025APEX-Agents
Bertie Vidgen, Austin Mann, Abby Fennelly +21
cs.CLcs.AIcs.LGarXiv:2601.14242v32026Can Language Models Encode Perceptual Structure Without Grounding? A Case Study in Color
Mostafa Abdou, Artur Kulmizev, Daniel Hershcovich +3
cs.CVcs.CLarXiv:2109.06129v22021Don't Act Blindly: Robust GUI Automation via Action-Effect Verification and Self-Correction
Yuzhe Zhang, Xianwei Xue, Xingyong Wu +8
cs.CLarXiv:2604.05477v12026Counterfactual Simulation Training for Chain-of-Thought Faithfulness
Peter Hase, Christopher Potts
cs.AIcs.CLarXiv:2602.20710v22026YuFeng-XGuard: A Reasoning-Centric, Interpretable, and Flexible Guardrail Model for Large Language Models
Junyu Lin, Meizhen Liu, Xiufeng Huang +12
cs.CLarXiv:2601.15588v22026Entropy-Adaptive Fine-Tuning: Resolving Confident Conflicts to Mitigate Forgetting
Muxi Diao, Lele Yang, Wuxuan Gong +6
cs.LGcs.AIcs.CLarXiv:2601.02151v12026Learning to Extract Coherent Summary via Deep Reinforcement Learning
Yuxiang Wu, Baotian Hu
cs.CLarXiv:1804.07036v12018LLM Agents Making Agent Tools
Georg Wölflein, Dyke Ferber, Daniel Truhn +2
cs.CLcs.AIcs.LGarXiv:2502.11705v22025Like trainer, like bot? Inheritance of bias in algorithmic content moderation
Reuben Binns, Michael Veale, Max Van Kleek +1
cs.CYcs.CLcs.LGarXiv:1707.01477v12017A Mutual Information Maximization Perspective of Language Representation Learning
Lingpeng Kong, Cyprien de Masson d'Autume, Wang Ling +3
cs.CLcs.LGarXiv:1910.08350v22019A Comprehensive Study of Knowledge Editing for Large Language Models
Ningyu Zhang, Yunzhi Yao, Bozhong Tian +19
cs.CLcs.AIcs.CVarXiv:2401.01286v52024A Unified View of Attention and Residual Sinks: Outlier-Driven Rescaling is Essential for Transformer Training
Zihan Qiu, Zeyu Huang, Kaiyue Wen +16
cs.CLarXiv:2601.22966v12026Mem-T: Densifying Rewards for Long-Horizon Memory Agents
Yanwei Yue, Boci Peng, Xuanbo Fan +3
cs.LGcs.CLarXiv:2601.23014v22026CM3: A Causal Masked Multimodal Model of the Internet
Armen Aghajanyan, Bernie Huang, Candace Ross +8
cs.CLarXiv:2201.07520v12022Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety
Xingjun Ma, Yifeng Gao, Yixu Wang +45
cs.CRcs.AIcs.CLarXiv:2502.05206v62025Fine-Grained Spoiler Detection from Large-Scale Review Corpora
Mengting Wan, Rishabh Misra, Ndapa Nakashole +1
cs.CLarXiv:1905.13416v12019PRO-Step: Step-level Process Reward Optimization for Retrieval-Augmented Generation
MinKeon Kim, Namjun Lee, Jaekwang Kim
cs.CLcs.AIarXiv:2609.01658v12026Full-Duplex-Bench-v3: Benchmarking Tool Use for Full-Duplex Voice Agents Under Real-World Disfluency
Guan-Ting Lin, Chen Chen, Zhehuai Chen +1
eess.AScs.CLarXiv:2604.04847v12026