Computation and Language
Papers filed under cs.CL on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
9,121 to 9,180 of 11,260
GameXpert-Bench: How Far Are Coding Agents from Expert Game Development?
Kun Chen, Haorong Hong, Peizhong Gao +7
cs.AIcs.CLarXiv:2608.21833v12026FinMCP-Bench: Benchmarking LLM Agents for Real-World Financial Tool Use under the Model Context Protocol
Jie Zhu, Yimin Tian, Boyang Li +8
cs.AIcs.CLarXiv:2603.24943v12026A Systematic Evaluation of Large Language Models of Code
Frank F. Xu, Uri Alon, Graham Neubig +1
cs.PLcs.CLarXiv:2202.13169v32022TIES-Merging: Resolving Interference When Merging Models
Prateek Yadav, Derek Tam, Leshem Choshen +2
cs.LGcs.AIcs.CLarXiv:2306.01708v22023CCTU: A Benchmark for Tool Use under Complex Constraints
Junjie Ye, Guoqiang Zhang, Wenjie Fu +3
cs.CLcs.AIarXiv:2603.15309v12026Arcee Trinity Large Technical Report
Varun Singh, Lucas Krauss, Sami Jaghouar +23
cs.LGcs.CLarXiv:2602.17004v12026Towards Automated Kernel Generation in the Era of LLMs
Yang Yu, Peiyu Zang, Chi Hsu Tsai +11
cs.LGcs.CLarXiv:2601.15727v32026Detection and Resolution of Rumours in Social Media: A Survey
Arkaitz Zubiaga, Ahmet Aker, Kalina Bontcheva +2
cs.CLcs.HCcs.IRarXiv:1704.00656v32017FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs
Qian Chen, Jinlan Fu, Changsong Li +3
cs.CLcs.CVcs.MMarXiv:2601.13836v22026Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Melanie Sclar, Yejin Choi, Yulia Tsvetkov +1
cs.CLcs.AIcs.LGarXiv:2310.11324v22023How Close is ChatGPT to Human Experts? Comparison Corpus, Evaluation, and Detection
Biyang Guo, Xin Zhang, Ziyuan Wang +5
cs.CLarXiv:2301.07597v12023Reward-free Alignment for Conflicting Objectives
Peter Chen, Xiaopeng Li, Xi Chen +1
cs.CLcs.AIcs.LGarXiv:2602.02495v32026PersonaVLM: Long-Term Personalized Multimodal LLMs
Chang Nie, Chaoyou Fu, Yifan Zhang +2
cs.CLcs.CVarXiv:2604.13074v12026MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms
Aida Amini, Saadia Gabriel, Peter Lin +3
cs.CLarXiv:1905.13319v12019ERNIE 2.0: A Continual Pre-training Framework for Language Understanding
Yu Sun, Shuohuan Wang, Yukun Li +4
cs.CLarXiv:1907.12412v22019Which Reasoning Trajectories Teach Students to Reason Better? A Simple Metric of Informative Alignment
Yuming Yang, Mingyoung Lai, Wanxu Zhao +13
cs.CLarXiv:2601.14249v52026Thinking in Frames: How Visual Context and Test-Time Scaling Empower Video Reasoning
Chengzu Li, Zanyi Wang, Jiaang Li +9
cs.LGcs.AIcs.CLarXiv:2601.21037v12026KAPSO: A Knowledge-grounded framework for Autonomous Program Synthesis and Optimization
Alireza Nadafian, Alireza Mohammadshahi, Majid Yazdani
cs.AIcs.CLcs.SEarXiv:2601.21526v22026Residual Context Diffusion Language Models
Yuezhou Hu, Harman Singh, Monishwaran Maheswaran +10
cs.CLcs.AIarXiv:2601.22954v22026HySparse: A Hybrid Sparse Attention Architecture with Oracle Token Selection and KV Cache Sharing
Yizhao Gao, Jianyu Wei, Qihao Zhang +11
cs.CLcs.AIarXiv:2602.03560v12026Linear representations in language models can change dramatically over a conversation
Andrew Kyle Lampinen, Yuxuan Li, Eghbal Hosseini +2
cs.CLcs.LGarXiv:2601.20834v22026C-Eval: A Multi-Level Multi-Discipline Chinese Evaluation Suite for Foundation Models
Yuzhen Huang, Yuzhuo Bai, Zhihao Zhu +10
cs.CLarXiv:2305.08322v32023Mechanistic Data Attribution: Tracing the Training Origins of Interpretable LLM Units
Jianhui Chen, Yuzhang Luo, Liangming Pan
cs.CLcs.AIcs.LGarXiv:2601.21996v22026BERT-ATTACK: Adversarial Attack Against BERT Using BERT
Linyang Li, Ruotian Ma, Qipeng Guo +2
cs.CLarXiv:2004.09984v32020Effective Use of Word Order for Text Categorization with Convolutional Neural Networks
Rie Johnson, Tong Zhang
cs.CLcs.LGstat.MLarXiv:1412.1058v22014TAPAS: Weakly Supervised Table Parsing via Pre-training
Jonathan Herzig, Paweł Krzysztof Nowak, Thomas Müller +2
cs.IRcs.AIcs.CLarXiv:2004.02349v22020A Neural Network Approach to Context-Sensitive Generation of Conversational Responses
Alessandro Sordoni, Michel Galley, Michael Auli +6
cs.CLcs.AIcs.LGarXiv:1506.06714v12015Lost in the Noise: How Reasoning Models Fail with Contextual Distractors
Seongyun Lee, Yongrae Jo, Minju Seo +2
cs.AIcs.CLarXiv:2601.07226v12026OpenAssistant Conversations -- Democratizing Large Language Model Alignment
Andreas Köpf, Yannic Kilcher, Dimitri von Rütte +15
cs.CLcs.AIarXiv:2304.07327v22023Cross-Task Generalization via Natural Language Crowdsourcing Instructions
Swaroop Mishra, Daniel Khashabi, Chitta Baral +1
cs.CLcs.AIcs.CVarXiv:2104.08773v42021MSA: Memory Sparse Attention for Efficient End-to-End Memory Model Scaling to 100M Tokens
Yu Chen, Runkai Chen, Sheng Yi +9
cs.CLcs.AIcs.IRarXiv:2603.23516v22026Texygen: A Benchmarking Platform for Text Generation Models
Yaoming Zhu, Sidi Lu, Lei Zheng +4
cs.CLcs.IRcs.LGarXiv:1802.01886v12018Understanding LSTM -- a tutorial into Long Short-Term Memory Recurrent Neural Networks
Ralf C. Staudemeyer, Eric Rothstein Morris
cs.NEcs.CLcs.LGarXiv:1909.09586v12019MemoBrain: Executive Memory as an Agentic Brain for Reasoning
Hongjin Qian, Zhao Cao, Zheng Liu
cs.AIcs.CLcs.IRarXiv:2601.08079v12026Prime Agent: A Self-Improving RLM Harness
Seth Karten, Alex L. Zhang, Kevin Thomas +8
cs.AIcs.CLcs.SEarXiv:2608.23552v12026Toward Efficient Agents: Memory, Tool learning, and Planning
Xiaofang Yang, Lijun Li, Heng Zhou +12
cs.AIcs.CLarXiv:2601.14192v22026Predicting the Type and Target of Offensive Posts in Social Media
Marcos Zampieri, Shervin Malmasi, Preslav Nakov +3
cs.CLarXiv:1902.09666v22019Self-Improving World Modelling with Latent Actions
Yifu Qiu, Zheng Zhao, Waylon Li +4
cs.LGcs.AIcs.CLarXiv:2602.06130v22026Adversarial Learning for Neural Dialogue Generation
Jiwei Li, Will Monroe, Tianlin Shi +3
cs.CLarXiv:1701.06547v52017Survey of the State of the Art in Natural Language Generation: Core tasks, applications and evaluation
Albert Gatt, Emiel Krahmer
cs.CLcs.AIcs.NEarXiv:1703.09902v42017LLMs Encode Their Failures: Predicting Success from Pre-Generation Activations
William Lugoloobi, Thomas Foster, William Bankes +1
cs.CLcs.AIcs.LGarXiv:2602.09924v42026SemEval-2017 Task 4: Sentiment Analysis in Twitter
Sara Rosenthal, Noura Farra, Preslav Nakov
cs.CLcs.IRcs.LGarXiv:1912.00741v12019Crowdsourcing Multiple Choice Science Questions
Johannes Welbl, Nelson F. Liu, Matt Gardner
cs.HCcs.AIcs.CLarXiv:1707.06209v12017Length-Unbiased Sequence Policy Optimization: Revealing and Controlling Response Length Variation in RLVR
Fanfan Liu, Youyang Yin, Peng Shi +3
cs.CLarXiv:2602.05261v12026Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
Yann Dubois, Balázs Galambosi, Percy Liang +1
cs.LGcs.AIcs.CLarXiv:2404.04475v22024Same Agent, Different Answers: A Repeat-Aware Audit of Corpus-Induced Answer Churn in Retrieval-Augmented QA
Jingjie Ning, Xueqi Li
cs.IRcs.CLarXiv:2608.22856v12026Temporal Pattern Attention for Multivariate Time Series Forecasting
Shun-Yao Shih, Fan-Keng Sun, Hung-yi Lee
cs.LGcs.CLstat.MLarXiv:1809.04206v32018Empty Shelves or Lost Keys? Recall Is the Bottleneck for Parametric Factuality
Nitay Calderon, Eyal Ben-David, Zorik Gekhman +2
cs.CLcs.AIarXiv:2602.14080v22026Instruction Tuning for Large Language Models: A Survey
Shengyu Zhang, Linfeng Dong, Xiaoya Li +8
cs.CLcs.AIcs.LGarXiv:2308.10792v102023UMEM: Unified Memory Extraction and Management Framework for Generalizable Memory
Yongshi Ye, Hui Jiang, Feihu Jiang +7
cs.CLarXiv:2602.10652v12026MatchTIR: Fine-Grained Supervision for Tool-Integrated Reasoning via Bipartite Matching
Changle Qu, Sunhao Dai, Hengyi Cai +3
cs.CLcs.AIarXiv:2601.10712v12026Agentic-R: Learning to Retrieve for Agentic Search
Wenhan Liu, Xinyu Ma, Yutao Zhu +4
cs.IRcs.CLarXiv:2601.11888v12026Explainability for Large Language Models: A Survey
Haiyan Zhao, Hanjie Chen, Fan Yang +6
cs.CLcs.AIcs.LGarXiv:2309.01029v32023LMEB: Long-horizon Memory Embedding Benchmark
Xinping Zhao, Xinshuo Hu, Jiaxin Xu +9
cs.CLarXiv:2603.12572v62026InCoder: A Generative Model for Code Infilling and Synthesis
Daniel Fried, Armen Aghajanyan, Jessy Lin +7
cs.SEcs.CLcs.LGarXiv:2204.05999v32022X-Coder: Advancing Competitive Programming with Fully Synthetic Tasks, Solutions, and Tests
Jie Wu, Haoling Li, Xin Zhang +7
cs.CLcs.LGarXiv:2601.06953v22026Apodex 1.1: Scaling Agentic Intelligence for Complex Work
Apodex Team, B. An, B. Li +68
cs.AIcs.CLcs.LGarXiv:2608.23283v12026Sentiment of Emojis
Petra Kralj Novak, Jasmina Smailović, Borut Sluban +1
cs.CLarXiv:1509.07761v22015How to Fine-Tune a Reasoning Model? A Teacher-Student Cooperation Framework to Synthesize Student-Consistent SFT Data
Zixian Huang, Kaichen Yang, Xu Huang +6
cs.CLarXiv:2604.14164v22026Reasoning Models Generate Societies of Thought
Junsol Kim, Shiyang Lai, Nino Scherrer +2
cs.CLcs.CYcs.LGarXiv:2601.10825v12026