Computation and Language

Papers filed under cs.CL on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

9,121 to 9,180 of 11,260

  1. GameXpert-Bench: How Far Are Coding Agents from Expert Game Development?

    Kun Chen, Haorong Hong, Peizhong Gao +7

    cs.AIcs.CLarXiv:2608.21833v12026
  2. FinMCP-Bench: Benchmarking LLM Agents for Real-World Financial Tool Use under the Model Context Protocol

    Jie Zhu, Yimin Tian, Boyang Li +8

    cs.AIcs.CLarXiv:2603.24943v12026
  3. A Systematic Evaluation of Large Language Models of Code

    Frank F. Xu, Uri Alon, Graham Neubig +1

    cs.PLcs.CLarXiv:2202.13169v32022
  4. TIES-Merging: Resolving Interference When Merging Models

    Prateek Yadav, Derek Tam, Leshem Choshen +2

    cs.LGcs.AIcs.CLarXiv:2306.01708v22023
  5. CCTU: A Benchmark for Tool Use under Complex Constraints

    Junjie Ye, Guoqiang Zhang, Wenjie Fu +3

    cs.CLcs.AIarXiv:2603.15309v12026
  6. Arcee Trinity Large Technical Report

    Varun Singh, Lucas Krauss, Sami Jaghouar +23

    cs.LGcs.CLarXiv:2602.17004v12026
  7. Towards Automated Kernel Generation in the Era of LLMs

    Yang Yu, Peiyu Zang, Chi Hsu Tsai +11

    cs.LGcs.CLarXiv:2601.15727v32026
  8. Detection and Resolution of Rumours in Social Media: A Survey

    Arkaitz Zubiaga, Ahmet Aker, Kalina Bontcheva +2

    cs.CLcs.HCcs.IRarXiv:1704.00656v32017
  9. FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs

    Qian Chen, Jinlan Fu, Changsong Li +3

    cs.CLcs.CVcs.MMarXiv:2601.13836v22026
  10. Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

    Melanie Sclar, Yejin Choi, Yulia Tsvetkov +1

    cs.CLcs.AIcs.LGarXiv:2310.11324v22023
  11. How Close is ChatGPT to Human Experts? Comparison Corpus, Evaluation, and Detection

    Biyang Guo, Xin Zhang, Ziyuan Wang +5

    cs.CLarXiv:2301.07597v12023
  12. Reward-free Alignment for Conflicting Objectives

    Peter Chen, Xiaopeng Li, Xi Chen +1

    cs.CLcs.AIcs.LGarXiv:2602.02495v32026
  13. PersonaVLM: Long-Term Personalized Multimodal LLMs

    Chang Nie, Chaoyou Fu, Yifan Zhang +2

    cs.CLcs.CVarXiv:2604.13074v12026
  14. MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms

    Aida Amini, Saadia Gabriel, Peter Lin +3

    cs.CLarXiv:1905.13319v12019
  15. ERNIE 2.0: A Continual Pre-training Framework for Language Understanding

    Yu Sun, Shuohuan Wang, Yukun Li +4

    cs.CLarXiv:1907.12412v22019
  16. Which Reasoning Trajectories Teach Students to Reason Better? A Simple Metric of Informative Alignment

    Yuming Yang, Mingyoung Lai, Wanxu Zhao +13

    cs.CLarXiv:2601.14249v52026
  17. Thinking in Frames: How Visual Context and Test-Time Scaling Empower Video Reasoning

    Chengzu Li, Zanyi Wang, Jiaang Li +9

    cs.LGcs.AIcs.CLarXiv:2601.21037v12026
  18. KAPSO: A Knowledge-grounded framework for Autonomous Program Synthesis and Optimization

    Alireza Nadafian, Alireza Mohammadshahi, Majid Yazdani

    cs.AIcs.CLcs.SEarXiv:2601.21526v22026
  19. Residual Context Diffusion Language Models

    Yuezhou Hu, Harman Singh, Monishwaran Maheswaran +10

    cs.CLcs.AIarXiv:2601.22954v22026
  20. HySparse: A Hybrid Sparse Attention Architecture with Oracle Token Selection and KV Cache Sharing

    Yizhao Gao, Jianyu Wei, Qihao Zhang +11

    cs.CLcs.AIarXiv:2602.03560v12026
  21. Linear representations in language models can change dramatically over a conversation

    Andrew Kyle Lampinen, Yuxuan Li, Eghbal Hosseini +2

    cs.CLcs.LGarXiv:2601.20834v22026
  22. C-Eval: A Multi-Level Multi-Discipline Chinese Evaluation Suite for Foundation Models

    Yuzhen Huang, Yuzhuo Bai, Zhihao Zhu +10

    cs.CLarXiv:2305.08322v32023
  23. Mechanistic Data Attribution: Tracing the Training Origins of Interpretable LLM Units

    Jianhui Chen, Yuzhang Luo, Liangming Pan

    cs.CLcs.AIcs.LGarXiv:2601.21996v22026
  24. BERT-ATTACK: Adversarial Attack Against BERT Using BERT

    Linyang Li, Ruotian Ma, Qipeng Guo +2

    cs.CLarXiv:2004.09984v32020
  25. Effective Use of Word Order for Text Categorization with Convolutional Neural Networks

    Rie Johnson, Tong Zhang

    cs.CLcs.LGstat.MLarXiv:1412.1058v22014
  26. TAPAS: Weakly Supervised Table Parsing via Pre-training

    Jonathan Herzig, Paweł Krzysztof Nowak, Thomas Müller +2

    cs.IRcs.AIcs.CLarXiv:2004.02349v22020
  27. A Neural Network Approach to Context-Sensitive Generation of Conversational Responses

    Alessandro Sordoni, Michel Galley, Michael Auli +6

    cs.CLcs.AIcs.LGarXiv:1506.06714v12015
  28. Lost in the Noise: How Reasoning Models Fail with Contextual Distractors

    Seongyun Lee, Yongrae Jo, Minju Seo +2

    cs.AIcs.CLarXiv:2601.07226v12026
  29. OpenAssistant Conversations -- Democratizing Large Language Model Alignment

    Andreas Köpf, Yannic Kilcher, Dimitri von Rütte +15

    cs.CLcs.AIarXiv:2304.07327v22023
  30. Cross-Task Generalization via Natural Language Crowdsourcing Instructions

    Swaroop Mishra, Daniel Khashabi, Chitta Baral +1

    cs.CLcs.AIcs.CVarXiv:2104.08773v42021
  31. MSA: Memory Sparse Attention for Efficient End-to-End Memory Model Scaling to 100M Tokens

    Yu Chen, Runkai Chen, Sheng Yi +9

    cs.CLcs.AIcs.IRarXiv:2603.23516v22026
  32. Texygen: A Benchmarking Platform for Text Generation Models

    Yaoming Zhu, Sidi Lu, Lei Zheng +4

    cs.CLcs.IRcs.LGarXiv:1802.01886v12018
  33. Understanding LSTM -- a tutorial into Long Short-Term Memory Recurrent Neural Networks

    Ralf C. Staudemeyer, Eric Rothstein Morris

    cs.NEcs.CLcs.LGarXiv:1909.09586v12019
  34. MemoBrain: Executive Memory as an Agentic Brain for Reasoning

    Hongjin Qian, Zhao Cao, Zheng Liu

    cs.AIcs.CLcs.IRarXiv:2601.08079v12026
  35. Prime Agent: A Self-Improving RLM Harness

    Seth Karten, Alex L. Zhang, Kevin Thomas +8

    cs.AIcs.CLcs.SEarXiv:2608.23552v12026
  36. Toward Efficient Agents: Memory, Tool learning, and Planning

    Xiaofang Yang, Lijun Li, Heng Zhou +12

    cs.AIcs.CLarXiv:2601.14192v22026
  37. Predicting the Type and Target of Offensive Posts in Social Media

    Marcos Zampieri, Shervin Malmasi, Preslav Nakov +3

    cs.CLarXiv:1902.09666v22019
  38. Self-Improving World Modelling with Latent Actions

    Yifu Qiu, Zheng Zhao, Waylon Li +4

    cs.LGcs.AIcs.CLarXiv:2602.06130v22026
  39. Adversarial Learning for Neural Dialogue Generation

    Jiwei Li, Will Monroe, Tianlin Shi +3

    cs.CLarXiv:1701.06547v52017
  40. Survey of the State of the Art in Natural Language Generation: Core tasks, applications and evaluation

    Albert Gatt, Emiel Krahmer

    cs.CLcs.AIcs.NEarXiv:1703.09902v42017
  41. LLMs Encode Their Failures: Predicting Success from Pre-Generation Activations

    William Lugoloobi, Thomas Foster, William Bankes +1

    cs.CLcs.AIcs.LGarXiv:2602.09924v42026
  42. SemEval-2017 Task 4: Sentiment Analysis in Twitter

    Sara Rosenthal, Noura Farra, Preslav Nakov

    cs.CLcs.IRcs.LGarXiv:1912.00741v12019
  43. Crowdsourcing Multiple Choice Science Questions

    Johannes Welbl, Nelson F. Liu, Matt Gardner

    cs.HCcs.AIcs.CLarXiv:1707.06209v12017
  44. Length-Unbiased Sequence Policy Optimization: Revealing and Controlling Response Length Variation in RLVR

    Fanfan Liu, Youyang Yin, Peng Shi +3

    cs.CLarXiv:2602.05261v12026
  45. Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

    Yann Dubois, Balázs Galambosi, Percy Liang +1

    cs.LGcs.AIcs.CLarXiv:2404.04475v22024
  46. Same Agent, Different Answers: A Repeat-Aware Audit of Corpus-Induced Answer Churn in Retrieval-Augmented QA

    Jingjie Ning, Xueqi Li

    cs.IRcs.CLarXiv:2608.22856v12026
  47. Temporal Pattern Attention for Multivariate Time Series Forecasting

    Shun-Yao Shih, Fan-Keng Sun, Hung-yi Lee

    cs.LGcs.CLstat.MLarXiv:1809.04206v32018
  48. Empty Shelves or Lost Keys? Recall Is the Bottleneck for Parametric Factuality

    Nitay Calderon, Eyal Ben-David, Zorik Gekhman +2

    cs.CLcs.AIarXiv:2602.14080v22026
  49. Instruction Tuning for Large Language Models: A Survey

    Shengyu Zhang, Linfeng Dong, Xiaoya Li +8

    cs.CLcs.AIcs.LGarXiv:2308.10792v102023
  50. UMEM: Unified Memory Extraction and Management Framework for Generalizable Memory

    Yongshi Ye, Hui Jiang, Feihu Jiang +7

    cs.CLarXiv:2602.10652v12026
  51. MatchTIR: Fine-Grained Supervision for Tool-Integrated Reasoning via Bipartite Matching

    Changle Qu, Sunhao Dai, Hengyi Cai +3

    cs.CLcs.AIarXiv:2601.10712v12026
  52. Agentic-R: Learning to Retrieve for Agentic Search

    Wenhan Liu, Xinyu Ma, Yutao Zhu +4

    cs.IRcs.CLarXiv:2601.11888v12026
  53. Explainability for Large Language Models: A Survey

    Haiyan Zhao, Hanjie Chen, Fan Yang +6

    cs.CLcs.AIcs.LGarXiv:2309.01029v32023
  54. LMEB: Long-horizon Memory Embedding Benchmark

    Xinping Zhao, Xinshuo Hu, Jiaxin Xu +9

    cs.CLarXiv:2603.12572v62026
  55. InCoder: A Generative Model for Code Infilling and Synthesis

    Daniel Fried, Armen Aghajanyan, Jessy Lin +7

    cs.SEcs.CLcs.LGarXiv:2204.05999v32022
  56. X-Coder: Advancing Competitive Programming with Fully Synthetic Tasks, Solutions, and Tests

    Jie Wu, Haoling Li, Xin Zhang +7

    cs.CLcs.LGarXiv:2601.06953v22026
  57. Apodex 1.1: Scaling Agentic Intelligence for Complex Work

    Apodex Team, B. An, B. Li +68

    cs.AIcs.CLcs.LGarXiv:2608.23283v12026
  58. Sentiment of Emojis

    Petra Kralj Novak, Jasmina Smailović, Borut Sluban +1

    cs.CLarXiv:1509.07761v22015
  59. How to Fine-Tune a Reasoning Model? A Teacher-Student Cooperation Framework to Synthesize Student-Consistent SFT Data

    Zixian Huang, Kaichen Yang, Xu Huang +6

    cs.CLarXiv:2604.14164v22026
  60. Reasoning Models Generate Societies of Thought

    Junsol Kim, Shiyang Lai, Nino Scherrer +2

    cs.CLcs.CYcs.LGarXiv:2601.10825v12026