Computation and Language

Papers filed under cs.CL on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

3,841 to 3,900 of 11,199

  1. MIRIX: Multi-Agent Memory System for LLM-Based Agents

    Yu Wang, Xi Chen

    cs.CLcs.AIarXiv:2507.07957v12025
  2. AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design

    Yaxin Luo, Haobin Jiang, Jialv Zou +11

    cs.CVcs.AIcs.CLarXiv:2608.13560v12026
    Summaries:한국어
  3. Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs

    Jasper Dekoninck, Nikola Jovanović, Tim Gehrunger +4

    cs.CLarXiv:2605.00674v22026
  4. Modular Cognitive Architecture Emerges in Large Language Models

    Pengrui Han, Jacob Andreas, Evelina Fedorenko +1

    cs.AIcs.CLcs.LGarXiv:2608.13567v12026
  5. When More is Less: Understanding Chain-of-Thought Length in LLMs

    Yuyang Wu, Yifei Wang, Ziyu Ye +3

    cs.AIcs.CLcs.LGarXiv:2502.07266v32025
  6. From Atomic Evidence to Logical Composition: Structured Compositional Reasoning over Compound Answer Options

    Obed Junias, Maria Leonor Pacheco

    cs.CLcs.AIarXiv:2608.12836v12026
  7. AVA-Encoder: Towards Agent-Native Video Representation Learning

    Chuyue Li, Jinpeng Yu, Haozhe Wang +7

    cs.CVcs.CLarXiv:2608.12313v12026
  8. Separating Syntax from Language: A Mechanistic Account of Translation in Multilingual LLMs

    Mikhail Sonkin, Tanja Baeumel, Daniil Gurgurov +2

    cs.CLarXiv:2609.01356v12026
  9. Preference Tree Optimization: Enhancing Goal-Oriented Dialogue with Look-Ahead Simulations

    Lior Baruch, Moshe Butman, Kfir Bar +1

    cs.CLcs.AIarXiv:2608.12062v12026
  10. ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents

    Yutao Mou, Pengfei Yang, Zhe Yin +6

    cs.CRcs.CLarXiv:2608.11878v12026
  11. AIMO-2 Winning Solution: Building State-of-the-Art Mathematical Reasoning Models with OpenMathReasoning dataset

    Ivan Moshkov, Darragh Hanley, Ivan Sorokin +5

    cs.AIcs.CLcs.LGarXiv:2504.16891v12025
  12. ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning

    Bill Yuchen Lin, Ronan Le Bras, Kyle Richardson +4

    cs.AIcs.CLcs.LGarXiv:2502.01100v22025
  13. InSight-doc: Agentic Visual Perception for Long-Document Understanding

    Kaican Li, Weiyan Xie, Lewei Yao +4

    cs.CVcs.CLcs.LGarXiv:2608.10628v12026
  14. DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments?

    Mizanur Rahman, Mohammed Saidul Islam, Ridwan Mahbub +3

    cs.AIcs.CLarXiv:2608.10366v12026
  15. SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring

    Yuling Shi, Jinghan Xu, Kelin Fu +12

    cs.CLcs.SEarXiv:2608.09802v12026
  16. Evo-Bench: Can Language Models Improve Agent Harness?

    Lisheng Huang, Chen Yang, Hao Zhou +6

    cs.CLarXiv:2608.09096v22026
  17. Cultivar: A Contrastive and Locale-Oriented Translation Benchmark for Investigating Contamination and Localisation Robustness

    Pinzhen Chen, Koel Dutta Chowdhury, Xiaoya Xu +20

    cs.CLcs.AIarXiv:2608.09766v12026
  18. LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers

    Tao Feng, Fangxu Yu, Haozhen Zhang +9

    cs.CLarXiv:2608.06867v12026
  19. DNA-GPT: Divergent N-Gram Analysis for Training-Free Detection of GPT-Generated Text

    Xianjun Yang, Wei Cheng, Yue Wu +3

    cs.CLcs.AIarXiv:2305.17359v22023
  20. RetroMAE: Pre-Training Retrieval-oriented Language Models Via Masked Auto-Encoder

    Shitao Xiao, Zheng Liu, Yingxia Shao +1

    cs.CLarXiv:2205.12035v22022
  21. General-Reasoner: Advancing LLM Reasoning Across All Domains

    Xueguang Ma, Qian Liu, Dongfu Jiang +3

    cs.CLarXiv:2505.14652v52025
  22. Delving into LLM-assisted writing in biomedical publications through excess vocabulary

    Dmitry Kobak, Rita González-Márquez, Emőke-Ágnes Horvát +1

    cs.CLcs.AIcs.CYarXiv:2406.07016v52024
  23. Teaching Nemotron Greek: Mining a Corpus, Adapting Retrieval, and Grounding Generation for Modern Greek across Specialist Domains

    Ayoub Kirouane, Christos Petrocheilos

    eess.AScs.AIcs.CLarXiv:2608.05138v12026
  24. VisualPatchWorld: Code World Models as Latent Structured Representations for Planning

    Jiaxin Bai, Jiaxuan Xiong

    cs.CLcs.ROarXiv:2607.25236v12026
  25. When Tokenization is Secretly Output Supervision

    Tanja Baeumel, Josef van Genabith, Simon Ostermann

    cs.CLarXiv:2609.01386v12026
  26. OWL: Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task Automation

    Mengkang Hu, Yuhang Zhou, Wendong Fan +13

    cs.AIcs.CLarXiv:2505.23885v22025
  27. ShortOPD: Recovering Pruned LLMs with Short-to-Long On-Policy Distillation

    Qingyu Zhang, Qianhao Yuan, Hongyu Lin +5

    cs.LGcs.AIcs.CLarXiv:2607.13124v22026
    Summaries:한국어
  28. Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation Models

    Yubo Wang, Jiarong Liang, Yuxuan Zhang +5

    cs.AIcs.CLarXiv:2607.12463v32026
  29. Classification and Clustering of Arguments with Contextualized Word Embeddings

    Nils Reimers, Benjamin Schiller, Tilman Beck +3

    cs.CLarXiv:1906.09821v12019
  30. Explore Before Committing: Hypothesis-Guided Search for Deep Research Agents

    Ruochen Zhou, Zhengyu Chen, Luan Zhang +3

    cs.CLarXiv:2609.01294v12026
  31. LLMPEDIA: Browsing, Verifying, and Comparing the Parametric Encyclopedic Knowledge of LLMs

    Muhammed Saeed, Simon Razniewski

    cs.CLarXiv:2609.01182v12026
  32. Accurate, Interdisciplinary and Transparent Structure-property Understanding with Deep Native Structural Reasoning

    Chen Tang, Yizhou Wang, Jianyu Wu +26

    cs.CLcs.AIcs.CEarXiv:2607.07708v12026
    Summaries:한국어
  33. What LLM Forecasters Know but Don't Say: Probing Internal Representations for Calibration and Faithfulness

    Raphaël Sarfati, Pratyush Ranjan Tiwari, Siddharth Boppana +3

    cs.CLcs.AIarXiv:2607.08046v12026
  34. Phrase-Localized Language-Contrastive Guidance: Training-Free Localized Accent Control for Code-Switching Text-to-Speech

    Che Hyun Lee, Sangkwon Park, Donghun Kang +4

    cs.CLcs.SDarXiv:2609.01016v12026
  35. PACE: A Proxy for Agentic Capability Evaluation

    Yueqi Song, Lintang Sutawika, Jiarui Liu +8

    cs.AIcs.CLarXiv:2607.02032v22026
  36. Logit-Contribution Scoring Identifies Non-Literal Retrieval Heads

    Aryo Pradipta Gema, Beatrice Alex, Pasquale Minervini

    cs.CLcs.AIcs.LGarXiv:2607.01002v12026
  37. QVal: Cheaply Evaluating Dense Supervision Signals for Long-Horizon LLM Agents

    Sergio Hernández-Gutiérrez, Matteo Merler, Ilze Amanda Auzina +3

    cs.LGcs.AIcs.CLarXiv:2606.32034v12026
  38. Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs

    Gabrielle Kaili-May Liu, Avi Caciularu, Gal Yona +2

    cs.CLcs.AIarXiv:2606.32032v12026
  39. JAMER: Project-Level Code Framework Dataset and Benchmark on Professional Game Engines

    Jianwen Sun, Chuanhao Li, Zizhen Li +5

    cs.SEcs.CLarXiv:2606.19830v22026
  40. Plans Don't Persist: Why Context Management Is Load Bearing for LLM Agents

    Aman Mehta, Anupam Datta

    cs.AIcs.CLarXiv:2606.22953v12026
  41. Reinforcement Learning from Rich Feedback with Distributional DAgger

    Rishabh Agrawal, Jacob Fein-Ashley, Paria Rashidinejad

    cs.LGcs.AIcs.CLarXiv:2606.05152v22026
  42. AI Research Agents Narrow Scientific Exploration

    Yixuan Tang, Yi Yang

    cs.CLarXiv:2605.27905v22026
  43. Efficient Agentic Reasoning Through Self-Regulated Simulative Planning

    Mingkai Deng, Jinyu Hou, Lara Sá Neves +4

    cs.AIcs.CLcs.LGarXiv:2605.22138v12026
  44. From Reasoning Chains to Verifiable Subproblems: Curriculum Reinforcement Learning Enables Credit Assignment for LLM Reasoning

    Xitai Jiang, Zihan Tang, Wenze Lin +3

    cs.LGcs.AIcs.CLarXiv:2605.22074v12026
  45. FutureSim: Replaying World Events to Evaluate Adaptive Agents

    Shashwat Goel, Nikhil Chandak, Arvindh Arun +5

    cs.LGcs.AIcs.CLarXiv:2605.15188v12026
  46. The Extrapolation Cliff in On-Policy Distillation of Near-Deterministic Structured Outputs

    Xin Li, Hao Jiang, Annan Wang +2

    cs.LGcs.CLarXiv:2605.08737v12026
  47. Going Full-TILT Boogie on Document Understanding with Text-Image-Layout Transformer

    Rafał Powalski, Łukasz Borchmann, Dawid Jurkiewicz +3

    cs.CLcs.LGarXiv:2102.09550v32021
  48. Automating Database-Native Function Code Synthesis with LLMs

    Wei Zhou, Xuanhe Zhou, Qikang He +4

    cs.DBcs.AIcs.CLarXiv:2604.06231v12026
  49. The Model Says Walk: How Surface Heuristics Override Implicit Constraints in LLM Reasoning

    Yubo Li, Lu Zhang, Tianchong Jiang +2

    cs.CLcs.AIarXiv:2603.29025v32026
  50. Omnilingual MT: Machine Translation for 1,600 Languages

    Omnilingual MT Team, Belen Alastruey, Niyati Bafna +29

    cs.CLarXiv:2603.16309v32026
  51. A Unified Mechanistic Analysis of Knowledge- and Safety-Based Refusals

    Yuri Son, Seunghee Kim, Hyuhng Joon Kim +1

    cs.CLarXiv:2609.00760v12026
  52. Aegis2.0: A Diverse AI Safety Dataset and Risks Taxonomy for Alignment of LLM Guardrails

    Shaona Ghosh, Prasoon Varshney, Makesh Narsimhan Sreedhar +4

    cs.CLarXiv:2501.09004v12025
  53. MMTEB: Massive Multilingual Text Embedding Benchmark

    Kenneth Enevoldsen, Isaac Chung, Imene Kerboua +83

    cs.CLcs.AIcs.IRarXiv:2502.13595v42025
  54. Context Length Alone Hurts LLM Performance Despite Perfect Retrieval

    Yufeng Du, Minyang Tian, Srikanth Ronanki +7

    cs.CLcs.AIarXiv:2510.05381v12025
  55. End-to-End Knowledge-Routed Relational Dialogue System for Automatic Diagnosis

    Lin Xu, Qixian Zhou, Ke Gong +3

    cs.CLarXiv:1901.10623v22019
  56. Motivation in Large Language Models

    Omer Nahum, Asael Sklar, Ariel Goldstein +1

    cs.CLcs.CYarXiv:2603.14347v12026
  57. MedMentions: A Large Biomedical Corpus Annotated with UMLS Concepts

    Sunil Mohan, Donghui Li

    cs.CLcs.LGarXiv:1902.09476v12019
  58. ELEPHANT: Measuring and understanding social sycophancy in LLMs

    Myra Cheng, Sunny Yu, Cinoo Lee +3

    cs.CLcs.AIcs.CYarXiv:2505.13995v22025
  59. WebDancer: Towards Autonomous Information Seeking Agency

    Jialong Wu, Baixuan Li, Runnan Fang +10

    cs.CLarXiv:2505.22648v32025
  60. Show, Control and Tell: A Framework for Generating Controllable and Grounded Captions

    Marcella Cornia, Lorenzo Baraldi, Rita Cucchiara

    cs.CVcs.CLarXiv:1811.10652v32018