Computation and Language
Papers filed under cs.CL on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
3,841 to 3,900 of 11,199
MIRIX: Multi-Agent Memory System for LLM-Based Agents
Yu Wang, Xi Chen
cs.CLcs.AIarXiv:2507.07957v12025AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design
Yaxin Luo, Haobin Jiang, Jialv Zou +11
cs.CVcs.AIcs.CLarXiv:2608.13560v12026Summaries:한국어Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs
Jasper Dekoninck, Nikola Jovanović, Tim Gehrunger +4
cs.CLarXiv:2605.00674v22026Modular Cognitive Architecture Emerges in Large Language Models
Pengrui Han, Jacob Andreas, Evelina Fedorenko +1
cs.AIcs.CLcs.LGarXiv:2608.13567v12026When More is Less: Understanding Chain-of-Thought Length in LLMs
Yuyang Wu, Yifei Wang, Ziyu Ye +3
cs.AIcs.CLcs.LGarXiv:2502.07266v32025From Atomic Evidence to Logical Composition: Structured Compositional Reasoning over Compound Answer Options
Obed Junias, Maria Leonor Pacheco
cs.CLcs.AIarXiv:2608.12836v12026AVA-Encoder: Towards Agent-Native Video Representation Learning
Chuyue Li, Jinpeng Yu, Haozhe Wang +7
cs.CVcs.CLarXiv:2608.12313v12026Separating Syntax from Language: A Mechanistic Account of Translation in Multilingual LLMs
Mikhail Sonkin, Tanja Baeumel, Daniil Gurgurov +2
cs.CLarXiv:2609.01356v12026Preference Tree Optimization: Enhancing Goal-Oriented Dialogue with Look-Ahead Simulations
Lior Baruch, Moshe Butman, Kfir Bar +1
cs.CLcs.AIarXiv:2608.12062v12026ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents
Yutao Mou, Pengfei Yang, Zhe Yin +6
cs.CRcs.CLarXiv:2608.11878v12026AIMO-2 Winning Solution: Building State-of-the-Art Mathematical Reasoning Models with OpenMathReasoning dataset
Ivan Moshkov, Darragh Hanley, Ivan Sorokin +5
cs.AIcs.CLcs.LGarXiv:2504.16891v12025ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning
Bill Yuchen Lin, Ronan Le Bras, Kyle Richardson +4
cs.AIcs.CLcs.LGarXiv:2502.01100v22025InSight-doc: Agentic Visual Perception for Long-Document Understanding
Kaican Li, Weiyan Xie, Lewei Yao +4
cs.CVcs.CLcs.LGarXiv:2608.10628v12026DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments?
Mizanur Rahman, Mohammed Saidul Islam, Ridwan Mahbub +3
cs.AIcs.CLarXiv:2608.10366v12026SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring
Yuling Shi, Jinghan Xu, Kelin Fu +12
cs.CLcs.SEarXiv:2608.09802v12026Evo-Bench: Can Language Models Improve Agent Harness?
Lisheng Huang, Chen Yang, Hao Zhou +6
cs.CLarXiv:2608.09096v22026Cultivar: A Contrastive and Locale-Oriented Translation Benchmark for Investigating Contamination and Localisation Robustness
Pinzhen Chen, Koel Dutta Chowdhury, Xiaoya Xu +20
cs.CLcs.AIarXiv:2608.09766v12026LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers
Tao Feng, Fangxu Yu, Haozhen Zhang +9
cs.CLarXiv:2608.06867v12026DNA-GPT: Divergent N-Gram Analysis for Training-Free Detection of GPT-Generated Text
Xianjun Yang, Wei Cheng, Yue Wu +3
cs.CLcs.AIarXiv:2305.17359v22023RetroMAE: Pre-Training Retrieval-oriented Language Models Via Masked Auto-Encoder
Shitao Xiao, Zheng Liu, Yingxia Shao +1
cs.CLarXiv:2205.12035v22022General-Reasoner: Advancing LLM Reasoning Across All Domains
Xueguang Ma, Qian Liu, Dongfu Jiang +3
cs.CLarXiv:2505.14652v52025Delving into LLM-assisted writing in biomedical publications through excess vocabulary
Dmitry Kobak, Rita González-Márquez, Emőke-Ágnes Horvát +1
cs.CLcs.AIcs.CYarXiv:2406.07016v52024Teaching Nemotron Greek: Mining a Corpus, Adapting Retrieval, and Grounding Generation for Modern Greek across Specialist Domains
Ayoub Kirouane, Christos Petrocheilos
eess.AScs.AIcs.CLarXiv:2608.05138v12026VisualPatchWorld: Code World Models as Latent Structured Representations for Planning
Jiaxin Bai, Jiaxuan Xiong
cs.CLcs.ROarXiv:2607.25236v12026When Tokenization is Secretly Output Supervision
Tanja Baeumel, Josef van Genabith, Simon Ostermann
cs.CLarXiv:2609.01386v12026OWL: Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task Automation
Mengkang Hu, Yuhang Zhou, Wendong Fan +13
cs.AIcs.CLarXiv:2505.23885v22025ShortOPD: Recovering Pruned LLMs with Short-to-Long On-Policy Distillation
Qingyu Zhang, Qianhao Yuan, Hongyu Lin +5
cs.LGcs.AIcs.CLarXiv:2607.13124v22026Summaries:한국어Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation Models
Yubo Wang, Jiarong Liang, Yuxuan Zhang +5
cs.AIcs.CLarXiv:2607.12463v32026Classification and Clustering of Arguments with Contextualized Word Embeddings
Nils Reimers, Benjamin Schiller, Tilman Beck +3
cs.CLarXiv:1906.09821v12019Explore Before Committing: Hypothesis-Guided Search for Deep Research Agents
Ruochen Zhou, Zhengyu Chen, Luan Zhang +3
cs.CLarXiv:2609.01294v12026LLMPEDIA: Browsing, Verifying, and Comparing the Parametric Encyclopedic Knowledge of LLMs
Muhammed Saeed, Simon Razniewski
cs.CLarXiv:2609.01182v12026Accurate, Interdisciplinary and Transparent Structure-property Understanding with Deep Native Structural Reasoning
Chen Tang, Yizhou Wang, Jianyu Wu +26
cs.CLcs.AIcs.CEarXiv:2607.07708v12026Summaries:한국어What LLM Forecasters Know but Don't Say: Probing Internal Representations for Calibration and Faithfulness
Raphaël Sarfati, Pratyush Ranjan Tiwari, Siddharth Boppana +3
cs.CLcs.AIarXiv:2607.08046v12026Phrase-Localized Language-Contrastive Guidance: Training-Free Localized Accent Control for Code-Switching Text-to-Speech
Che Hyun Lee, Sangkwon Park, Donghun Kang +4
cs.CLcs.SDarXiv:2609.01016v12026PACE: A Proxy for Agentic Capability Evaluation
Yueqi Song, Lintang Sutawika, Jiarui Liu +8
cs.AIcs.CLarXiv:2607.02032v22026Logit-Contribution Scoring Identifies Non-Literal Retrieval Heads
Aryo Pradipta Gema, Beatrice Alex, Pasquale Minervini
cs.CLcs.AIcs.LGarXiv:2607.01002v12026QVal: Cheaply Evaluating Dense Supervision Signals for Long-Horizon LLM Agents
Sergio Hernández-Gutiérrez, Matteo Merler, Ilze Amanda Auzina +3
cs.LGcs.AIcs.CLarXiv:2606.32034v12026Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs
Gabrielle Kaili-May Liu, Avi Caciularu, Gal Yona +2
cs.CLcs.AIarXiv:2606.32032v12026JAMER: Project-Level Code Framework Dataset and Benchmark on Professional Game Engines
Jianwen Sun, Chuanhao Li, Zizhen Li +5
cs.SEcs.CLarXiv:2606.19830v22026Plans Don't Persist: Why Context Management Is Load Bearing for LLM Agents
Aman Mehta, Anupam Datta
cs.AIcs.CLarXiv:2606.22953v12026Reinforcement Learning from Rich Feedback with Distributional DAgger
Rishabh Agrawal, Jacob Fein-Ashley, Paria Rashidinejad
cs.LGcs.AIcs.CLarXiv:2606.05152v22026AI Research Agents Narrow Scientific Exploration
Yixuan Tang, Yi Yang
cs.CLarXiv:2605.27905v22026Efficient Agentic Reasoning Through Self-Regulated Simulative Planning
Mingkai Deng, Jinyu Hou, Lara Sá Neves +4
cs.AIcs.CLcs.LGarXiv:2605.22138v12026From Reasoning Chains to Verifiable Subproblems: Curriculum Reinforcement Learning Enables Credit Assignment for LLM Reasoning
Xitai Jiang, Zihan Tang, Wenze Lin +3
cs.LGcs.AIcs.CLarXiv:2605.22074v12026FutureSim: Replaying World Events to Evaluate Adaptive Agents
Shashwat Goel, Nikhil Chandak, Arvindh Arun +5
cs.LGcs.AIcs.CLarXiv:2605.15188v12026The Extrapolation Cliff in On-Policy Distillation of Near-Deterministic Structured Outputs
Xin Li, Hao Jiang, Annan Wang +2
cs.LGcs.CLarXiv:2605.08737v12026Going Full-TILT Boogie on Document Understanding with Text-Image-Layout Transformer
Rafał Powalski, Łukasz Borchmann, Dawid Jurkiewicz +3
cs.CLcs.LGarXiv:2102.09550v32021Automating Database-Native Function Code Synthesis with LLMs
Wei Zhou, Xuanhe Zhou, Qikang He +4
cs.DBcs.AIcs.CLarXiv:2604.06231v12026The Model Says Walk: How Surface Heuristics Override Implicit Constraints in LLM Reasoning
Yubo Li, Lu Zhang, Tianchong Jiang +2
cs.CLcs.AIarXiv:2603.29025v32026Omnilingual MT: Machine Translation for 1,600 Languages
Omnilingual MT Team, Belen Alastruey, Niyati Bafna +29
cs.CLarXiv:2603.16309v32026A Unified Mechanistic Analysis of Knowledge- and Safety-Based Refusals
Yuri Son, Seunghee Kim, Hyuhng Joon Kim +1
cs.CLarXiv:2609.00760v12026Aegis2.0: A Diverse AI Safety Dataset and Risks Taxonomy for Alignment of LLM Guardrails
Shaona Ghosh, Prasoon Varshney, Makesh Narsimhan Sreedhar +4
cs.CLarXiv:2501.09004v12025MMTEB: Massive Multilingual Text Embedding Benchmark
Kenneth Enevoldsen, Isaac Chung, Imene Kerboua +83
cs.CLcs.AIcs.IRarXiv:2502.13595v42025Context Length Alone Hurts LLM Performance Despite Perfect Retrieval
Yufeng Du, Minyang Tian, Srikanth Ronanki +7
cs.CLcs.AIarXiv:2510.05381v12025End-to-End Knowledge-Routed Relational Dialogue System for Automatic Diagnosis
Lin Xu, Qixian Zhou, Ke Gong +3
cs.CLarXiv:1901.10623v22019Motivation in Large Language Models
Omer Nahum, Asael Sklar, Ariel Goldstein +1
cs.CLcs.CYarXiv:2603.14347v12026MedMentions: A Large Biomedical Corpus Annotated with UMLS Concepts
Sunil Mohan, Donghui Li
cs.CLcs.LGarXiv:1902.09476v12019ELEPHANT: Measuring and understanding social sycophancy in LLMs
Myra Cheng, Sunny Yu, Cinoo Lee +3
cs.CLcs.AIcs.CYarXiv:2505.13995v22025WebDancer: Towards Autonomous Information Seeking Agency
Jialong Wu, Baixuan Li, Runnan Fang +10
cs.CLarXiv:2505.22648v32025Show, Control and Tell: A Framework for Generating Controllable and Grounded Captions
Marcella Cornia, Lorenzo Baraldi, Rita Cucchiara
cs.CVcs.CLarXiv:1811.10652v32018