Computation and Language
Papers filed under cs.CL on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
3,301 to 3,360 of 11,253
Aristotle: IMO-level Automated Theorem Proving
Tudor Achim, Alex Best, Alberto Bietti +20
cs.AIcs.CLarXiv:2510.01346v22025dParallel: Learnable Parallel Decoding for dLLMs
Zigeng Chen, Gongfan Fang, Xinyin Ma +2
cs.CLarXiv:2509.26488v12025Vision Is Not Overhead: One-Pass Block Drafting for Lossless Speculative Decoding in Vision-Language Models
Jungseob Lee, Seongtae Hong, Dongyub Jude Lee +4
cs.AIcs.CLcs.CVarXiv:2609.00355v12026A Survey of Word Embeddings Evaluation Methods
Amir Bakarov
cs.CLarXiv:1801.09536v12018DiTAR: Diffusion Transformer Autoregressive Modeling for Speech Generation
Dongya Jia, Zhuo Chen, Jiawei Chen +8
eess.AScs.AIcs.CLarXiv:2502.03930v42025Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL
Weizhen Li, Jianbo Lin, Zhuosong Jiang +27
cs.AIcs.CLarXiv:2508.13167v12025SafeArena: Evaluating the Safety of Autonomous Web Agents
Ada Defne Tur, Nicholas Meade, Xing Han Lù +6
cs.LGcs.AIcs.CLarXiv:2503.04957v12025Token Assorted: Mixing Latent and Text Tokens for Improved Language Model Reasoning
DiJia Su, Hanlin Zhu, Yingchen Xu +3
cs.CLcs.AIcs.LGarXiv:2502.03275v22025AdaEvolve: Adaptive LLM Driven Zeroth-Order Optimization
Mert Cemri, Shubham Agrawal, Akshat Gupta +9
cs.NEcs.AIcs.CLarXiv:2602.20133v12026Vision-Language Models Do Not Understand Negation
Kumail Alhamoud, Shaden Alshammari, Yonglong Tian +4
cs.CVcs.CLarXiv:2501.09425v22025MMC: Advancing Multimodal Chart Understanding with Large-scale Instruction Tuning
Fuxiao Liu, Xiaoyang Wang, Wenlin Yao +5
cs.CLcs.AIarXiv:2311.10774v22023Investigating Cultural Alignment of Large Language Models
Badr AlKhamissi, Muhammad ElNokrashy, Mai AlKhamissi +1
cs.CLcs.CYarXiv:2402.13231v22024LLaMA-Omni2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis
Qingkai Fang, Yan Zhou, Shoutao Guo +2
cs.CLcs.AIcs.SDarXiv:2505.02625v12025Same Semantics, Different Outcome: On the Modality Robustness of Multimodal LLMs under Knowledge Conflict
Jungyeon Lee, Yejin Yoon, Taeuk Kim
cs.CLarXiv:2609.00550v12026Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning
Vaishnavi Shrivastava, Ahmed Awadallah, Vidhisha Balachandran +3
cs.CLcs.LGarXiv:2508.09726v12025AgentEvolver: Towards Efficient Self-Evolving Agent System
Yunpeng Zhai, Shuchang Tao, Cheng Chen +10
cs.LGcs.AIcs.CLarXiv:2511.10395v12025Neural Semantic Role Labeling with Dependency Path Embeddings
Michael Roth, Mirella Lapata
cs.CLarXiv:1605.07515v22016A Sober Look at Progress in Language Model Reasoning: Pitfalls and Paths to Reproducibility
Andreas Hochlehnert, Hardik Bhatnagar, Vishaal Udandarao +3
cs.LGcs.CLarXiv:2504.07086v22025Multi-Agent Design: Optimizing Agents with Better Prompts and Topologies
Han Zhou, Xingchen Wan, Ruoxi Sun +5
cs.LGcs.AIcs.CLarXiv:2502.02533v22025MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation
Weihao Xuan, Rui Yang, Heli Qi +29
cs.CLarXiv:2503.10497v22025Why Is Spatial Reasoning Hard for VLMs? An Attention Mechanism Perspective on Focus Areas
Shiqi Chen, Tongyao Zhu, Ruochen Zhou +7
cs.CLarXiv:2503.01773v32025MLGym: A New Framework and Benchmark for Advancing AI Research Agents
Deepak Nathani, Lovish Madaan, Nicholas Roberts +14
cs.CLcs.AIcs.LGarXiv:2502.14499v12025Inducing Programmatic Skills for Agentic Tasks
Zora Zhiruo Wang, Apurva Gandhi, Graham Neubig +1
cs.CLarXiv:2504.06821v22025When Modality Gap Reduction Fails: Prediction-Level Hubness in CLIP
Shota Sato, Hajime Kiyama, Tosho Hirasawa +1
cs.CLcs.CVarXiv:2609.01103v12026ReCLIP: A Strong Zero-Shot Baseline for Referring Expression Comprehension
Sanjay Subramanian, William Merrill, Trevor Darrell +3
cs.CVcs.CLarXiv:2204.05991v22022Text Readability Assessment for Second Language Learners
Menglin Xia, Ekaterina Kochmar, Ted Briscoe
cs.CLarXiv:1906.07580v12019How to (Properly) Evaluate Cross-Lingual Word Embeddings: On Strong Baselines, Comparative Analyses, and Some Misconceptions
Goran Glavas, Robert Litschko, Sebastian Ruder +1
cs.CLarXiv:1902.00508v12019Improving Multi-Task Deep Neural Networks via Knowledge Distillation for Natural Language Understanding
Xiaodong Liu, Pengcheng He, Weizhu Chen +1
cs.CLarXiv:1904.09482v12019The Geometry of Refusal in Large Language Models: Concept Cones and Representational Independence
Tom Wollschläger, Jannes Elstner, Simon Geisler +3
cs.LGcs.AIcs.CLarXiv:2502.17420v22025Memp: Exploring Agent Procedural Memory
Runnan Fang, Yuan Liang, Xiaobin Wang +6
cs.CLcs.AIcs.LGarXiv:2508.06433v42025Perception-R1: Pioneering Perception Policy with Reinforcement Learning
En Yu, Kangheng Lin, Liang Zhao +11
cs.CVcs.CLarXiv:2504.07954v12025Multiagent Finetuning: Self Improvement with Diverse Reasoning Chains
Vighnesh Subramaniam, Yilun Du, Joshua B. Tenenbaum +3
cs.CLcs.AIcs.LGarXiv:2501.05707v22025Beneath the Diff: Diagnosing and Mitigating Algorithmic Mode Collapse in Code-Level Autonomous Research Loops
Bowei He, Weixu Zhang, Yili Jin +1
cs.CLcs.SEarXiv:2609.00077v12026NSIDDx: A Design Framework for Neuro-Symbolic, Practitioner-First Differential Diagnosis in Low-Resource Settings
Aarav Singh
cs.CLarXiv:2609.00256v12026Deep Video Discovery: Agentic Search with Tool Use for Long-form Video Understanding
Xiaoyi Zhang, Zhaoyang Jia, Zongyu Guo +4
cs.CVcs.AIcs.CLarXiv:2505.18079v42025Towards Thinking-Optimal Scaling of Test-Time Compute for LLM Reasoning
Wenkai Yang, Shuming Ma, Yankai Lin +1
cs.CLcs.AIarXiv:2502.18080v22025Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation
Sayash Kapoor, Benedikt Stroebl, Peter Kirgis +28
cs.AIcs.CLarXiv:2510.11977v12025LLM Generated Persona is a Promise with a Catch
Ang Li, Haozhe Chen, Hongseok Namkoong +1
cs.CLcs.AIcs.CYarXiv:2503.16527v12025Quasar: Datasets for Question Answering by Search and Reading
Bhuwan Dhingra, Kathryn Mazaitis, William W. Cohen
cs.CLcs.IRcs.LGarXiv:1707.03904v22017MTRAG: A Multi-Turn Conversational Benchmark for Evaluating Retrieval-Augmented Generation Systems
Yannis Katsis, Sara Rosenthal, Kshitij Fadnis +7
cs.CLcs.AIarXiv:2501.03468v12025NoLiMa: Long-Context Evaluation Beyond Literal Matching
Ali Modarressi, Hanieh Deilamsalehy, Franck Dernoncourt +4
cs.CLarXiv:2502.05167v32025Agentic Reasoning: A Streamlined Framework for Enhancing LLM Reasoning with Agentic Tools
Junde Wu, Jiayuan Zhu, Yuyuan Liu +2
cs.AIcs.CLarXiv:2502.04644v22025Rank1: Test-Time Compute for Reranking in Information Retrieval
Orion Weller, Kathryn Ricci, Eugene Yang +3
cs.IRcs.CLcs.LGarXiv:2502.18418v22025EditReward: A Human-Aligned Reward Model for Instruction-Guided Image Editing
Keming Wu, Sicong Jiang, Max Ku +3
cs.CVcs.AIcs.CLarXiv:2509.26346v22025jina-embeddings-v3: Multilingual Embeddings With Task LoRA
Saba Sturua, Isabelle Mohr, Mohammad Kalim Akram +8
cs.CLcs.AIcs.IRarXiv:2409.10173v32024Supply-Chain Poisoning Attacks Against LLM Coding Agent Skill Ecosystems
Yubin Qu, Yi Liu, Tongcheng Geng +5
cs.CRcs.AIcs.CLarXiv:2604.03081v12026An LLM Compiler for Parallel Function Calling
Sehoon Kim, Suhong Moon, Ryan Tabrizi +4
cs.CLarXiv:2312.04511v32023Agentic AI for Scientific Discovery: A Survey of Progress, Challenges, and Future Directions
Mourad Gridach, Jay Nanavati, Khaldoun Zine El Abidine +2
cs.CLarXiv:2503.08979v12025A Web of Hate: Tackling Hateful Speech in Online Social Spaces
Haji Mohammad Saleem, Kelly P Dillon, Susan Benesch +1
cs.CLarXiv:1709.10159v12017Composite Task-Completion Dialogue Policy Learning via Hierarchical Deep Reinforcement Learning
Baolin Peng, Xiujun Li, Lihong Li +4
cs.CLcs.AIcs.LGarXiv:1704.03084v32017CausaLM: Causal Model Explanation Through Counterfactual Language Models
Amir Feder, Nadav Oved, Uri Shalit +1
cs.CLcs.AIcs.LGarXiv:2005.13407v52020The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution
Junlong Li, Wenshuo Zhao, Jian Zhao +18
cs.CLcs.AIarXiv:2510.25726v22025Continual Learning for Large Language Models: A Survey
Tongtong Wu, Linhao Luo, Yuan-Fang Li +3
cs.CLcs.LGarXiv:2402.01364v22024Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models
Alon Albalak, Duy Phung, Nathan Lile +8
cs.LGcs.AIcs.CLarXiv:2502.17387v12025SpatialLadder: Progressive Training for Spatial Reasoning in Vision-Language Models
Hongxing Li, Dingming Li, Zixuan Wang +7
cs.CVcs.AIcs.CLarXiv:2510.08531v12025WMT24++: Expanding the Language Coverage of WMT24 to 55 Languages & Dialects
Daniel Deutsch, Eleftheria Briakou, Isaac Caswell +14
cs.CLarXiv:2502.12404v12025SPICE: Self-Play In Corpus Environments Improves Reasoning
Bo Liu, Chuanyang Jin, Seungone Kim +7
cs.CLarXiv:2510.24684v12025TWIX: a Two-Stage Approach for End-To-End Named Entity Recognition and Relation Extraction
Marco Martinelli, Laura Menotti
cs.CLarXiv:2609.00832v12026DeepSeekMath-V2: Towards Self-Verifiable Mathematical Reasoning
Zhihong Shao, Yuxiang Luo, Chengda Lu +6
cs.AIcs.CLarXiv:2511.22570v12025MVAN: Multi-View Attention Networks for Fake News Detection on Social Media
Shiwen Ni, Jiawen Li, Hung-Yu Kao
cs.CLarXiv:2506.01627v12025