Computation and Language

Papers filed under cs.CL on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

3,721 to 3,780 of 11,333

  1. GeoAgent: Evaluating VLM Geolocalization Through Embodied Navigation

    Arka Mukherjee, Soham Roy, Kartikeya Trivedi +1

    cs.CVcs.CLarXiv:2608.29483v12026
  2. Agentic Large Language Models, a survey

    Aske Plaat, Max van Duijn, Niki van Stein +3

    cs.AIcs.CLcs.LGarXiv:2503.23037v32025
  3. Twin Worlds: Equivariance-Based Abstention for Evidence-Grounded Reasoning

    Vy Nguyen, Ziqi Xu, Jeffrey Chan +5

    cs.CLcs.AIcs.LGarXiv:2608.28018v12026
  4. TEMPLAR Wales: A georeferenced environmental and toponymic dataset of Welsh settlements

    Oktay Karakuş, Can Eyupoglu

    cs.LGcs.CLarXiv:2608.26970v12026
  5. Why Didn't It Check? Unsupported Final Claims and Their Repair in Two Tool-Equipped Language Models

    Justin Bronder

    cs.AIcs.CLarXiv:2608.27768v12026
  6. Cluster, Route, Escalate: Cascaded Framework for Cost-Aware LLM Serving

    Yasmin Moslem, Magdalena Kacmajor, Vasudevan Nedumpozhimana +11

    cs.PFcs.CLarXiv:2606.27457v12026
  7. MemRL: Self-Evolving Agents via Runtime Reinforcement Learning on Episodic Memory

    Shengtao Zhang, Jiaqian Wang, Ruiwen Zhou +11

    cs.CLarXiv:2601.03192v22026
  8. Dependency-Aware Revocable Decoding for Efficient Diffusion Large Language Model Inference

    Wooje Park, Insu Lee, Minyoung Noh +4

    cs.CLarXiv:2608.26574v12026
  9. Thoughts Are All Over the Place: On the Underthinking of o1-Like LLMs

    Yue Wang, Qiuzhi Liu, Jiahao Xu +11

    cs.CLarXiv:2501.18585v22025
  10. PACEShop: Evaluating Personalized, Actionable, Compositional, and Evidence-grounded Shopping Assistants

    Weimin Lyu, Chen Luo, Guangrui Li +9

    cs.CLcs.AIarXiv:2608.26180v12026
  11. On Scope Classification and Current Knowledge-Editing Benchmarks: A Negative Result, with INLAY as a Gradient-Free Case Study

    Aditya Pratap Singh

    cs.CLcs.AIcs.LGarXiv:2608.26292v12026
  12. Which India Survives Translation? Narrative Homogenisation Across Indian Oral Traditions in LLMs

    Paarth Singh Rathore

    cs.CLarXiv:2608.26123v12026
  13. Learning Language Games through Interaction

    Sida I. Wang, Percy Liang, Christopher D. Manning

    cs.CLcs.AIarXiv:1606.02447v12016
  14. IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

    Siyi Zhou, Yiquan Zhou, Yi He +4

    cs.CLcs.AIcs.SDarXiv:2506.21619v22025
  15. The 'Problem' of Human Label Variation: On Ground Truth in Data, Modeling and Evaluation

    Barbara Plank

    cs.CLcs.LGarXiv:2211.02570v12022
  16. The Multilingual FrameNet Corpus

    Beatrice Fiumanò, Nicolas Lazzari, Simone Paolo Ponzetto +1

    cs.CLcs.AIarXiv:2608.23037v12026
  17. Improving Few-Step Language Flows with Untied Self-Conditioning

    Bocheng Li, Linli Xu

    cs.CLcs.AIcs.LGarXiv:2608.22244v12026
  18. RpBERT: A Text-image Relation Propagation-based BERT Model for Multimodal NER

    Lin Sun, Jiquan Wang, Kai Zhang +2

    cs.CLcs.LGarXiv:2102.02967v12021
  19. Scaling Unsupervised Word Alignment to Documents via Structural Constraints

    Michelle Wastl, Jannis Vamvas, Rico Sennrich

    cs.CLarXiv:2608.21023v12026
  20. TURINGBENCH: A Benchmark Environment for Turing Test in the Age of Neural Text Generation

    Adaku Uchendu, Zeyu Ma, Thai Le +2

    cs.CLarXiv:2109.13296v12021
  21. Towards Facilitating Empathic Conversations in Online Mental Health Support: A Reinforcement Learning Approach

    Ashish Sharma, Inna W. Lin, Adam S. Miner +2

    cs.CLcs.SIarXiv:2101.07714v32021
  22. Beyond Recall: Behavioral Specification as an Interpretive Layer for AI Personalization

    Aarik Gulaya

    cs.CLcs.AIcs.HCarXiv:2605.28969v22026
  23. Directional Contextual Representations for Dependency Relations: Why Cross-Direction Pairing Fails

    Sai Krishna Arthanari, JaeHyeong Chang, Chengzhe Sun +1

    cs.CLarXiv:2608.20647v12026
  24. PRISM: A Multi-Dimensional Benchmark for Evaluating LLM Peer Reviewers

    Ngoc Phan Phuoc Loc, Toan Huynh La Viet, Thanh Tran Khanh +8

    cs.CLarXiv:2605.26730v22026
  25. Institutional Newspapers Pipeline: Deriving billions of high quality tokens from historical newspapers

    Matteo Cargnelutti, Catherine Brobston, Eben English +6

    cs.CLcs.DLarXiv:2608.18972v12026
  26. FastKernels: Benchmarking GPU Kernel Generation in Production

    Gabriele Oliaro, Yichao Fu, May Jiang +5

    cs.LGcs.AIcs.CLarXiv:2605.23215v12026
  27. From Storage to Access: Verifiable Activation of Parametric Knowledge in LLMs via Explicit Priming and Implicit Reasoning

    Zuocheng Ying, Yang Yang, Yumou Wu +6

    cs.CLcs.AIcs.LGarXiv:2608.18581v12026
  28. Compress and Forget: bitsandbytes Quantization Amplifies Proactive Interference in LLMs

    Shayan Shahrabi-Farahani, Dara Rahmati

    cs.CLcs.LGarXiv:2608.18578v12026
  29. LASE: Language-Adversarial Speaker Encoding for Indic Cross-Script Identity Preservation

    Venkata Pushpak Teja Menta

    cs.SDcs.CLeess.ASarXiv:2605.00777v12026
  30. Syntactic Data Augmentation Increases Robustness to Inference Heuristics

    Junghyun Min, R. Thomas McCoy, Dipanjan Das +2

    cs.CLarXiv:2004.11999v12020
  31. HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents

    Qianchu Liu, Sheng Zhang, Guanghui Qin +16

    cs.AIcs.CLcs.CVarXiv:2606.31179v12026
    Summaries:한국어
  32. Scaling Long-Horizon LLM Agent via Context-Folding

    Weiwei Sun, Miao Lu, Zhan Ling +4

    cs.CLcs.LGarXiv:2510.11967v12025
  33. LongCat-Next: Lexicalizing Modalities as Discrete Tokens

    Meituan LongCat Team, Bin Xiao, Chao Wang +86

    cs.CVcs.CLarXiv:2603.27538v12026
  34. Frames2LoRA: Parametric Video Internalization for Vision-Language Models

    Manan Suri, Sarvesh Baskar, Dinesh Manocha

    cs.CVcs.CLarXiv:2606.04351v22026
  35. PaperMentor: A Human-Centered Multi-Agent Writing Tutor for AI Research Papers on Overleaf

    Jiarui Liu, Terry Jingchen Zhang, Ryan Faulkner +17

    cs.CLarXiv:2606.08857v12026
  36. Conditional Hypothesis Generation for LLM-Based Text Analysis with Researcher-Specified Covariates

    Paiheng Xu, Jing Liu, Wei Ai

    cs.CLcs.AIarXiv:2606.03029v12026
  37. Model-Based Quality Assessment for Massively Multilingual Parallel Data

    Abdelaziz M. A. Ibrahim, Zihao Li, Jörg Tiedemann +1

    cs.CLarXiv:2606.00285v12026
  38. CroCo: Cross-Lingual Contrastive Preference Tuning on Self-Generations

    Mike Zhang, Ali Basirat, Desmond Elliott

    cs.CLcs.AIarXiv:2605.26293v12026
  39. MobileGym: A Verifiable and Highly Parallel Simulation Platform for Mobile GUI Agent Research

    Dingbang Wu, Rui Hao, Haiyang Wang +8

    cs.AIcs.CLarXiv:2605.26114v22026
  40. CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows?

    Haolin Chen, Deon Metelski, Leon Qi +30

    cs.CLcs.AIarXiv:2605.16679v22026
  41. UniPrefill: Universal Long-Context Prefill Acceleration via Block-wise Dynamic Sparsification

    Qihang Fan, Huaibo Huang, Zhiying Wu +2

    cs.CLarXiv:2605.06221v12026
  42. LightMem: Lightweight and Efficient Memory-Augmented Generation

    Jizhan Fang, Xinle Deng, Haoming Xu +9

    cs.CLcs.AIcs.CVarXiv:2510.18866v42025
  43. Polyglot Teachers: Evaluating Language Models for Multilingual Synthetic Data Generation

    Lester James V. Miranda, Ivan Vulić, Anna Korhonen

    cs.CLarXiv:2604.11290v22026
  44. Model Capability Dominates: Inference-Time Optimization Lessons from AIMO 3

    Natapong Nitarach

    cs.CLarXiv:2603.27844v22026
  45. S0 Tuning: Zero-Overhead Adaptation of Hybrid Recurrent-Attention Models

    Jack Young

    cs.CLcs.LGarXiv:2604.01168v22026
  46. Reaching Beyond the Mode: RL for Distributional Reasoning in Language Models

    Isha Puri, Mehul Damani, Idan Shenfeld +3

    cs.LGcs.AIcs.CLarXiv:2603.24844v12026
  47. Memento-Skills: Let Agents Design Agents

    Huichi Zhou, Siyuan Guo, Anjie Liu +14

    cs.AIcs.CLcs.LGarXiv:2603.18743v12026
  48. Tabular LLMs for Interpretable Few-Shot Alzheimer's Disease Prediction with Multimodal Biomedical Data

    Sophie Kearney, Shu Yang, Zixuan Wen +8

    cs.CLcs.LGq-bio.QMarXiv:2603.17191v12026
  49. Unified Vision-Language Modeling via Concept Space Alignment

    Yifu Qiu, Paul-Ambroise Duquenne, Holger Schwenk

    cs.CVcs.AIcs.CLarXiv:2603.01096v12026
  50. CiteAudit: You Cited It, But Did You Read It? A Benchmark for Verifying Scientific References in the LLM Era

    Kaiwen Shi, Weixiang Sun, Zheyuan Zhang +3

    cs.CLcs.DLarXiv:2602.23452v32026
  51. References Improve LLM Alignment in Non-Verifiable Domains

    Kejian Shi, Yixin Liu, Peifeng Wang +3

    cs.CLcs.AIcs.LGarXiv:2602.16802v12026
  52. s1: Simple test-time scaling

    Niklas Muennighoff, Zitong Yang, Weijia Shi +7

    cs.CLcs.AIcs.LGarXiv:2501.19393v32025
  53. Seeing is Coding: On the Effectiveness of Vision Language Models in Code Understanding

    Yuling Shi, Chaoxiang Xie, Zhensu Sun +7

    cs.CLcs.SEarXiv:2602.01785v32026
  54. TTCS: Test-Time Curriculum Synthesis for Self-Evolving

    Chengyi Yang, Zhishang Xiang, Yunbo Tang +5

    cs.LGcs.AIcs.CLarXiv:2601.22628v12026
  55. AgentLongBench: A Controllable Long Benchmark For Long-Contexts Agents via Environment Rollouts

    Shicheng Fang, Yuxin Wang, Xiaoran Liu +6

    cs.CLarXiv:2601.20730v32026
  56. REASONING GYM: Reasoning Environments for Reinforcement Learning with Verifiable Rewards

    Zafir Stojanovski, Oliver Stanley, Joe Sharratt +4

    cs.LGcs.AIcs.CLarXiv:2505.24760v22025
  57. Multiplex Thinking: Reasoning via Token-wise Branch-and-Merge

    Yao Tang, Li Dong, Yaru Hao +3

    cs.CLcs.AIcs.LGarXiv:2601.08808v12026
  58. FlashLabs Chroma 1.0: A Real-Time End-to-End Spoken Dialogue Model with Personalized Voice Cloning

    Tanyu Chen, Tairan Chen, Kai Shen +4

    cs.SDcs.CLeess.ASarXiv:2601.11141v12026
  59. SampoNLP: A Self-Referential Toolkit for Morphological Analysis of Subword Tokenizers

    Iaroslav Chelombitko, Ekaterina Chelombitko, Aleksey Komissarov

    cs.CLcs.IRcs.LGarXiv:2601.04469v12026
  60. MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

    Kunlun Zhu, Hongyi Du, Zhaochen Hong +8

    cs.MAcs.AIcs.CLarXiv:2503.01935v12025