Artificial Intelligence

Papers filed under cs.AI on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

121 to 180 of 15,205

  1. LUCAID: Agentic Multimodal AI for Lung Cancer Precision Pathology

    Marie-Lisa Eich, Kai Standvoss, Timo Milbich +30

    cs.CVcs.AIcs.LGarXiv:2608.23803v12026
  2. Less Noise, More Voice: Reinforcement Learning for Reasoning via Instruction Purification

    Yiju Guo, Tianyi Hu, Zexu Sun +1

    cs.LGcs.AIcs.CLarXiv:2601.21244v32026
  3. Mem2ActBench: A Benchmark for Evaluating Long-Term Memory Utilization in Task-Oriented Autonomous Agents

    Yiting Shen, Kun Li, Wei Zhou +1

    cs.CLcs.AIarXiv:2601.19935v12026
  4. From Generation to Simulation: How Far Are World Models from Being True Simulators?

    Tong Wang, Huan Deng, Mucheng Yang +3

    cs.AIcs.CVarXiv:2608.23070v12026
  5. Why Can't I Open My Drawer? Mitigating Object-Driven Shortcuts in Zero-Shot Compositional Action Recognition

    Geo Ahn, Inwoong Lee, Taeoh Kim +3

    cs.CVcs.AIarXiv:2601.16211v32026
    Summaries:한국어
  6. Deep Knowledge Tracing

    Chris Piech, Jonathan Spencer, Jonathan Huang +4

    cs.AIcs.CYcs.LGarXiv:1506.05908v12015
  7. Budget-Constrained Embodied Perception: Four Resource Walls and a Pre-Registered Evaluation of Access-Structured Perception on Open Models at less than 31B

    Defu Lin, Wenhui Chen, Ziyao Lin +3

    cs.AIarXiv:2608.22975v12026
  8. Barycentric Fused Gromov-Wasserstein Balancing for Causal Inference under Multiple Treatments

    Yuki Murakami, Takumi Hattori, Kohsuke Kubota

    stat.MEcs.AIcs.LGarXiv:2608.22024v12026
  9. Transition Matching Distillation for Fast Video Generation

    Weili Nie, Julius Berner, Nanye Ma +3

    cs.CVcs.AIcs.LGarXiv:2601.09881v22026
  10. MirrorBench: A Benchmark to Evaluate Conversational User-Proxy Agents for Human-Likeness

    Ashutosh Hathidara, Julien Yu, Vaishali Senthil +2

    cs.AIcs.LGarXiv:2601.08118v32026
  11. On the Role of Citations in Preference Data

    Yu Hou, Hal Daumé, Rachel Rudinger +1

    cs.CLcs.AIarXiv:2608.21376v12026
  12. Beyond Hard Masks: Progressive Token Evolution for Diffusion Language Models

    Linhao Zhong, Linyu Wu, Bozhen Fang +6

    cs.CLcs.AIarXiv:2601.07351v22026
  13. CellPath-Bench: A Multidimensional Benchmark for Whole-Slide Cellular Representations in Pathology Foundation Models

    Bokai Zhao, Yiyang Zhang, Hanqing Chao +6

    cs.AIcs.CVarXiv:2608.21060v12026
  14. MemGovern: Enhancing Code Agents through Learning from Governed Human Experiences

    Qihao Wang, Ziming Cheng, Shuo Zhang +12

    cs.SEcs.AIarXiv:2601.06789v22026
  15. Beyond Static Summarization: Proactive Memory Extraction for LLM Agents

    Chengyuan Yang, Zequn Sun, Wei Wei +1

    cs.CLcs.AIarXiv:2601.04463v22026
  16. MentorPulse: Refreshing Cross-Model Latent Guidance for Long-Form Generation

    Ziwu Liu, Guozhong Li, Chen Qiu +2

    cs.CLcs.AIarXiv:2608.20927v12026
  17. CDRL: Certification-Driven Reinforcement Learning for Neutrino Flavor Model Discovery

    Piyush Jha, Jake Rudolph, Victoria Knapp-Pérez +3

    cs.AIcs.LGcs.LOarXiv:2608.20686v12026
  18. The Principles of Diffusion Models

    Chieh-Hsin Lai, Yang Song, Dongjun Kim +2

    cs.LGcs.AIcs.GRarXiv:2510.21890v32025
  19. Toward Auto-Research: Mining Falsifiable Research Ideas from Paper Knowledge Graphs with Categorical Structure

    Yuchen Wang, Zhongzhi Luan

    cs.CLcs.AIarXiv:2608.20361v12026
  20. How to Train a Real-World Silicon Concierge? Internalizing Complex Business Workflow to Only OneModel

    Chang Liu, Chaoyang Ning, Dayi Jiang +32

    cs.CLcs.AIarXiv:2608.20350v12026
  21. Repo0: Design-Driven Zero-to-All Code Generation

    Silin Chen, Haoyi Teng, Xiaodong Gu +5

    cs.SEcs.AIarXiv:2608.19854v12026
  22. What is Missing from AI Post-Training AI: An Empirical Analysis

    Joy Jia Yin Lim, Xin Huang, Hao Peng +5

    cs.AIcs.CLcs.LGarXiv:2608.19072v12026
  23. Flama: a Python framework for development and deployment of production-ready APIs, machine learning, and LLM services

    José A. Perdiguero López, Miguel A. Durán-Olivencia

    cs.SEcs.AIcs.LGarXiv:2608.18733v12026
  24. Partition the Support, Reconstruct the Residual: Training-Free Sparse Attention for Video Generation and World Models

    Pardis Taghavi, Reza Langari, Gaurav Pandey

    cs.CVcs.AIcs.LGarXiv:2608.18484v12026
  25. Formal Verification of Romanov's Triplet Logic: A Verified Filter for Sliding-window 3-CNF with Application to Structured Formulas

    Dmitry V. Alexandrov

    cs.LOcs.AIcs.CCarXiv:2608.18445v12026
  26. Position: AI Leaderboards Are Underserving the Global South: A Case Study from India

    Sourav Banerjee, Saikat Saha

    cs.AIecon.GNarXiv:2608.18117v12026
  27. SAGE: Self-Evolving Storyboard Skills via Attribution-Guided Rule Evolution

    Maolin Ran, Xiaoyang Lu, Jiaqi Liu +5

    cs.AIarXiv:2608.17468v12026
  28. MapAnything: Universal Feed-Forward Metric 3D Reconstruction

    Nikhil Keetha, Norman Müller, Johannes Schönberger +14

    cs.CVcs.AIcs.LGarXiv:2509.13414v32025
  29. DeAR: Decentralized Agentic Reasoning via Capability Grounding and Collaborative Thought Navigation

    Xing Wei, Changmeng Zheng, XiaoYong Wei +2

    cs.AIarXiv:2608.17282v12026
  30. Learning Where and What to Lift for Bi-planar X-ray-to-CT Reconstruction

    Yifei Wu, Yicheng Wu, Qiang Ma +5

    cs.CVcs.AIarXiv:2608.17255v12026
  31. A Survey of Reinforcement Learning for Large Reasoning Models

    Kaiyan Zhang, Yuxin Zuo, Bingxiang He +36

    cs.CLcs.AIcs.LGarXiv:2509.08827v32025
  32. HAF: Adapting Generalist VLAs to Humanoid Whole-Body Loco-manipulation via Hierarchical Action Flow and Spectral Latent RL

    Langzhe Gu, Chengkai Hou, Meng Li +14

    cs.ROcs.AIarXiv:2608.16837v12026
  33. Degradation-Aligned Self-Supervised Learning for State of Health Estimation of Lithium-Ion Batteries under Label Sparsity

    Jiaqi Yao, Julia Kowal

    eess.SPcs.AIcs.LGarXiv:2608.16612v12026
  34. Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

    Gheorghe Comanici, Eric Bieber, Mike Schaekermann +3432

    cs.CLcs.AIarXiv:2507.06261v62025
  35. Discrete Diffusion in Large Language and Multimodal Models: A Survey

    Runpeng Yu, Qi Li, Xinchao Wang

    cs.LGcs.AIarXiv:2506.13759v52025
  36. AlphaEvolve: A coding agent for scientific and algorithmic discovery

    Alexander Novikov, Ngân Vũ, Marvin Eisenberger +15

    cs.AIcs.LGcs.NEarXiv:2506.13131v12025
  37. Pre-Training BERT on Arabic Tweets: Practical Considerations

    Ahmed Abdelali, Sabit Hassan, Hamdy Mubarak +2

    cs.CLcs.AIarXiv:2102.10684v12021
  38. Moral Stories: Situated Reasoning about Norms, Intents, Actions, and their Consequences

    Denis Emelin, Ronan Le Bras, Jena D. Hwang +2

    cs.CLcs.AIarXiv:2012.15738v12020
  39. J1: Incentivizing Thinking in LLM-as-a-Judge via Reinforcement Learning

    Chenxi Whitehouse, Tianlu Wang, Ping Yu +4

    cs.CLcs.AIcs.LGarXiv:2505.10320v32025
  40. MLLM-Guided Semantic Correction for Text-to-Video Generation

    Junhao Chen, Zheqi Lv, Keting Yin +6

    cs.CVcs.AIarXiv:2608.16513v12026
  41. Absolute Zero: Reinforced Self-play Reasoning with Zero Data

    Andrew Zhao, Yiran Wu, Yang Yue +8

    cs.LGcs.AIcs.CLarXiv:2505.03335v32025
  42. STAIR: Semantic-Temporal Automaton for Interpretable Reasoning in Temporal Question Answering

    Xinlong Dai, Jinchuan Zhang, Lei Gao +3

    cs.CLcs.AIarXiv:2608.16224v12026
  43. VARM-Bench: Benchmarking Verifiable Structured Reasoning in Chinese Abusive Speech Moderation

    Mingyu Yuan, Shengtao Wen, Lingbing Guo +2

    cs.AIarXiv:2608.15600v12026
  44. Social Chemistry 101: Learning to Reason about Social and Moral Norms

    Maxwell Forbes, Jena D. Hwang, Vered Shwartz +2

    cs.CLcs.AIarXiv:2011.00620v32020
  45. When Entropy Is Not Enough: Reclaiming Lost Semantics in LLM Output Length Prediction

    Feiyang Ren, Shengtao Wen, Lingbing Guo +3

    cs.AIarXiv:2608.15592v12026
  46. Open Question Answering over Tables and Text

    Wenhu Chen, Ming-Wei Chang, Eva Schlinger +2

    cs.CLcs.AIarXiv:2010.10439v22020
  47. Dynamic Early Exit in Reasoning Models

    Chenxu Yang, Qingyi Si, Yongjie Duan +6

    cs.CLcs.AIarXiv:2504.15895v32025
  48. DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

    Zhiwei He, Tian Liang, Jiahao Xu +12

    cs.CLcs.AIarXiv:2504.11456v22025
  49. SEAL: Steerable Reasoning Calibration of Large Language Models for Free

    Runjin Chen, Zhenyu Zhang, Junyuan Hong +2

    cs.CLcs.AIarXiv:2504.07986v32025
  50. The Quality of Claude AI-authored Python Tests Is Not Weaker Than Human-authored Tests

    Douglas J. Leith

    cs.SEcs.AIarXiv:2608.15188v12026
  51. Calibrated Trust, Not Sharper Prediction: An Empirical Test of Uncertainty Fusion

    Surya Saka

    cs.LGcs.AIcs.CLarXiv:2608.14617v12026
  52. Plausible but Not Valid: A Psychometric Audit of LLMs as Synthetic Survey Respondents

    Mantas Lukauskas, Viktorija Šarkauskaitė

    cs.CYcs.AIcs.CLarXiv:2608.14606v12026
  53. HiCo-GS: Hierarchical Context Aggregation and Geometric Consistency for Octree Gaussian Splatting

    Wei Zhang, Shengkai Yu, Shiqiang Gong +3

    cs.CVcs.AIarXiv:2608.14136v12026
  54. MedAgent-Pro: Towards Evidence-based Multi-modal Medical Diagnosis via Reasoning Agentic Workflow

    Ziyue Wang, Junde Wu, Linghan Cai +4

    cs.AIarXiv:2503.18968v32025
  55. Online Bayesian Goal Inference for Boundedly-Rational Planning Agents

    Tan Zhi-Xuan, Jordyn L. Mann, Tom Silver +2

    cs.AIarXiv:2006.07532v22020
  56. TokenSqueeze: Performance-Preserving Compression for Reasoning LLMs

    Yuxiang Zhang, Zhengxu Yu, Weihang Pan +5

    cs.LGcs.AIarXiv:2511.13223v12025
  57. TxAgent: An AI Agent for Therapeutic Reasoning Across a Universe of Tools

    Shanghua Gao, Richard Zhu, Zhenglun Kong +5

    cs.AIcs.LGarXiv:2503.10970v12025
  58. Evidence-RL: Towards Evidence-intensive Visual Reasoning

    Haojie Huang, Xinlei Yu, Chengming Xu +6

    cs.CVcs.AIarXiv:2608.08021v12026
  59. Multi-Agent Forensic Reasoning for Generalizable Deepfake Video Detection

    Xuechao Zou, Shun Zhang, Kai Li +6

    cs.CVcs.AIcs.MAarXiv:2608.06865v12026
  60. EnvACE: Internalizing Environment Dynamics via World Rehearsal for Agentic Reinforcement Learning

    Zishan Xu, Zhiyuan Yao, Yuxin Chen +9

    cs.AIarXiv:2608.06197v12026