Artificial Intelligence

Papers filed under cs.AI on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

3,061 to 3,120 of 15,502

  1. RubRIX: Rubric-Driven Risk Mitigation in Caregiver-AI Interactions

    Drishti Goel, Jeongah Lee, Qiuyue Joy Zhong +5

    cs.HCcs.AIcs.CLarXiv:2601.13235v12026
  2. Parallelism and Generation Order in Masked Diffusion Language Models: Limits Today, Potential Tomorrow

    Yangyang Zhong, Yanmei Gu, Zhengqing Zang +14

    cs.CLcs.AIcs.LGarXiv:2601.15593v22026
  3. MMR-Bench: A Comprehensive Benchmark for Multimodal LLM Routing

    Haoxuan Ma, Guannan Lai, Han-Jia Ye

    cs.AIarXiv:2601.17814v12026
  4. MMErroR: A Benchmark for Erroneous Reasoning in Vision-Language Models

    Yang Shi, Yifeng Xie, Minzhe Guo +6

    cs.CVcs.AIcs.LGarXiv:2601.03331v22026
  5. CircuChain: Disentangling Competence and Compliance in LLM Circuit Analysis

    Mayank Ravishankara

    cs.SEcs.AIarXiv:2602.15037v12026
  6. Probability-Entropy Calibration: An Elastic Indicator for Adaptive Fine-tuning

    Wenhao Yu, Shaohang Wei, Jiahong Liu +5

    cs.LGcs.AIarXiv:2602.01745v22026
  7. Vision-Language Introspection: Mitigating Overconfident Hallucinations in MLLMs via Interpretable Bi-Causal Steering

    Shuliang Liu, Songbo Yang, Dong Fang +7

    cs.CVcs.AIarXiv:2601.05159v12026
  8. Match-SRNN: Modeling the Recursive Matching Structure with Spatial RNN

    Shengxian Wan, Yanyan Lan, Jun Xu +3

    cs.CLcs.AIcs.LGarXiv:1604.04378v12016
  9. Beyond Precision: Training-Inference Mismatch is an Optimization Problem and Simple LR Scheduling Fixes It

    Yaxiang Zhang, Yingru Li, Jiacai Liu +4

    cs.LGcs.AIarXiv:2602.01826v12026
  10. Multiplex Behavioral Relation Learning for Recommendation via Memory Augmented Transformer Network

    Lianghao Xia, Chao Huang, Yong Xu +3

    cs.IRcs.AIarXiv:2110.04002v12021
  11. NextMem: Towards Latent Factual Memory for LLM-based Agents

    Zeyu Zhang, Rui Li, Xiaoyan Zhao +4

    cs.AIcs.IRcs.LGarXiv:2603.15634v12026
  12. Small Language Models: Survey, Measurements, and Insights

    Zhenyan Lu, Xiang Li, Dongqi Cai +5

    cs.CLcs.AIcs.LGarXiv:2409.15790v32024
  13. Do Latent-CoT Models Think Step-by-Step? A Mechanistic Study on Sequential Reasoning Tasks

    Jia Liang, Liangming Pan

    cs.AIcs.LGarXiv:2602.00449v12026
  14. Flow as the Cross-Domain Manipulation Interface

    Mengda Xu, Zhenjia Xu, Yinghao Xu +4

    cs.ROcs.AIarXiv:2407.15208v22024
  15. Penalizing Gradient Norm for Efficiently Improving Generalization in Deep Learning

    Yang Zhao, Hao Zhang, Xiuyuan Hu

    cs.LGcs.AIarXiv:2202.03599v32022
  16. AdaMARP: An Adaptive Multi-Agent Interaction Framework for General Immersive Role-Playing

    Zhenhua Xu, Dongsheng Chen, Shuo Wang +4

    cs.AIcs.CLarXiv:2601.11007v22026
  17. Reducing False Positives in Static Bug Detection with LLMs: An Empirical Study in Industry

    Xueying Du, Jiayi Feng, Yi Zou +6

    cs.SEcs.AIarXiv:2601.18844v12026
  18. GR2: Generative Reasoning Re-ranker

    Mingfu Liang, Yufei Li, Jay Xu +20

    cs.IRcs.AIarXiv:2602.07774v62026
  19. Dynamics Within Latent Chain-of-Thought: An Empirical Study of Causal Structure

    Zirui Li, Xuefeng Bai, Kehai Chen +4

    cs.AIcs.CLarXiv:2602.08783v32026
  20. Image Generators with Conditionally-Independent Pixel Synthesis

    Ivan Anokhin, Kirill Demochkin, Taras Khakhulin +3

    cs.CVcs.AIcs.LGarXiv:2011.13775v12020
  21. SWE Context Bench: A Benchmark for Context Learning in Coding

    Jiayuan Zhu, Junde Wu, Minhao Hu +9

    cs.SEcs.AIarXiv:2602.08316v32026
  22. Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

    Xiong Wang, Yangze Li, Chaoyou Fu +5

    cs.SDcs.AIcs.CLarXiv:2411.00774v52024
  23. Semi-Supervised Learning of Visual Features by Non-Parametrically Predicting View Assignments with Support Samples

    Mahmoud Assran, Mathilde Caron, Ishan Misra +4

    cs.CVcs.AIcs.LGarXiv:2104.13963v32021
  24. Just-In-Time Reinforcement Learning: Continual Learning in LLM Agents Without Gradient Updates

    Yibo Li, Zijie Lin, Ailin Deng +5

    cs.LGcs.AIarXiv:2601.18510v32026
  25. HiPER: Hierarchical Reinforcement Learning with Explicit Credit Assignment for Large Language Model Agents

    Jiangweizhi Peng, Yuanxin Liu, Ruida Zhou +4

    cs.LGcs.AIarXiv:2602.16165v22026
  26. Low-Dimensional and Transversely Curved Optimization Dynamics in Grokking

    Yongzhong Xu

    cs.LGcs.AIarXiv:2602.16746v32026
  27. ES-MemEval: Benchmarking Conversational Agents on Personalized Long-Term Emotional Support

    Tiantian Chen, Jiaqi Lu, Ying Shen +1

    cs.CLcs.AIarXiv:2602.01885v12026
  28. The Generative AI Paradox: GenAI and the Erosion of Trust, the Corrosion of Information Verification, and the Demise of Truth

    Emilio Ferrara

    cs.CYcs.AIcs.HCarXiv:2601.00306v12026
  29. AvatarPoser: Articulated Full-Body Pose Tracking from Sparse Motion Sensing

    Jiaxi Jiang, Paul Streli, Huajian Qiu +4

    cs.CVcs.AIcs.GRarXiv:2207.13784v12022
  30. AdaPoinTr: Diverse Point Cloud Completion with Adaptive Geometry-Aware Transformers

    Xumin Yu, Yongming Rao, Ziyi Wang +2

    cs.CVcs.AIarXiv:2301.04545v12023
  31. Industrialized Deception: The Collateral Effects of LLM-Generated Misinformation on Digital Ecosystems

    Alexander Loth, Martin Kappes, Marc-Oliver Pahl

    cs.CYcs.AIcs.CLarXiv:2601.21963v22026
  32. QUASAR: A Universal Autonomous System for Atomistic Simulation and a Benchmark of Its Capabilities

    Fengxu Yang, Jack D. Evans

    cond-mat.mtrl-scics.AIarXiv:2602.00185v22026
  33. Deep Whole-body Parkour

    Ziwen Zhuang, Shaoting Zhu, Mengjie Zhao +1

    cs.ROcs.AIarXiv:2601.07701v12026
  34. Beyond Dialogue Time: Temporal Semantic Memory for Personalized LLM Agents

    Miao Su, Yucan Guo, Zhongni Hou +8

    cs.AIarXiv:2601.07468v22026
  35. Tackling Data Heterogeneity in Federated Learning with Class Prototypes

    Yutong Dai, Zeyuan Chen, Junnan Li +3

    cs.LGcs.AIarXiv:2212.02758v22022
  36. MogaNet: Multi-order Gated Aggregation Network

    Siyuan Li, Zedong Wang, Zicheng Liu +6

    cs.CVcs.AIarXiv:2211.03295v42022
  37. Understanding Neural Networks via Feature Visualization: A survey

    Anh Nguyen, Jason Yosinski, Jeff Clune

    cs.LGcs.AIcs.CVarXiv:1904.08939v12019
  38. Ethics of AI: A Systematic Literature Review of Principles and Challenges

    Arif Ali Khan, Sher Badshah, Peng Liang +4

    cs.CYcs.AIarXiv:2109.07906v12021
  39. BiasScope: Towards Automated Detection of Bias in LLM-as-a-Judge Evaluation

    Peng Lai, Zhihao Ou, Yong Wang +4

    cs.CLcs.AIcs.SEarXiv:2602.09383v12026
  40. Ask don't tell: Reducing sycophancy in large language models

    Magda Dubois, Cozmin Ududec, Christopher Summerfield +1

    cs.HCcs.AIarXiv:2602.23971v42026
  41. HFedMoE: Resource-aware Heterogeneous Federated Learning with Mixture-of-Experts

    Zihan Fang, Zheng Lin, Senkang Hu +5

    cs.LGcs.AIcs.NIarXiv:2601.00583v12026
  42. Improving Neural Question Generation using Answer Separation

    Yanghoon Kim, Hwanhee Lee, Joongbo Shin +1

    cs.CLcs.AIcs.NEarXiv:1809.02393v22018
  43. ASTRA-bench: Evaluating Tool-Use Agent Reasoning and Action Planning with Personal User Context

    Zidi Xiu, David Q. Sun, Kevin Cheng +9

    cs.AIarXiv:2603.01357v12026
  44. On the origin of neural scaling laws: from random graphs to natural language

    Maissam Barkeshli, Alberto Alfarano, Andrey Gromov

    cs.LGcond-mat.dis-nncs.AIarXiv:2601.10684v12026
  45. ArchAgent: Agentic AI-driven Computer Architecture Discovery

    Raghav Gupta, Akanksha Jain, Abraham Gonzalez +10

    cs.AIcs.ARarXiv:2602.22425v12026
  46. Context Matters: Repository-Aware Security Analysis of the Agent Skill Ecosystem

    Florian Holzbauer, David Schmidt, Gabriel Gegenhuber +2

    cs.CRcs.AIarXiv:2603.16572v22026
  47. SleepLM: Natural-Language Intelligence for Human Sleep

    Zongzhe Xu, Zitao Shuai, Eideen Mozaffari +3

    cs.AIarXiv:2602.23605v22026
  48. DeepSteal: Advanced Model Extractions Leveraging Efficient Weight Stealing in Memories

    Adnan Siraj Rakin, Md Hafizul Islam Chowdhuryy, Fan Yao +1

    cs.CRcs.AIcs.CVarXiv:2111.04625v12021
  49. Auditing Multi-Agent LLM Reasoning Trees Outperforms Majority Vote and LLM-as-Judge

    Wei Yang, Shixuan Li, Heng Ping +3

    cs.AIarXiv:2602.09341v22026
  50. WebCoderBench: Benchmarking Web Application Generation with Comprehensive and Interpretable Evaluation Metrics

    Chenxu Liu, Yingjie Fu, Wei Yang +2

    cs.SEcs.AIarXiv:2601.02430v32026
  51. AQAScore: Evaluating Semantic Alignment in Text-to-Audio Generation via Audio Question Answering

    Chun-Yi Kuan, Kai-Wei Chang, Hung-yi Lee

    eess.AScs.AIcs.CLarXiv:2601.14728v12026
  52. Designing Interpretable ML System to Enhance Trust in Healthcare: A Systematic Review to Proposed Responsible Clinician-AI-Collaboration Framework

    Elham Nasarian, Roohallah Alizadehsani, U. Rajendra Acharya +1

    cs.AIcs.HCcs.LGarXiv:2311.11055v22023
  53. Hypothesize-Then-Verify: Speculative Root Cause Analysis for Microservices with Pathwise Parallelism

    Lingzhe Zhang, Tong Jia, Yunpeng Zhai +5

    cs.SEcs.AIarXiv:2601.02736v12026
  54. LLaTTE: Scaling Laws for Multi-Stage Sequence Modeling in Large-Scale Ads Recommendation

    Lee Xiong, Zhirong Chen, Rahul Mayuranath +17

    cs.IRcs.AIcs.LGarXiv:2601.20083v12026
  55. ENEAS: Embedding-guided Neural Ensemble for Adaptive Segmentation

    Javier del Pino, Salvador Rodríguez, Alejandro Garabito +2

    cs.CVcs.AIarXiv:2609.03756v12026
  56. IT-OSE: Exploring Optimal Sample Size for Industrial Data Augmentation

    Mingchun Sun, Rongqiang Zhao, Zhennan Huang +2

    cs.LGcs.AIarXiv:2602.15878v12026
  57. Stop Uploading Test Data in Plain Text: Practical Strategies for Mitigating Data Contamination by Evaluation Benchmarks

    Alon Jacovi, Avi Caciularu, Omer Goldman +1

    cs.CLcs.AIarXiv:2305.10160v22023
  58. Knowledge distillation from multi-modal to mono-modal segmentation networks

    Minhao Hu, Matthis Maillard, Ya Zhang +4

    cs.CVcs.AIstat.MLarXiv:2106.09564v12021
  59. How AI Coding Agents Modify Code: A Large-Scale Study of GitHub Pull Requests

    Daniel Ogenrwot, John Businge

    cs.SEcs.AIarXiv:2601.17581v32026
  60. An improvement of the convergence proof of the ADAM-Optimizer

    Sebastian Bock, Josef Goppold, Martin Weiß

    cs.LGcs.AIstat.MLarXiv:1804.10587v12018