Artificial Intelligence
Papers filed under cs.AI on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
61 to 120 of 15,189
NovBench: Evaluating Large Language Models on Academic Paper Novelty Assessment
Wenqing Wu, Yi Zhao, Yuzhuo Wang +4
cs.CLcs.AIcs.DLarXiv:2604.11543v12026Trustworthy AI: From Principles to Practices
Bo Li, Peng Qi, Bo Liu +5
cs.AIcs.LGarXiv:2110.01167v22021Diagnosing with Insights: Structured Analysis of Agent Failures via Behavioral Abstractions
Jiayi Bi, Yanjie Gao, Yuanmin Xie +4
cs.AIarXiv:2609.02371v12026ClearSight: Visual Signal Enhancement for Object Hallucination Mitigation in Multimodal Large language Models
Hao Yin, Guangzong Si, Zilei Wang
cs.CVcs.AIarXiv:2503.13107v22025Log-based Anomaly Detection Without Log Parsing
Van-Hoang Le, Hongyu Zhang
cs.SEcs.AIarXiv:2108.01955v32021Graph-of-Skills: Dependency-Aware Structural Retrieval for Massive Agent Skills
Dawei Liu, Zongxia Li, Hongyang Du +4
cs.AIarXiv:2604.05333v32026Combee: Scaling Prompt Learning for Self-Improving Language Model Agents
Hanchen Li, Runyuan He, Qizheng Zhang +11
cs.AIcs.CLcs.LGarXiv:2604.04247v12026Bayesian Deep Convolutional Networks with Many Channels are Gaussian Processes
Roman Novak, Lechao Xiao, Jaehoon Lee +6
stat.MLcs.AIcs.LGarXiv:1810.05148v42018Sparse Readout Prism: Explaining Logit-Lens Scores in Features Instead of Tokens
Matteo He, William F. Shen, Xinchi Qiu +1
cs.CLcs.AIcs.LGarXiv:2609.01936v12026Part-based Graph Convolutional Network for Action Recognition
Kalpit Thakkar, P J Narayanan
cs.CVcs.AIarXiv:1809.04983v12018A Gentle Introduction to Conformal Prediction and Distribution-Free Uncertainty Quantification
Anastasios N. Angelopoulos, Stephen Bates
cs.LGcs.AImath.STarXiv:2107.07511v62021Probing Factual Knowledge Transfer with Training Data Interventions
Romina Oji, Marc Braun, Marcel Bollmann +2
cs.CLcs.AIarXiv:2609.01341v12026World Action Verifier: Self-Improving World Models via Forward-Inverse Asymmetry
Yuejiang Liu, Fan Feng, Lingjing Kong +6
cs.LGcs.AIcs.ROarXiv:2604.01985v220263D Consistent & Robust Segmentation of Cardiac Images by Deep Learning with Spatial Propagation
Qiao Zheng, Hervé Delingette, Nicolas Duchateau +1
cs.CVcs.AIcs.LGarXiv:1804.09400v12018Seq2Seq-Vis: A Visual Debugging Tool for Sequence-to-Sequence Models
Hendrik Strobelt, Sebastian Gehrmann, Michael Behrisch +3
cs.CLcs.AIcs.NEarXiv:1804.09299v22018Automatic Image-Level Morphological Trait Annotation for Organismal Images
Vardaan Pahuja, Samuel Stevens, Alyson East +2
cs.CVcs.AIarXiv:2604.01619v32026WHALE: A Simple Recipe for Joint Harness-Weight Optimization
Haechan Kim, Yoonho Lee, Gisang Lee +2
cs.LGcs.AIarXiv:2609.00196v12026Learning a SAT Solver from Single-Bit Supervision
Daniel Selsam, Matthew Lamm, Benedikt Bünz +3
cs.AIcs.LGcs.LOarXiv:1802.03685v42018Experiential Reflective Learning for Self-Improving LLM Agents
Marc-Antoine Allard, Arnaud Teinturier, Victor Xing +1
cs.LGcs.AIarXiv:2603.24639v22026VFIG: Vectorizing Complex Figures in SVG with Vision-Language Models
Qijia He, Xunmei Liu, Hammaad Memon +6
cs.CVcs.AIarXiv:2603.24575v22026Uncertainty-guided Compositional Alignment with Part-to-Whole Semantic Representativeness in Hyperbolic Vision-Language Models
Hayeon Kim, Ji Ha Jang, Junghun James Kim +1
cs.CVcs.AIarXiv:2603.22042v32026Perceive to Hypothesize, Verify to Ground: An Agentic Reasoning Framework for Open-World Geo-Localization
Yutian Jiang, Ruijie Li, Sisuo Lyu +4
cs.AIcs.MMarXiv:2608.29880v12026AgentHER: Hindsight Experience Replay for LLM Agent Trajectory Relabeling
Liang Ding
cs.AIcs.CLarXiv:2603.21357v42026Background-Free Objectness Learning for Class-Agnostic Detection
Dania Batool, Liliana Lo Presti, Marco La Cascia +1
cs.CVcs.AIcs.ROarXiv:2608.29232v12026AgentProcessBench: Diagnosing Step-Level Process Quality in Tool-Using Agents
Shengda Fan, Xuyan Ye, Yupeng Huo +9
cs.AIarXiv:2603.14465v22026Active learning machine learns to create new quantum experiments
Alexey A. Melnikov, Hendrik Poulsen Nautrup, Mario Krenn +4
quant-phcs.AIstat.MLarXiv:1706.00868v32017MuseControlLite: Multifunctional Music Generation with Lightweight Conditioners
Fang-Duo Tsai, Shih-Lun Wu, Weijaw Lee +4
cs.SDcs.AIeess.ASarXiv:2506.18729v22025MiniAppBench: Evaluating the Shift from Text to Interactive HTML Responses in LLM-Powered Assistants
Zuhao Zhang, Chengyue Yu, Yuante Li +3
cs.AIarXiv:2603.09652v32026Daydreaming: Stealing Hidden Agent Skills through Black-Box Task Interaction
Yu-Lin Tsai, Yu-An Lu, Ci-Yang Tsai +3
cs.CRcs.AIcs.LGarXiv:2608.26733v12026FinToolBench: Evaluating LLM Agents for Real-World Financial Tool Use
Jiaxuan Lu, Kong Wang, Yemin Wang +9
cs.AIarXiv:2603.08262v22026Candidate supply and answer selection shape the value of LLM judging in multi-agent systems
Jia-Hao Ji, Sijie Li, Jiabei Cheng +3
cs.AIcs.MAarXiv:2608.25937v12026AdapTools: Adaptive Tool-based Indirect Prompt Injection Attacks on Agentic LLMs
Che Wang, Jiaming Zhang, Ziqi Zhang +6
cs.CRcs.AIarXiv:2602.20720v12026ReIn: Conversational Error Recovery with Reasoning Inception
Takyoung Kim, Jinseok Nam, Chandrayee Basu +5
cs.CLcs.AIarXiv:2602.17022v12026AI Slop and Hallucinations in Vulnerability Assessment: A Survey on Reasoning Failures and Trustworthy Mitigation
Junchen Ding, Jialiang Dong, Yichen Zhu +5
cs.CRcs.AIarXiv:2608.25667v12026GENIUS: Generative Fluid Intelligence Evaluation Suite
Ruichuan An, Sihan Yang, Ziyu Guo +8
cs.LGcs.AIcs.CVarXiv:2602.11144v12026AudioLens: Multi-Perspective Speech Clustering with Reasoning Audio-Language Models
Wenjun Huang, Qiaosong Chu, Tiger Shao +11
cs.SDcs.AIarXiv:2608.25177v12026FuzzingBrain-Bench V1: Evaluating Open-Ended Bug Discovery by LLMs
Ze Sheng, Aleksandar Kezic, Zhicheng Chen +1
cs.AIcs.CRcs.LGarXiv:2608.25158v12026Towards Agentic Intelligence for Materials Science
Huan Zhang, Yizhan Li, Wenhao Huang +18
cond-mat.mtrl-scics.AIarXiv:2602.00169v22026The Hot Mess of AI: How Does Misalignment Scale With Model Intelligence and Task Complexity?
Alexander Hägele, Aryo Pradipta Gema, Henry Sleight +2
cs.AIarXiv:2601.23045v22026What is Wrong with Topic Modeling? (and How to Fix it Using Search-based Software Engineering)
Amritanshu Agrawal, Wei Fu, Tim Menzies
cs.SEcs.AIcs.CLarXiv:1608.08176v42016More Rejective, Not More Discriminative: The Unit of Verification in Pre-Execution LLM Oversight
Yuchen Han, Cheng Yan, Wuyang Zhang
cs.AIarXiv:2608.23941v12026LUCAID: Agentic Multimodal AI for Lung Cancer Precision Pathology
Marie-Lisa Eich, Kai Standvoss, Timo Milbich +30
cs.CVcs.AIcs.LGarXiv:2608.23803v12026Less Noise, More Voice: Reinforcement Learning for Reasoning via Instruction Purification
Yiju Guo, Tianyi Hu, Zexu Sun +1
cs.LGcs.AIcs.CLarXiv:2601.21244v32026Mem2ActBench: A Benchmark for Evaluating Long-Term Memory Utilization in Task-Oriented Autonomous Agents
Yiting Shen, Kun Li, Wei Zhou +1
cs.CLcs.AIarXiv:2601.19935v12026From Generation to Simulation: How Far Are World Models from Being True Simulators?
Tong Wang, Huan Deng, Mucheng Yang +3
cs.AIcs.CVarXiv:2608.23070v12026Why Can't I Open My Drawer? Mitigating Object-Driven Shortcuts in Zero-Shot Compositional Action Recognition
Geo Ahn, Inwoong Lee, Taeoh Kim +3
cs.CVcs.AIarXiv:2601.16211v32026Summaries:한국어Deep Knowledge Tracing
Chris Piech, Jonathan Spencer, Jonathan Huang +4
cs.AIcs.CYcs.LGarXiv:1506.05908v12015Budget-Constrained Embodied Perception: Four Resource Walls and a Pre-Registered Evaluation of Access-Structured Perception on Open Models at less than 31B
Defu Lin, Wenhui Chen, Ziyao Lin +3
cs.AIarXiv:2608.22975v12026Barycentric Fused Gromov-Wasserstein Balancing for Causal Inference under Multiple Treatments
Yuki Murakami, Takumi Hattori, Kohsuke Kubota
stat.MEcs.AIcs.LGarXiv:2608.22024v12026Transition Matching Distillation for Fast Video Generation
Weili Nie, Julius Berner, Nanye Ma +3
cs.CVcs.AIcs.LGarXiv:2601.09881v22026MirrorBench: A Benchmark to Evaluate Conversational User-Proxy Agents for Human-Likeness
Ashutosh Hathidara, Julien Yu, Vaishali Senthil +2
cs.AIcs.LGarXiv:2601.08118v32026On the Role of Citations in Preference Data
Yu Hou, Hal Daumé, Rachel Rudinger +1
cs.CLcs.AIarXiv:2608.21376v12026Beyond Hard Masks: Progressive Token Evolution for Diffusion Language Models
Linhao Zhong, Linyu Wu, Bozhen Fang +6
cs.CLcs.AIarXiv:2601.07351v22026CellPath-Bench: A Multidimensional Benchmark for Whole-Slide Cellular Representations in Pathology Foundation Models
Bokai Zhao, Yiyang Zhang, Hanqing Chao +6
cs.AIcs.CVarXiv:2608.21060v12026Beyond Static Summarization: Proactive Memory Extraction for LLM Agents
Chengyuan Yang, Zequn Sun, Wei Wei +1
cs.CLcs.AIarXiv:2601.04463v22026MentorPulse: Refreshing Cross-Model Latent Guidance for Long-Form Generation
Ziwu Liu, Guozhong Li, Chen Qiu +2
cs.CLcs.AIarXiv:2608.20927v12026CDRL: Certification-Driven Reinforcement Learning for Neutrino Flavor Model Discovery
Piyush Jha, Jake Rudolph, Victoria Knapp-Pérez +3
cs.AIcs.LGcs.LOarXiv:2608.20686v12026The Principles of Diffusion Models
Chieh-Hsin Lai, Yang Song, Dongjun Kim +2
cs.LGcs.AIcs.GRarXiv:2510.21890v32025Toward Auto-Research: Mining Falsifiable Research Ideas from Paper Knowledge Graphs with Categorical Structure
Yuchen Wang, Zhongzhi Luan
cs.CLcs.AIarXiv:2608.20361v12026How to Train a Real-World Silicon Concierge? Internalizing Complex Business Workflow to Only OneModel
Chang Liu, Chaoyang Ning, Dayi Jiang +32
cs.CLcs.AIarXiv:2608.20350v12026