Artificial Intelligence
Papers filed under cs.AI on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
10,501 to 10,560 of 15,398
Toxicity in ChatGPT: Analyzing Persona-assigned Language Models
Ameet Deshpande, Vishvak Murahari, Tanmay Rajpurohit +2
cs.CLcs.AIcs.LGarXiv:2304.05335v12023Language Models can Solve Computer Tasks
Geunwoo Kim, Pierre Baldi, Stephen McAleer
cs.CLcs.AIcs.HCarXiv:2303.17491v32023AI and the Everything in the Whole Wide World Benchmark
Inioluwa Deborah Raji, Emily M. Bender, Amandalynne Paullada +2
cs.LGcs.AIcs.PFarXiv:2111.15366v12021DL3DV-10K: A Large-Scale Scene Dataset for Deep Learning-based 3D Vision
Lu Ling, Yichen Sheng, Zhi Tu +17
cs.CVcs.AIarXiv:2312.16256v22023Goodput Maximization for Large Language Model Edge Inference: A Two-Phase Maskable PPO Approach
Xiaojing Chen, Qi Zhang, Wei Ni +2
eess.SYcs.AIarXiv:2608.25543v12026Revelation Control
Qinyou Wang
cs.LGcs.AIstat.MLarXiv:2608.23860v12026Summaries:한국어FAMPWQ: Fisher Information-based Adaptive Mixed Precision Weight Quantization for Effective LLM Inference
Gongwei Lee, Ji Liu, Juncheng Jia +1
cs.LGcs.AIcs.DCarXiv:2608.24945v12026Summaries:한국어Spending Scarce Confirmatory PET Measurements: Target-Aligned Validation in A4/LEARN
Eliuvish Han Cui
stat.APcs.AIcs.LGarXiv:2608.22223v12026Right-Sizing LLM-Agent Decomposition in VAT Determination: A Pilot Controlled Sweep
Pedro Santos
cs.MAcs.AIcs.SEarXiv:2608.23395v12026Names Can Hurt: Spotting Slopsquatting Risks Caused by Package Name Hallucinations in Local Coding LLMs
Akash Raj, Sargam Sahu
cs.CLcs.AIarXiv:2608.23897v12026CVE-SAI: Counterfactual Visual Evidence-Guided Selective Attribute Indexing for Risk-Controlled E-commerce Search
Xiaolong Sun, Qichao Wang, Hangyu Li +1
cs.AIcs.CVarXiv:2608.25023v12026CacheRouter: A Dual-Path Tool Routing Architecture with Cache-Preserving Main-Model Isolation for Long-Tail Tool Discovery
Donghui Zha, Lingwei Xu, Linxiao Wu +2
cs.AIarXiv:2608.22708v12026Correcting a learned physical invariant improves world-model rollouts
Richard Bao
cs.AIarXiv:2608.23526v12026Why Does Robustness Reduce Superposition?
Adam Elimadi
cs.LGcs.AIarXiv:2608.22155v12026STAGE: Stateful Translation to Agentic Graph Execution with Policy-Scoped Context and Deterministic Control
Mengxi Luo, Changjia Chen, An Cao +2
cs.AIarXiv:2608.22538v12026ExFold: Unified Expert Folding for Training-Free MoE Prefill-Decode Acceleration
Juntong Wu, Yifei Liu, Junyi Chen +6
cs.LGcs.AIarXiv:2608.24938v12026When Does Frequency Decomposition Benefit Physics-Informed Neural Networks? A Preliminary Ablation Study
Shubham Rai
cs.LGcs.AIarXiv:2608.24940v12026A Mathematical Theory of Interpretation: Rational Entropy, Spectral Readout, and Confusability as a Resource
Blake Reynolds
cs.ITcs.AIarXiv:2608.23892v12026Learning to Grade Efficiently: A Bandit-Driven Prompt-Selection Framework for Low-Cost LLM Essay Scoring
Olga Manakina, Igor Bogdanov
cs.LGcs.AIcs.CLarXiv:2608.23814v12026Discovering Cross-Language Reasoning Invariance in LLMs with Geometry-Invariant Sparse Autoencoders
Igor Bogdanov, Changcheng Huang
cs.LGcs.AIcs.CLarXiv:2608.23809v12026What Reaches Expert Review? Representation, Structural Screening, and Candidate-Form Dependence in AI-Assisted Item Development
Christopher Brooks
cs.CLcs.AIarXiv:2608.23766v12026Beyond Executable Models: The Pufibara Agent Harness and the Modelica Agent Workflow Benchmark for Physical System Modeling
Zizhe Wang
cs.SEcs.AIarXiv:2608.23653v12026Feedback That Backfires: Why Small Language Model Agents Repeat the Call They Just Watched Fail
Esmail Gumaan
cs.SEcs.AIarXiv:2608.23651v12026Retrieval-augmented generation vs. deterministic tax computation in multi-agent financial advisory: A 2x2 factorial experiment
Aryan Brar, Justin Du, Avery Lor +2
cs.AIarXiv:2608.23908v12026Semantic Overlays: Mitigating Prompt Injection with Annotations Beyond Tokens and Steering Vectors
Joshua Penman
cs.AIcs.CLcs.CRarXiv:2608.23873v12026SyPS: Measuring Sycophancy Prompt Sensitivity in Large Language Models
Lijia Huang, Yao Fu, Sihao Ren
cs.AIarXiv:2608.23837v12026When Names Cross Scripts: A Source-Grounded Benchmark for Historical Entity Reconciliation in the Mongol World
Xiang Chen, Zeyu Zhang
cs.CLcs.AIarXiv:2608.23507v12026Adversarial Entropy Inflation Against Gumbel-Based Inference Verification
Nikita Kezins
cs.CRcs.AIarXiv:2608.23375v12026Hypergraph Embedding Indexing for Efficient Dense Vector Retrieval
Kishore Konda
cs.IRcs.AIarXiv:2608.22980v12026Targeting the Attention Heads Behind Object Hallucination in LLaVA
Armaan Sandhu, Abhilasha Senapati, Hima Kammachi
cs.CVcs.AIarXiv:2608.24966v12026Drift Variation Autoencoder: Unifying Generation and Representation Learning through Conditional Posterior Flow Matching
Jiarui Cao
cs.LGcs.AIarXiv:2608.25138v12026Resource-Efficient Pruning for Transformer via Low-Rank Importance Estimation
Peng Liu, Huibing Zeng, Yiqun Zhang +2
cs.LGcs.AIcs.ETarXiv:2608.24973v12026The Invisible Editorial Layer: Formalizing Undisclosed Inference-Time Steering, Probability Placement, and the Attribution Problem in Deployed Language Models
Augusto Camargo
cs.AIcs.CLcs.CYarXiv:2608.24662v22026Unmatched Does Not Mean False: Incomplete Reference Sets Can Reverse Calibration Rankings in Open-Ended Theory-of-Mind Tracking
Zhexi Feng, Wuxi Chen, Bingrui Zhang
cs.CLcs.AIarXiv:2608.25654v12026Output Dilution: Redundant but Fragile Representations in MoE Models
Orion Reblitz-Richardson
cs.LGcs.AIcs.CLarXiv:2608.25231v12026Reading Is Not Using: Retrieval, Judgment, and the Design of AI Financial Research Workflows
Miao Liu, Zhizhe Liu
cs.CLcs.AIarXiv:2608.24842v12026Knowing When to Ask for Help: Bayesian Self-Escalation in Hierarchical LLM Agents
Nadeem Shaikh
cs.LGcs.AIstat.MLarXiv:2608.24087v12026Constrained Entity Selection under Partial Knowledge for LLM-Based Knowledge Graph QA
Emanuel Kitzelmann
cs.AIarXiv:2608.24824v12026Confident at the moment of action: belief miscalibration in LLM play under hidden information
Bhushan Kashinath Joshi
cs.AIcs.CLcs.LGarXiv:2608.24691v12026Do Recipes Have Personas? Characterizing and Generating Creator Style in Attributed Procedural Graphs
Lei Jiang
cs.AIarXiv:2608.24369v12026Benchmarking LLM Judges for Voice-Agent Evaluation: Reliability, Calibration, and Human Oversight
Anupam Purwar, Shashank Singh, Kritika Srivastava
cs.AIcs.ETarXiv:2608.24314v12026VideoHarness-RSI: Recursive Harness Self-Improvement for Long-Video Understanding with Frozen Vision-Language Models
Guoyang Xu, Hao Chen
cs.AIarXiv:2608.24302v12026Giraffe: A Mapping Architecture from Hidden Text Representations to Visual Embeddings for Efficient Graphic Design
Nejla Ghaboosi
cs.AIcs.LGarXiv:2608.23970v12026MMJailBench: A Factorized Benchmark for Disentangling Multimodal Jailbreak Vulnerabilities
Tianshi Wang, Jingsong Wang, Yafei Huang +3
cs.CRcs.AIcs.MMarXiv:2608.25490v12026RotDroid: Cross-Orientation State Equivalence Testing for Detecting GUI Rotation Bugs in Android Apps
Mengdi Qin, Bo Jiang
cs.SEcs.AIarXiv:2608.25425v12026A Tendon-Driven Five-Fingered Hand with Distributed Tactile Perception for Dexterous Manipulation
Huayang Chen, Longhui Qin
cs.ROcs.AIarXiv:2608.25547v12026PIVOT: A Multi-Trajectory Dataset and Testbed for Pose, Intrinsics, and Novel Viewpoint Evaluation in Real-World 3D Reconstruction
Mary Raymond
cs.CVcs.AIarXiv:2608.25401v12026Neither Precision Nor Architecture Alone: Controlled Tests of Failure Remedies for Physics-Informed Neural Networks
Jinyuan Zhang, Peng He, He Hu +2
cs.LGcs.AIarXiv:2608.25327v12026PonsRAG: A Pons-Inspired RAG Bridging Cognitive Islands for Coordinated Long Narrative Reasoning
Rongchen Zhao, Yu Chen, Juyuan Wang +4
cs.AIcs.CLarXiv:2608.25486v12026Escaping Low-Dimensional Overlap: Multi-Task Model Merging via High-Dimensional Sparse Disentanglement
Yihang Zhang, Shengke Sun, Junjie Wen +1
cs.LGcs.AIcs.CLarXiv:2608.25354v12026When Less Is More: An Empirical Study of Minimal Responses in Counseling Dialogues and the Behavior of LLMs
Zhiyang Qi
cs.CLcs.AIarXiv:2608.24080v12026Repair or Resample? Rethinking Failure Debugging in LLM Multi-Agent Systems
Zhongwen Luan, Xiaoyu Zhang, Ming Hu +3
cs.AIcs.SEarXiv:2608.25920v12026Summaries:简体中文A Statistical Audit of Physical AI Benchmark Redundancy
Zaruhi Navasardyan, Hrant Davtyan
cs.ROcs.AIarXiv:2608.25940v12026A Visual Dependence-Aware Framework for Multimodal Unsupervised Continual Post-Training
Kaichen Li, Zhilin Zhu, Jianhao Huang +7
cs.CVcs.AIarXiv:2608.26095v12026TAU-Agent: An Agentic Retrieval-Augmented Framework for Traffic Anomaly Understanding
Yuqiang Lin, Yan Shi, Sam Lockyer +5
cs.CVcs.AIarXiv:2608.25935v12026From General Agents to RCA Experts: A Self-Evolving Harness for Root Cause Analysis
Haiyu Huang, Jiewei Lyu, Zhihan Jiang +5
cs.SEcs.AIarXiv:2608.25661v12026Summaries:简体中文Choose Your Game Wisely: Measuring Game-Theoretic Structures in Real-World Vehicle Interactions
Yueyuan Li, Rongcheng Nie, Weijie Xi +4
cs.AIcs.ROarXiv:2608.25917v12026Gating Before Commitment: Anticipating Intent Divergence to Prevent Post-Interaction Decision Failures in Autonomous Driving
Cong Xu, Ravi Sankar
cs.ROcs.AIarXiv:2608.26074v12026SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks
Alexander Robey, Eric Wong, Hamed Hassani +1
cs.LGcs.AIstat.MLarXiv:2310.03684v42023FRAME: separating sampling variation from representational cause in medical imaging fairness
Mahshad Lotfinia, Daniel Truhn, Andreas Maier +1
cs.CVcs.AIcs.LGarXiv:2608.25981v12026