Artificial Intelligence

Papers filed under cs.AI on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

10,501 to 10,560 of 15,398

  1. Toxicity in ChatGPT: Analyzing Persona-assigned Language Models

    Ameet Deshpande, Vishvak Murahari, Tanmay Rajpurohit +2

    cs.CLcs.AIcs.LGarXiv:2304.05335v12023
  2. Language Models can Solve Computer Tasks

    Geunwoo Kim, Pierre Baldi, Stephen McAleer

    cs.CLcs.AIcs.HCarXiv:2303.17491v32023
  3. AI and the Everything in the Whole Wide World Benchmark

    Inioluwa Deborah Raji, Emily M. Bender, Amandalynne Paullada +2

    cs.LGcs.AIcs.PFarXiv:2111.15366v12021
  4. DL3DV-10K: A Large-Scale Scene Dataset for Deep Learning-based 3D Vision

    Lu Ling, Yichen Sheng, Zhi Tu +17

    cs.CVcs.AIarXiv:2312.16256v22023
  5. Goodput Maximization for Large Language Model Edge Inference: A Two-Phase Maskable PPO Approach

    Xiaojing Chen, Qi Zhang, Wei Ni +2

    eess.SYcs.AIarXiv:2608.25543v12026
  6. Revelation Control

    Qinyou Wang

    cs.LGcs.AIstat.MLarXiv:2608.23860v12026
    Summaries:한국어
  7. FAMPWQ: Fisher Information-based Adaptive Mixed Precision Weight Quantization for Effective LLM Inference

    Gongwei Lee, Ji Liu, Juncheng Jia +1

    cs.LGcs.AIcs.DCarXiv:2608.24945v12026
    Summaries:한국어
  8. Spending Scarce Confirmatory PET Measurements: Target-Aligned Validation in A4/LEARN

    Eliuvish Han Cui

    stat.APcs.AIcs.LGarXiv:2608.22223v12026
  9. Right-Sizing LLM-Agent Decomposition in VAT Determination: A Pilot Controlled Sweep

    Pedro Santos

    cs.MAcs.AIcs.SEarXiv:2608.23395v12026
  10. Names Can Hurt: Spotting Slopsquatting Risks Caused by Package Name Hallucinations in Local Coding LLMs

    Akash Raj, Sargam Sahu

    cs.CLcs.AIarXiv:2608.23897v12026
  11. CVE-SAI: Counterfactual Visual Evidence-Guided Selective Attribute Indexing for Risk-Controlled E-commerce Search

    Xiaolong Sun, Qichao Wang, Hangyu Li +1

    cs.AIcs.CVarXiv:2608.25023v12026
  12. CacheRouter: A Dual-Path Tool Routing Architecture with Cache-Preserving Main-Model Isolation for Long-Tail Tool Discovery

    Donghui Zha, Lingwei Xu, Linxiao Wu +2

    cs.AIarXiv:2608.22708v12026
  13. Correcting a learned physical invariant improves world-model rollouts

    Richard Bao

    cs.AIarXiv:2608.23526v12026
  14. Why Does Robustness Reduce Superposition?

    Adam Elimadi

    cs.LGcs.AIarXiv:2608.22155v12026
  15. STAGE: Stateful Translation to Agentic Graph Execution with Policy-Scoped Context and Deterministic Control

    Mengxi Luo, Changjia Chen, An Cao +2

    cs.AIarXiv:2608.22538v12026
  16. ExFold: Unified Expert Folding for Training-Free MoE Prefill-Decode Acceleration

    Juntong Wu, Yifei Liu, Junyi Chen +6

    cs.LGcs.AIarXiv:2608.24938v12026
  17. When Does Frequency Decomposition Benefit Physics-Informed Neural Networks? A Preliminary Ablation Study

    Shubham Rai

    cs.LGcs.AIarXiv:2608.24940v12026
  18. A Mathematical Theory of Interpretation: Rational Entropy, Spectral Readout, and Confusability as a Resource

    Blake Reynolds

    cs.ITcs.AIarXiv:2608.23892v12026
  19. Learning to Grade Efficiently: A Bandit-Driven Prompt-Selection Framework for Low-Cost LLM Essay Scoring

    Olga Manakina, Igor Bogdanov

    cs.LGcs.AIcs.CLarXiv:2608.23814v12026
  20. Discovering Cross-Language Reasoning Invariance in LLMs with Geometry-Invariant Sparse Autoencoders

    Igor Bogdanov, Changcheng Huang

    cs.LGcs.AIcs.CLarXiv:2608.23809v12026
  21. What Reaches Expert Review? Representation, Structural Screening, and Candidate-Form Dependence in AI-Assisted Item Development

    Christopher Brooks

    cs.CLcs.AIarXiv:2608.23766v12026
  22. Beyond Executable Models: The Pufibara Agent Harness and the Modelica Agent Workflow Benchmark for Physical System Modeling

    Zizhe Wang

    cs.SEcs.AIarXiv:2608.23653v12026
  23. Feedback That Backfires: Why Small Language Model Agents Repeat the Call They Just Watched Fail

    Esmail Gumaan

    cs.SEcs.AIarXiv:2608.23651v12026
  24. Retrieval-augmented generation vs. deterministic tax computation in multi-agent financial advisory: A 2x2 factorial experiment

    Aryan Brar, Justin Du, Avery Lor +2

    cs.AIarXiv:2608.23908v12026
  25. Semantic Overlays: Mitigating Prompt Injection with Annotations Beyond Tokens and Steering Vectors

    Joshua Penman

    cs.AIcs.CLcs.CRarXiv:2608.23873v12026
  26. SyPS: Measuring Sycophancy Prompt Sensitivity in Large Language Models

    Lijia Huang, Yao Fu, Sihao Ren

    cs.AIarXiv:2608.23837v12026
  27. When Names Cross Scripts: A Source-Grounded Benchmark for Historical Entity Reconciliation in the Mongol World

    Xiang Chen, Zeyu Zhang

    cs.CLcs.AIarXiv:2608.23507v12026
  28. Adversarial Entropy Inflation Against Gumbel-Based Inference Verification

    Nikita Kezins

    cs.CRcs.AIarXiv:2608.23375v12026
  29. Hypergraph Embedding Indexing for Efficient Dense Vector Retrieval

    Kishore Konda

    cs.IRcs.AIarXiv:2608.22980v12026
  30. Targeting the Attention Heads Behind Object Hallucination in LLaVA

    Armaan Sandhu, Abhilasha Senapati, Hima Kammachi

    cs.CVcs.AIarXiv:2608.24966v12026
  31. Drift Variation Autoencoder: Unifying Generation and Representation Learning through Conditional Posterior Flow Matching

    Jiarui Cao

    cs.LGcs.AIarXiv:2608.25138v12026
  32. Resource-Efficient Pruning for Transformer via Low-Rank Importance Estimation

    Peng Liu, Huibing Zeng, Yiqun Zhang +2

    cs.LGcs.AIcs.ETarXiv:2608.24973v12026
  33. The Invisible Editorial Layer: Formalizing Undisclosed Inference-Time Steering, Probability Placement, and the Attribution Problem in Deployed Language Models

    Augusto Camargo

    cs.AIcs.CLcs.CYarXiv:2608.24662v22026
  34. Unmatched Does Not Mean False: Incomplete Reference Sets Can Reverse Calibration Rankings in Open-Ended Theory-of-Mind Tracking

    Zhexi Feng, Wuxi Chen, Bingrui Zhang

    cs.CLcs.AIarXiv:2608.25654v12026
  35. Output Dilution: Redundant but Fragile Representations in MoE Models

    Orion Reblitz-Richardson

    cs.LGcs.AIcs.CLarXiv:2608.25231v12026
  36. Reading Is Not Using: Retrieval, Judgment, and the Design of AI Financial Research Workflows

    Miao Liu, Zhizhe Liu

    cs.CLcs.AIarXiv:2608.24842v12026
  37. Knowing When to Ask for Help: Bayesian Self-Escalation in Hierarchical LLM Agents

    Nadeem Shaikh

    cs.LGcs.AIstat.MLarXiv:2608.24087v12026
  38. Constrained Entity Selection under Partial Knowledge for LLM-Based Knowledge Graph QA

    Emanuel Kitzelmann

    cs.AIarXiv:2608.24824v12026
  39. Confident at the moment of action: belief miscalibration in LLM play under hidden information

    Bhushan Kashinath Joshi

    cs.AIcs.CLcs.LGarXiv:2608.24691v12026
  40. Do Recipes Have Personas? Characterizing and Generating Creator Style in Attributed Procedural Graphs

    Lei Jiang

    cs.AIarXiv:2608.24369v12026
  41. Benchmarking LLM Judges for Voice-Agent Evaluation: Reliability, Calibration, and Human Oversight

    Anupam Purwar, Shashank Singh, Kritika Srivastava

    cs.AIcs.ETarXiv:2608.24314v12026
  42. VideoHarness-RSI: Recursive Harness Self-Improvement for Long-Video Understanding with Frozen Vision-Language Models

    Guoyang Xu, Hao Chen

    cs.AIarXiv:2608.24302v12026
  43. Giraffe: A Mapping Architecture from Hidden Text Representations to Visual Embeddings for Efficient Graphic Design

    Nejla Ghaboosi

    cs.AIcs.LGarXiv:2608.23970v12026
  44. MMJailBench: A Factorized Benchmark for Disentangling Multimodal Jailbreak Vulnerabilities

    Tianshi Wang, Jingsong Wang, Yafei Huang +3

    cs.CRcs.AIcs.MMarXiv:2608.25490v12026
  45. RotDroid: Cross-Orientation State Equivalence Testing for Detecting GUI Rotation Bugs in Android Apps

    Mengdi Qin, Bo Jiang

    cs.SEcs.AIarXiv:2608.25425v12026
  46. A Tendon-Driven Five-Fingered Hand with Distributed Tactile Perception for Dexterous Manipulation

    Huayang Chen, Longhui Qin

    cs.ROcs.AIarXiv:2608.25547v12026
  47. PIVOT: A Multi-Trajectory Dataset and Testbed for Pose, Intrinsics, and Novel Viewpoint Evaluation in Real-World 3D Reconstruction

    Mary Raymond

    cs.CVcs.AIarXiv:2608.25401v12026
  48. Neither Precision Nor Architecture Alone: Controlled Tests of Failure Remedies for Physics-Informed Neural Networks

    Jinyuan Zhang, Peng He, He Hu +2

    cs.LGcs.AIarXiv:2608.25327v12026
  49. PonsRAG: A Pons-Inspired RAG Bridging Cognitive Islands for Coordinated Long Narrative Reasoning

    Rongchen Zhao, Yu Chen, Juyuan Wang +4

    cs.AIcs.CLarXiv:2608.25486v12026
  50. Escaping Low-Dimensional Overlap: Multi-Task Model Merging via High-Dimensional Sparse Disentanglement

    Yihang Zhang, Shengke Sun, Junjie Wen +1

    cs.LGcs.AIcs.CLarXiv:2608.25354v12026
  51. When Less Is More: An Empirical Study of Minimal Responses in Counseling Dialogues and the Behavior of LLMs

    Zhiyang Qi

    cs.CLcs.AIarXiv:2608.24080v12026
  52. Repair or Resample? Rethinking Failure Debugging in LLM Multi-Agent Systems

    Zhongwen Luan, Xiaoyu Zhang, Ming Hu +3

    cs.AIcs.SEarXiv:2608.25920v12026
    Summaries:简体中文
  53. A Statistical Audit of Physical AI Benchmark Redundancy

    Zaruhi Navasardyan, Hrant Davtyan

    cs.ROcs.AIarXiv:2608.25940v12026
  54. A Visual Dependence-Aware Framework for Multimodal Unsupervised Continual Post-Training

    Kaichen Li, Zhilin Zhu, Jianhao Huang +7

    cs.CVcs.AIarXiv:2608.26095v12026
  55. TAU-Agent: An Agentic Retrieval-Augmented Framework for Traffic Anomaly Understanding

    Yuqiang Lin, Yan Shi, Sam Lockyer +5

    cs.CVcs.AIarXiv:2608.25935v12026
  56. From General Agents to RCA Experts: A Self-Evolving Harness for Root Cause Analysis

    Haiyu Huang, Jiewei Lyu, Zhihan Jiang +5

    cs.SEcs.AIarXiv:2608.25661v12026
    Summaries:简体中文
  57. Choose Your Game Wisely: Measuring Game-Theoretic Structures in Real-World Vehicle Interactions

    Yueyuan Li, Rongcheng Nie, Weijie Xi +4

    cs.AIcs.ROarXiv:2608.25917v12026
  58. Gating Before Commitment: Anticipating Intent Divergence to Prevent Post-Interaction Decision Failures in Autonomous Driving

    Cong Xu, Ravi Sankar

    cs.ROcs.AIarXiv:2608.26074v12026
  59. SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks

    Alexander Robey, Eric Wong, Hamed Hassani +1

    cs.LGcs.AIstat.MLarXiv:2310.03684v42023
  60. FRAME: separating sampling variation from representational cause in medical imaging fairness

    Mahshad Lotfinia, Daniel Truhn, Andreas Maier +1

    cs.CVcs.AIcs.LGarXiv:2608.25981v12026