Software Engineering

Papers filed under cs.SE on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

361 to 420 of 1,405

  1. AICD Bench: A Challenging Benchmark for AI-Generated Code Detection

    Daniil Orel, Dilshod Azizov, Indraneil Paul +3

    cs.LGcs.SEarXiv:2602.02079v12026
  2. Formal Scenario-Based Testing of Autonomous Vehicles: From Simulation to the Real World

    Daniel J. Fremont, Edward Kim, Yash Vardhan Pant +7

    eess.SYcs.LGcs.LOarXiv:2003.07739v22020
  3. BenchGuard: Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks

    Xinming Tu, Tianze Wang, Yingzhou +4

    cs.CLcs.AIcs.SEarXiv:2604.24955v12026
  4. Building the Truman Show: A TrustZone-Based Framework for Lightweight Out-of-band Kernel Security Monitoring

    Zhenling Duan, Pan Dong, Renshuang Jiang +2

    cs.CRcs.SEarXiv:2608.29758v12026
  5. Cost-Effective Repository Exploration for Agentic Issue Localization

    Mohammad Nour Al Awad, Sergey Ivanov

    cs.SEcs.AIarXiv:2608.29675v12026
  6. A Survey of Learning-based Automated Program Repair

    Quanjun Zhang, Chunrong Fang, Yuxiang Ma +2

    cs.SEarXiv:2301.03270v32023
  7. Model Context Protocol Threat Modeling and Analyzing Vulnerabilities to Prompt Injection with Tool Poisoning

    Charoes Huang, Xin Huang, Ngoc Phu Tran +1

    cs.CRcs.SEarXiv:2603.22489v12026
  8. From Technical Debt to Cognitive and Intent Debt: Rethinking Software Health in the Age of AI

    Margaret-Anne Storey

    cs.SEarXiv:2603.22106v42026
  9. Debt Behind the AI Boom: A Large-Scale Empirical Study of AI-Generated Code in the Wild

    Yue Liu, Ratnadira Widyasari, Yanjie Zhao +3

    cs.SEarXiv:2603.28592v22026
  10. Automatically Discovering, Reporting and Reproducing Android Application Crashes

    Kevin Moran, Mario Linares-Vásquez, Carlos Bernal-Cárdenas +2

    cs.SEarXiv:1706.01130v12017
  11. Exploring Quantum Software Testing Across Research and Practice: Emerging Results from a Multivocal Literature Review

    Rodolfo Gil-Pereira, Ronnie de Souza Santos, Cleyton Magalhaes +1

    cs.SEarXiv:2609.00354v12026
  12. Clawdrain: Exploiting Tool-Calling Chains for Stealthy Token Exhaustion in OpenClaw Agents

    Ben Dong, Hui Feng, Qian Wang

    cs.CRcs.SEarXiv:2603.00902v12026
  13. Auditing Anonymous AI Models: A Four-Stage Protocol for Black-Box Identity Verification

    Yisen Xi

    cs.SEcs.AIcs.CRarXiv:2608.31142v12026
  14. EffiSkill: Agent Skill Based Automated Code Efficiency Optimization

    Zimu Wang, Yuling Shi, Mengfan Li +4

    cs.SEcs.CLarXiv:2603.27850v12026
  15. Evaluating Tiny Recursive Models Across Training for Code Generation

    Anjani Sirivella, Aanisha Newaz, Glaucia Melo

    cs.AIcs.LGcs.SEarXiv:2608.29376v12026
  16. SkillProbe: Security Auditing for Emerging Agent Skill Marketplaces via Multi-Agent Collaboration

    Zihan Guo, Zhiyu Chen, Xiaohang Nie +3

    cs.CRcs.SEarXiv:2603.21019v12026
  17. Pythia: AI-assisted Code Completion System

    Alexey Svyatkovskiy, Ying Zhao, Shengyu Fu +1

    cs.SEcs.LGarXiv:1912.00742v12019
  18. Structurally Aligned Subtask-Level Memory for Software Engineering Agents

    Kangning Shen, Jingyuan Zhang, Chenxi Sun +2

    cs.SEcs.AIarXiv:2602.21611v12026
  19. Real Money, Fake Models: Deceptive Model Claims in Shadow APIs

    Yage Zhang, Yukun Jiang, Zeyuan Chen +3

    cs.CRcs.AIcs.SEarXiv:2603.01919v22026
  20. SWE-bench Goes Live!

    Linghao Zhang, Shilin He, Chaoyun Zhang +12

    cs.SEcs.AIcs.CLarXiv:2505.23419v22025
  21. ARGUS: Defending LLM Agents Against Context-Aware Prompt Injection

    Shihao Weng, Yang Feng, Jinrui Zhang +3

    cs.CRcs.SEarXiv:2605.03378v22026
  22. AgentTrace: A Structured Logging Framework for Agent System Observability

    Adam AlSayyad, Kelvin Yuxiang Huang, Richik Pal

    cs.SEcs.AIarXiv:2602.10133v12026
  23. On the Impact of AGENTS.md Files on the Efficiency of AI Coding Agents

    Jai Lal Lulla, Seyedmoein Mohsenimofidi, Matthias Galster +3

    cs.SEcs.AIcs.ETarXiv:2601.20404v22026
  24. Fault Localization with Code Coverage Representation Learning

    Yi Li, Shaohua Wang, Tien N. Nguyen

    cs.SEarXiv:2103.00270v12021
  25. From Silicon to Boot Code: Extending Automated Program Repair to Firmware-Layer Security Workarounds

    Maisha Mastora, Dean Sullivan

    cs.SEarXiv:2609.01769v12026
  26. Modelstamp: Pre-Deserialization Verification of Machine-Learning Artifacts and Runtime Environment State

    Anagha Dhekne

    cs.SEcs.CRarXiv:2609.01781v12026
  27. Solver-Aided Verification of Policy Compliance in Tool-Augmented LLM Agents

    Cailin Winston, Claris Winston, René Just

    cs.SEcs.AIarXiv:2603.20449v12026
  28. MedPerf: Open Benchmarking Platform for Medical Artificial Intelligence using Federated Evaluation

    Alexandros Karargyris, Renato Umeton, Micah J. Sheller +39

    cs.LGcs.DCcs.PFarXiv:2110.01406v32021
  29. When Safe Skills Collide: Measuring Compositional Risk in Agent Skill Ecosystems

    Su Wang, Pin Qian, Yihang Chen +6

    cs.SEcs.AIcs.CRarXiv:2606.00448v12026
  30. GUI Agents for Continual Game Generation

    Yixu Huang, Bo Li, Na Li +8

    cs.SEcs.AIcs.CVarXiv:2605.28258v12026
  31. Software Model Checking via Large-Block Encoding

    Dirk Beyer, Alessandro Cimatti, Alberto Griggio +2

    cs.SEcs.PLarXiv:0904.4709v12009
  32. Towards Effective Bug Triage with Towards Effective Bug Triage with Software Data Reduction Techniques

    Jifeng Xuan, He Jiang, Yan Hu +4

    cs.SEarXiv:1704.04761v12017
  33. A Software Engineering Perspective on Engineering Machine Learning Systems: State of the Art and Challenges

    Görkem Giray

    cs.SEcs.LGarXiv:2012.07919v32020
  34. Humans are Missing from AI Coding Agent Research

    Zora Z. Wang, John Yang, Kilian Lieret +10

    cs.HCcs.AIcs.SEarXiv:2608.12355v12026
  35. Inside the Scaffold: A Source-Code Taxonomy of Coding Agent Architectures

    Benjamin Rombaut

    cs.SEcs.AIcs.ETarXiv:2604.03515v22026
  36. T(r)opical Islands: Visualizing & Understanding Socio-Technical Artifacts

    Adam Štěpánek, Marco Raglianti, Vít Rusňák +3

    cs.SEarXiv:2609.05254v12026
  37. Stop Comparing LLM Agents Without Disclosing the Harness

    Yunbei Zhang, Janet Wang, Yingqiang Ge +3

    cs.AIcs.SEarXiv:2605.23950v12026
  38. An Empirical Study on Learning Paths and Gender Dynamics in Scrum Master Roles

    Manuela Petrescu, Paul Razvan Petrescu

    cs.SEcs.CYarXiv:2609.05186v12026
  39. ASTRA - Agentic System for Ticket Resolution and Analysis

    Shashidhar Reddy Javaji, Mohamed Trabelsi, Jin Cao +1

    cs.MAcs.AIcs.IRarXiv:2608.28790v12026
  40. SkillOps: Managing LLM Agent Skill Libraries as Self-Maintaining Software Ecosystems

    Hongji Pu, Xinyuan Song, Liang Zhao

    cs.SEcs.MAarXiv:2605.13716v12026
  41. Rubric Is All You Need: Enhancing LLM-based Code Evaluation With Question-Specific Rubrics

    Aditya Pathak, Rachit Gandhi, Vaibhav Uttam +11

    cs.SEcs.AIarXiv:2503.23989v32025
  42. Automated Fixing of Programs with Contracts

    Yu Pei, Carlo A. Furia, Martin Nordio +3

    cs.SEarXiv:1403.1117v32014
  43. Unsafer in Many Turns: Benchmarking and Defending Multi-Turn Safety Risks in Tool-Using Agents

    Xu Li, Simon Yu, Minzhou Pan +5

    cs.CRcs.AIcs.CLarXiv:2602.13379v22026
  44. Flow-of-Action: SOP Enhanced LLM-Based Multi-Agent System for Root Cause Analysis

    Changhua Pei, Zexin Wang, Fengrui Liu +9

    cs.SEarXiv:2502.08224v12025
  45. Evaluating Agent-based Program Repair at Google

    Pat Rondon, Renyao Wei, José Cambronero +5

    cs.SEcs.AIarXiv:2501.07531v12025
  46. Out of the BLEU: how should we assess quality of the Code Generation models?

    Mikhail Evtikhiev, Egor Bogomolov, Yaroslav Sokolov +1

    cs.SEcs.LGarXiv:2208.03133v22022
  47. Revisiting the Core Ontology and Problem in Requirements Engineering

    Ivan Jureta, John Mylopoulos, Stephane Faulkner

    cs.SEarXiv:0811.4364v12008
  48. Defining Smart Contract Defects on Ethereum

    Jiachi Chen, Xin Xia, David Lo +3

    cs.SEarXiv:1905.01467v32019
  49. Code to Think, Think to Code: A Survey on Code-Enhanced Reasoning and Reasoning-Driven Code Intelligence in LLMs

    Dayu Yang, Tianyang Liu, Daoan Zhang +8

    cs.CLcs.AIcs.LGarXiv:2502.19411v12025
  50. SWE-MeM: Learning Adaptive Memory Management for Long-Horizon Coding Agents

    Shuzheng Gao, Wenhao Zeng, Zhaojian Yu +5

    cs.SEcs.AIarXiv:2606.28434v12026
  51. Agent-Driven Verification of Memory Safety for liblzma Decoder Components with VST

    Prokhor Shlyakhtun, Alexander Gryzlov, Vladimir Kukharenko +5

    cs.SEarXiv:2608.29716v12026
  52. InteractBench: Benchmarking LLMs on Competitive Programming under Unrevealed Information

    Jiaze Li, Aocheng Shen, Bing Liu +4

    cs.SEcs.AIcs.CLarXiv:2608.29632v12026
  53. Rust's Type Checker Implementation Is Unsound: An Empirical Study on Soundness Bugs in rustc

    Yusung Sim, Sukyoung Ryu, Jaemin Hong

    cs.SEcs.PLarXiv:2608.28713v12026
  54. trajectory-judge: What Outcome-Only LLM Judges Miss on Agent Trajectories

    Hadi Mohammadi

    cs.CLcs.AIcs.SEarXiv:2609.00038v12026
  55. Towards Human-Bot Collaborative Software Architecting with ChatGPT

    Aakash Ahmad, Muhammad Waseem, Peng Liang +3

    cs.SEcs.AIarXiv:2302.14600v12023
  56. SkillCraft: Can LLM Agents Learn to Use Tools Skillfully?

    Shiqi Chen, Jingze Gai, Ruochen Zhou +13

    cs.CLcs.SEarXiv:2603.00718v22026
  57. FlowCheck: Helping End-Users Specify and Verify Intent in Vibe-Coded Web Apps

    Reya Vir, Lydia Chilton, Zhuo Zhang +1

    cs.SEcs.HCcs.PLarXiv:2608.28880v12026
  58. Exploring the Responses of Large Language Models to Beginner Programmers' Help Requests

    Arto Hellas, Juho Leinonen, Sami Sarsa +3

    cs.CYcs.AIcs.CLarXiv:2306.05715v12023
  59. Formal Analysis and Supply Chain Security for Agentic AI Skills

    Varun Pratap Bhardwaj

    cs.CRcs.AIcs.SEarXiv:2603.00195v22026
  60. Understanding Flaky Tests: The Developer's Perspective

    Moritz Eck, Fabio Palomba, Marco Castelluccio +1

    cs.SEarXiv:1907.01466v12019