Software Engineering

Papers filed under cs.SE on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

301 to 360 of 1,389

  1. Intent Formalization: A Grand Challenge for Reliable Coding in the Age of AI Agents

    Shuvendu K. Lahiri

    cs.SEcs.AIcs.PLarXiv:2603.17150v12026
  2. RealSec-bench: A Benchmark for Evaluating Secure Code Generation in Real-World Repositories

    Yanlin Wang, Ziyao Zhang, Chong Wang +5

    cs.CRcs.SEarXiv:2601.22706v12026
  3. DafnyPro: LLM-Assisted Automated Verification for Dafny Programs

    Debangshu Banerjee, Olivier Bouissou, Stefan Zetzsche

    cs.SEarXiv:2601.05385v12026
  4. Programming by Chat: A Large-Scale Behavioral Analysis of 11,579 Real-World AI-Assisted IDE Sessions

    Ningzhi Tang, Chaoran Chen, Zihan Fang +6

    cs.SEcs.HCarXiv:2604.00436v22026
  5. PlotChain: Deterministic Checkpointed Evaluation of Multimodal LLMs on Engineering Plot Reading

    Mayank Ravishankara

    cs.AIcs.SEarXiv:2602.13232v12026
  6. 6-Layer Model for a Structured Description and Categorization of Urban Traffic and Environment

    Maike Scholtes, Lukas Westhofen, Lara Ruth Turner +11

    cs.OHcs.AIcs.SEarXiv:2012.06319v22020
  7. Edge Impulse: An MLOps Platform for Tiny Machine Learning

    Shawn Hymel, Colby Banbury, Daniel Situnayake +13

    cs.DCcs.LGcs.SEarXiv:2212.03332v32022
  8. Where Do AI Coding Agents Fail? An Empirical Study of Failed Agentic Pull Requests in GitHub

    Ramtin Ehsani, Sakshi Pathak, Shriya Rawal +3

    cs.SEcs.AIarXiv:2601.15195v12026
  9. Immersion in the GitHub Universe: Scaling Coding Agents to Mastery

    Jiale Zhao, Guoxin Chen, Fanzhe Meng +11

    cs.SEarXiv:2602.09892v42026
  10. How does information access affect LLM monitors' ability to detect sabotage?

    Rauno Arike, Raja Mehta Moreno, Rohan Subramani +2

    cs.AIcs.SEarXiv:2601.21112v22026
  11. What do we know about software development in startups?

    Carmine Giardino, Michael Unterkalmsteiner, Nicolò Paternoster +2

    cs.SEarXiv:2307.13707v12023
  12. The History Is the Detector: Executing CVE Patch History, End-to-End

    Qiushi Wu, Kevin Eykholt, Youngja Park +4

    cs.CRcs.AIcs.SEarXiv:2609.05335v12026
  13. Compressing Code Context for LLM-based Issue Resolution

    Haoxiang Jia, Earl T. Barr, Sergey Mechtaev

    cs.SEarXiv:2603.28119v12026
  14. Goedel-Code-Prover: Hierarchical Proof Search for Open State-of-the-Art Code Verification

    Zenan Li, Ziran Yang, Deyuan He +7

    cs.SEcs.AIarXiv:2603.19329v32026
  15. Evaluating and Improving Automated Repository-Level Rust Issue Resolution with LLM-based Agents

    Jiahong Xiang, Wenxiao He, Xihua Wang +2

    cs.SEarXiv:2602.22764v12026
  16. ARIA - An Agentic Framework for Autonomous Testing of Infotainment Systems

    António Azevedo, Bruno Lima, João Pascoal Faria

    cs.SEcs.AIarXiv:2609.04913v12026
  17. Better Understanding, Better Fixes? A Study of Hallucination in LLM-based Automated Program Repair

    Xuemeng Cai, Jiakun Liu, Linhan Yang +2

    cs.SEcs.AIarXiv:2609.04909v12026
  18. ContractSkill: Repairable Contract-Based Skills for Multimodal Web Agents

    Zijian Lu, Yiping Zuo, Yupeng Nie +4

    cs.SEcs.AIarXiv:2603.20340v32026
  19. Characterizing Faults in Agentic AI: A Taxonomy of Types, Symptoms, and Root Causes

    Mehil B Shah, Mohammad Mehdi Morovati, Mohammad Masudur Rahman +1

    cs.SEarXiv:2603.06847v22026
  20. Beyond Code Generation: Reliability, Verification, and Cost Economics in the Agentic Software Development Lifecycle

    Happy Bhati

    cs.SEcs.AIarXiv:2609.04681v12026
  21. In Line with Context: Repository-Level Code Generation via Context Inlining

    Chao Hu, Wenhao Zeng, Yuling Shi +2

    cs.SEcs.AIarXiv:2601.00376v32026
  22. Dynamic Adaptation of the LLM Context for Generating Routines with Coupled Semantics

    Gnaneswar Villuri, Hashmath Shaik, Alex Doboli

    cs.SEcs.AIarXiv:2609.04570v12026
  23. Building a research-software catalog with a coding agent: from hackathon prototype to public deployment

    Kazuyoshi Yoshimi, Satoshi Terasaki, Gotai Yamada

    cs.SEcs.AIcs.CYarXiv:2609.04711v12026
  24. Advancing Requirements Engineering through Generative AI: Assessing the Role of LLMs

    Chetan Arora, John Grundy, Mohamed Abdelrazek

    cs.SEarXiv:2310.13976v22023
  25. Software Development in Startup Companies: The Greenfield Startup Model

    Carmine Giardino, Nicolò Paternoster, Michael Unterkalmsteiner +2

    cs.SEarXiv:2308.09438v12023
  26. Trace2Tower: Transition-Aware EigenTrace Induction of Multi-Level Skills for LLM Agents

    Jiazheng Sun, Boyu Yang, Binhao Yuan +2

    cs.AIcs.SEarXiv:2609.05261v12026
  27. Cloud-OpsBench: A Reproducible Benchmark for Agentic Root Cause Analysis in Cloud Systems

    Yilun Wang, Guangba Yu, Haiyu Huang +4

    cs.SEarXiv:2603.00468v22026
  28. SWE-ABS: Adversarial Benchmark Strengthening Exposes Inflated Success Rates on Test-based Benchmark

    Boxi Yu, Yang Cao, Yuzhong Zhang +9

    cs.SEarXiv:2603.00520v12026
  29. Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents?

    Thibaud Gloaguen, Niels Mündler, Mark Müller +2

    cs.SEcs.AIarXiv:2602.11988v22026
  30. Library Drift: Diagnosing and Fixing a Silent Failure Mode in Self-Evolving LLM Skill Libraries

    Xing Zhang, Yanwei Cui, Guanghui Wang +4

    cs.AIcs.CLcs.SEarXiv:2605.19576v32026
  31. Is GitHub's Copilot as Bad as Humans at Introducing Vulnerabilities in Code?

    Owura Asare, Meiyappan Nagappan, N. Asokan

    cs.SEcs.CRarXiv:2204.04741v52022
  32. Recommender System for Online Dating Service

    Lukas Brozovsky, Vaclav Petricek

    cs.IRcs.SEarXiv:cs/0703042v12007
  33. AI IDEs or Autonomous Agents? Measuring the Impact of Coding Agents on Software Development

    Shyam Agarwal, Hao He, Bogdan Vasilescu

    cs.SEarXiv:2601.13597v22026
  34. Deterministic LLM Inference Across GPU Kernels: Power-of-Two INT8 Quantization Scales and the Limits of Tolerance-Based Conformance

    Teng-Ruei Chen

    cs.LGcs.SEarXiv:2609.00363v12026
  35. AgentWorm: Self-Propagating Attacks Across LLM Agent Ecosystems

    Yihao Zhang, Zeming Wei, Xiaokun Luan +7

    cs.CRcs.AIcs.LGarXiv:2603.15727v32026
  36. Large Language Model for Vulnerability Detection: Emerging Results and Future Directions

    Xin Zhou, Ting Zhang, David Lo

    cs.SEarXiv:2401.15468v12024
  37. The Web-CLI: Verifiable Privacy for Tools, Models, and Inference Engines in the Browser

    Tejaswi Gowda

    cs.HCcs.CRcs.SEarXiv:2608.28950v22026
  38. SWE-Refactor: A Repository-Level Benchmark for Real-World LLM-Based Code Refactoring

    Yisen Xu, Jinqiu Yang, Tse-Hsun +1

    cs.SEarXiv:2602.03712v12026
  39. The reach of a verification tool decides its value: A controlled study of verification surface, artifact quality, and cost in AI coding agents

    Achint Mehta

    cs.SEcs.AIarXiv:2608.28795v12026
  40. Spec-Driven Development:From Code to Contract in the Age of AI Coding Assistants

    Deepak Babu Piskala

    cs.SEcs.AIarXiv:2602.00180v12026
  41. Wink: Recovering from Misbehaviors in Coding Agents

    Rahul Nanda, Chandra Maddila, Smriti Jha +3

    cs.SEcs.AIcs.HCarXiv:2602.17037v22026
  42. Large Language Models for Code Generation: A Comprehensive Survey of Challenges, Techniques, Evaluation, and Applications

    Nam Huynh, Beiyu Lin

    cs.SEcs.LGarXiv:2503.01245v22025
  43. From Flat Logs to Causal Graphs: Hierarchical Failure Attribution for LLM-based Multi-Agent Systems

    Yawen Wang, Wenjie Wu, Junjie Wang +1

    cs.AIcs.SEarXiv:2602.23701v12026
  44. Promises, Perils, and (Timely) Heuristics for Mining Coding Agent Activity

    Romain Robbes, Théo Matricon, Thomas Degueule +2

    cs.SEarXiv:2601.18345v12026
  45. AICD Bench: A Challenging Benchmark for AI-Generated Code Detection

    Daniil Orel, Dilshod Azizov, Indraneil Paul +3

    cs.LGcs.SEarXiv:2602.02079v12026
  46. Formal Scenario-Based Testing of Autonomous Vehicles: From Simulation to the Real World

    Daniel J. Fremont, Edward Kim, Yash Vardhan Pant +7

    eess.SYcs.LGcs.LOarXiv:2003.07739v22020
  47. BenchGuard: Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks

    Xinming Tu, Tianze Wang, Yingzhou +4

    cs.CLcs.AIcs.SEarXiv:2604.24955v12026
  48. Building the Truman Show: A TrustZone-Based Framework for Lightweight Out-of-band Kernel Security Monitoring

    Zhenling Duan, Pan Dong, Renshuang Jiang +2

    cs.CRcs.SEarXiv:2608.29758v12026
  49. Cost-Effective Repository Exploration for Agentic Issue Localization

    Mohammad Nour Al Awad, Sergey Ivanov

    cs.SEcs.AIarXiv:2608.29675v12026
  50. A Survey of Learning-based Automated Program Repair

    Quanjun Zhang, Chunrong Fang, Yuxiang Ma +2

    cs.SEarXiv:2301.03270v32023
  51. Model Context Protocol Threat Modeling and Analyzing Vulnerabilities to Prompt Injection with Tool Poisoning

    Charoes Huang, Xin Huang, Ngoc Phu Tran +1

    cs.CRcs.SEarXiv:2603.22489v12026
  52. From Technical Debt to Cognitive and Intent Debt: Rethinking Software Health in the Age of AI

    Margaret-Anne Storey

    cs.SEarXiv:2603.22106v42026
  53. Debt Behind the AI Boom: A Large-Scale Empirical Study of AI-Generated Code in the Wild

    Yue Liu, Ratnadira Widyasari, Yanjie Zhao +3

    cs.SEarXiv:2603.28592v22026
  54. Automatically Discovering, Reporting and Reproducing Android Application Crashes

    Kevin Moran, Mario Linares-Vásquez, Carlos Bernal-Cárdenas +2

    cs.SEarXiv:1706.01130v12017
  55. Exploring Quantum Software Testing Across Research and Practice: Emerging Results from a Multivocal Literature Review

    Rodolfo Gil-Pereira, Ronnie de Souza Santos, Cleyton Magalhaes +1

    cs.SEarXiv:2609.00354v12026
  56. Clawdrain: Exploiting Tool-Calling Chains for Stealthy Token Exhaustion in OpenClaw Agents

    Ben Dong, Hui Feng, Qian Wang

    cs.CRcs.SEarXiv:2603.00902v12026
  57. Auditing Anonymous AI Models: A Four-Stage Protocol for Black-Box Identity Verification

    Yisen Xi

    cs.SEcs.AIcs.CRarXiv:2608.31142v12026
  58. EffiSkill: Agent Skill Based Automated Code Efficiency Optimization

    Zimu Wang, Yuling Shi, Mengfan Li +4

    cs.SEcs.CLarXiv:2603.27850v12026
  59. Evaluating Tiny Recursive Models Across Training for Code Generation

    Anjani Sirivella, Aanisha Newaz, Glaucia Melo

    cs.AIcs.LGcs.SEarXiv:2608.29376v12026
  60. SkillProbe: Security Auditing for Emerging Agent Skill Marketplaces via Multi-Agent Collaboration

    Zihan Guo, Zhiyu Chen, Xiaohang Nie +3

    cs.CRcs.SEarXiv:2603.21019v12026