Software Engineering
Papers filed under cs.SE on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
301 to 360 of 1,389
Intent Formalization: A Grand Challenge for Reliable Coding in the Age of AI Agents
Shuvendu K. Lahiri
cs.SEcs.AIcs.PLarXiv:2603.17150v12026RealSec-bench: A Benchmark for Evaluating Secure Code Generation in Real-World Repositories
Yanlin Wang, Ziyao Zhang, Chong Wang +5
cs.CRcs.SEarXiv:2601.22706v12026DafnyPro: LLM-Assisted Automated Verification for Dafny Programs
Debangshu Banerjee, Olivier Bouissou, Stefan Zetzsche
cs.SEarXiv:2601.05385v12026Programming by Chat: A Large-Scale Behavioral Analysis of 11,579 Real-World AI-Assisted IDE Sessions
Ningzhi Tang, Chaoran Chen, Zihan Fang +6
cs.SEcs.HCarXiv:2604.00436v22026PlotChain: Deterministic Checkpointed Evaluation of Multimodal LLMs on Engineering Plot Reading
Mayank Ravishankara
cs.AIcs.SEarXiv:2602.13232v120266-Layer Model for a Structured Description and Categorization of Urban Traffic and Environment
Maike Scholtes, Lukas Westhofen, Lara Ruth Turner +11
cs.OHcs.AIcs.SEarXiv:2012.06319v22020Edge Impulse: An MLOps Platform for Tiny Machine Learning
Shawn Hymel, Colby Banbury, Daniel Situnayake +13
cs.DCcs.LGcs.SEarXiv:2212.03332v32022Where Do AI Coding Agents Fail? An Empirical Study of Failed Agentic Pull Requests in GitHub
Ramtin Ehsani, Sakshi Pathak, Shriya Rawal +3
cs.SEcs.AIarXiv:2601.15195v12026Immersion in the GitHub Universe: Scaling Coding Agents to Mastery
Jiale Zhao, Guoxin Chen, Fanzhe Meng +11
cs.SEarXiv:2602.09892v42026How does information access affect LLM monitors' ability to detect sabotage?
Rauno Arike, Raja Mehta Moreno, Rohan Subramani +2
cs.AIcs.SEarXiv:2601.21112v22026What do we know about software development in startups?
Carmine Giardino, Michael Unterkalmsteiner, Nicolò Paternoster +2
cs.SEarXiv:2307.13707v12023The History Is the Detector: Executing CVE Patch History, End-to-End
Qiushi Wu, Kevin Eykholt, Youngja Park +4
cs.CRcs.AIcs.SEarXiv:2609.05335v12026Compressing Code Context for LLM-based Issue Resolution
Haoxiang Jia, Earl T. Barr, Sergey Mechtaev
cs.SEarXiv:2603.28119v12026Goedel-Code-Prover: Hierarchical Proof Search for Open State-of-the-Art Code Verification
Zenan Li, Ziran Yang, Deyuan He +7
cs.SEcs.AIarXiv:2603.19329v32026Evaluating and Improving Automated Repository-Level Rust Issue Resolution with LLM-based Agents
Jiahong Xiang, Wenxiao He, Xihua Wang +2
cs.SEarXiv:2602.22764v12026ARIA - An Agentic Framework for Autonomous Testing of Infotainment Systems
António Azevedo, Bruno Lima, João Pascoal Faria
cs.SEcs.AIarXiv:2609.04913v12026Better Understanding, Better Fixes? A Study of Hallucination in LLM-based Automated Program Repair
Xuemeng Cai, Jiakun Liu, Linhan Yang +2
cs.SEcs.AIarXiv:2609.04909v12026ContractSkill: Repairable Contract-Based Skills for Multimodal Web Agents
Zijian Lu, Yiping Zuo, Yupeng Nie +4
cs.SEcs.AIarXiv:2603.20340v32026Characterizing Faults in Agentic AI: A Taxonomy of Types, Symptoms, and Root Causes
Mehil B Shah, Mohammad Mehdi Morovati, Mohammad Masudur Rahman +1
cs.SEarXiv:2603.06847v22026Beyond Code Generation: Reliability, Verification, and Cost Economics in the Agentic Software Development Lifecycle
Happy Bhati
cs.SEcs.AIarXiv:2609.04681v12026In Line with Context: Repository-Level Code Generation via Context Inlining
Chao Hu, Wenhao Zeng, Yuling Shi +2
cs.SEcs.AIarXiv:2601.00376v32026Dynamic Adaptation of the LLM Context for Generating Routines with Coupled Semantics
Gnaneswar Villuri, Hashmath Shaik, Alex Doboli
cs.SEcs.AIarXiv:2609.04570v12026Building a research-software catalog with a coding agent: from hackathon prototype to public deployment
Kazuyoshi Yoshimi, Satoshi Terasaki, Gotai Yamada
cs.SEcs.AIcs.CYarXiv:2609.04711v12026Advancing Requirements Engineering through Generative AI: Assessing the Role of LLMs
Chetan Arora, John Grundy, Mohamed Abdelrazek
cs.SEarXiv:2310.13976v22023Software Development in Startup Companies: The Greenfield Startup Model
Carmine Giardino, Nicolò Paternoster, Michael Unterkalmsteiner +2
cs.SEarXiv:2308.09438v12023Trace2Tower: Transition-Aware EigenTrace Induction of Multi-Level Skills for LLM Agents
Jiazheng Sun, Boyu Yang, Binhao Yuan +2
cs.AIcs.SEarXiv:2609.05261v12026Cloud-OpsBench: A Reproducible Benchmark for Agentic Root Cause Analysis in Cloud Systems
Yilun Wang, Guangba Yu, Haiyu Huang +4
cs.SEarXiv:2603.00468v22026SWE-ABS: Adversarial Benchmark Strengthening Exposes Inflated Success Rates on Test-based Benchmark
Boxi Yu, Yang Cao, Yuzhong Zhang +9
cs.SEarXiv:2603.00520v12026Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents?
Thibaud Gloaguen, Niels Mündler, Mark Müller +2
cs.SEcs.AIarXiv:2602.11988v22026Library Drift: Diagnosing and Fixing a Silent Failure Mode in Self-Evolving LLM Skill Libraries
Xing Zhang, Yanwei Cui, Guanghui Wang +4
cs.AIcs.CLcs.SEarXiv:2605.19576v32026Is GitHub's Copilot as Bad as Humans at Introducing Vulnerabilities in Code?
Owura Asare, Meiyappan Nagappan, N. Asokan
cs.SEcs.CRarXiv:2204.04741v52022Recommender System for Online Dating Service
Lukas Brozovsky, Vaclav Petricek
cs.IRcs.SEarXiv:cs/0703042v12007AI IDEs or Autonomous Agents? Measuring the Impact of Coding Agents on Software Development
Shyam Agarwal, Hao He, Bogdan Vasilescu
cs.SEarXiv:2601.13597v22026Deterministic LLM Inference Across GPU Kernels: Power-of-Two INT8 Quantization Scales and the Limits of Tolerance-Based Conformance
Teng-Ruei Chen
cs.LGcs.SEarXiv:2609.00363v12026AgentWorm: Self-Propagating Attacks Across LLM Agent Ecosystems
Yihao Zhang, Zeming Wei, Xiaokun Luan +7
cs.CRcs.AIcs.LGarXiv:2603.15727v32026Large Language Model for Vulnerability Detection: Emerging Results and Future Directions
Xin Zhou, Ting Zhang, David Lo
cs.SEarXiv:2401.15468v12024The Web-CLI: Verifiable Privacy for Tools, Models, and Inference Engines in the Browser
Tejaswi Gowda
cs.HCcs.CRcs.SEarXiv:2608.28950v22026SWE-Refactor: A Repository-Level Benchmark for Real-World LLM-Based Code Refactoring
Yisen Xu, Jinqiu Yang, Tse-Hsun +1
cs.SEarXiv:2602.03712v12026The reach of a verification tool decides its value: A controlled study of verification surface, artifact quality, and cost in AI coding agents
Achint Mehta
cs.SEcs.AIarXiv:2608.28795v12026Spec-Driven Development:From Code to Contract in the Age of AI Coding Assistants
Deepak Babu Piskala
cs.SEcs.AIarXiv:2602.00180v12026Wink: Recovering from Misbehaviors in Coding Agents
Rahul Nanda, Chandra Maddila, Smriti Jha +3
cs.SEcs.AIcs.HCarXiv:2602.17037v22026Large Language Models for Code Generation: A Comprehensive Survey of Challenges, Techniques, Evaluation, and Applications
Nam Huynh, Beiyu Lin
cs.SEcs.LGarXiv:2503.01245v22025From Flat Logs to Causal Graphs: Hierarchical Failure Attribution for LLM-based Multi-Agent Systems
Yawen Wang, Wenjie Wu, Junjie Wang +1
cs.AIcs.SEarXiv:2602.23701v12026Promises, Perils, and (Timely) Heuristics for Mining Coding Agent Activity
Romain Robbes, Théo Matricon, Thomas Degueule +2
cs.SEarXiv:2601.18345v12026AICD Bench: A Challenging Benchmark for AI-Generated Code Detection
Daniil Orel, Dilshod Azizov, Indraneil Paul +3
cs.LGcs.SEarXiv:2602.02079v12026Formal Scenario-Based Testing of Autonomous Vehicles: From Simulation to the Real World
Daniel J. Fremont, Edward Kim, Yash Vardhan Pant +7
eess.SYcs.LGcs.LOarXiv:2003.07739v22020BenchGuard: Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks
Xinming Tu, Tianze Wang, Yingzhou +4
cs.CLcs.AIcs.SEarXiv:2604.24955v12026Building the Truman Show: A TrustZone-Based Framework for Lightweight Out-of-band Kernel Security Monitoring
Zhenling Duan, Pan Dong, Renshuang Jiang +2
cs.CRcs.SEarXiv:2608.29758v12026Cost-Effective Repository Exploration for Agentic Issue Localization
Mohammad Nour Al Awad, Sergey Ivanov
cs.SEcs.AIarXiv:2608.29675v12026A Survey of Learning-based Automated Program Repair
Quanjun Zhang, Chunrong Fang, Yuxiang Ma +2
cs.SEarXiv:2301.03270v32023Model Context Protocol Threat Modeling and Analyzing Vulnerabilities to Prompt Injection with Tool Poisoning
Charoes Huang, Xin Huang, Ngoc Phu Tran +1
cs.CRcs.SEarXiv:2603.22489v12026From Technical Debt to Cognitive and Intent Debt: Rethinking Software Health in the Age of AI
Margaret-Anne Storey
cs.SEarXiv:2603.22106v42026Debt Behind the AI Boom: A Large-Scale Empirical Study of AI-Generated Code in the Wild
Yue Liu, Ratnadira Widyasari, Yanjie Zhao +3
cs.SEarXiv:2603.28592v22026Automatically Discovering, Reporting and Reproducing Android Application Crashes
Kevin Moran, Mario Linares-Vásquez, Carlos Bernal-Cárdenas +2
cs.SEarXiv:1706.01130v12017Exploring Quantum Software Testing Across Research and Practice: Emerging Results from a Multivocal Literature Review
Rodolfo Gil-Pereira, Ronnie de Souza Santos, Cleyton Magalhaes +1
cs.SEarXiv:2609.00354v12026Clawdrain: Exploiting Tool-Calling Chains for Stealthy Token Exhaustion in OpenClaw Agents
Ben Dong, Hui Feng, Qian Wang
cs.CRcs.SEarXiv:2603.00902v12026Auditing Anonymous AI Models: A Four-Stage Protocol for Black-Box Identity Verification
Yisen Xi
cs.SEcs.AIcs.CRarXiv:2608.31142v12026EffiSkill: Agent Skill Based Automated Code Efficiency Optimization
Zimu Wang, Yuling Shi, Mengfan Li +4
cs.SEcs.CLarXiv:2603.27850v12026Evaluating Tiny Recursive Models Across Training for Code Generation
Anjani Sirivella, Aanisha Newaz, Glaucia Melo
cs.AIcs.LGcs.SEarXiv:2608.29376v12026SkillProbe: Security Auditing for Emerging Agent Skill Marketplaces via Multi-Agent Collaboration
Zihan Guo, Zhiyu Chen, Xiaohang Nie +3
cs.CRcs.SEarXiv:2603.21019v12026