Software Engineering

Papers filed under cs.SE on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

1 to 60 of 1,389

  1. Owl Eyes: Spotting UI Display Issues via Visual Understanding

    Zhe Liu, Chunyang Chen, Junjie Wang +3

    cs.SEarXiv:2009.01417v22020
  2. Learning Performance-Improving Code Edits

    Alexander Shypula, Aman Madaan, Yimeng Zeng +7

    cs.SEcs.AIcs.LGarXiv:2302.07867v52023
  3. Verus-SpecGym: An Agentic Environment for Evaluating Specification Autoformalization

    Anmol Agarwal, Natalie Neamtu, Pranjal Aggarwal +6

    cs.SEcs.AIcs.CLarXiv:2605.26457v12026
  4. TBar: Revisiting Template-based Automated Program Repair

    Kui Liu, Anil Koyuncu, Dongsun Kim +1

    cs.SEarXiv:1903.08409v22019
  5. VLAA-GUI: Knowing When to Stop, Recover, and Search, A Modular Framework for GUI Automation

    Qijun Han, Haoqin Tu, Zijun Wang +11

    cs.CLcs.AIcs.SEarXiv:2604.21375v22026
  6. Scaling Test-Time Compute for Agentic Coding

    Joongwon Kim, Wannan Yang, Kelvin Niu +13

    cs.SEcs.AIcs.CLarXiv:2604.16529v12026
  7. Log-based Anomaly Detection Without Log Parsing

    Van-Hoang Le, Hongyu Zhang

    cs.SEcs.AIarXiv:2108.01955v32021
  8. Developer Attitudes and Practices Towards Optimizing Software Energy Consumption

    Max Weber, Alina Mailach, Florian Sattler +2

    cs.SEarXiv:2608.30527v12026
  9. FuzzingBrain-Bench V1: Evaluating Open-Ended Bug Discovery by LLMs

    Ze Sheng, Aleksandar Kezic, Zhicheng Chen +1

    cs.AIcs.CRcs.LGarXiv:2608.25158v12026
  10. Secret MCP: Evidence-Bounded and Context-Isolated Design Specification Generation from Web Screenshots

    Yeongjin Jo

    cs.SEarXiv:2608.24944v12026
  11. What is Wrong with Topic Modeling? (and How to Fix it Using Search-based Software Engineering)

    Amritanshu Agrawal, Wei Fu, Tim Menzies

    cs.SEcs.AIcs.CLarXiv:1608.08176v42016
  12. MemGovern: Enhancing Code Agents through Learning from Governed Human Experiences

    Qihao Wang, Ziming Cheng, Shuo Zhang +12

    cs.SEcs.AIarXiv:2601.06789v22026
  13. Repo0: Design-Driven Zero-to-All Code Generation

    Silin Chen, Haoyi Teng, Xiaodong Gu +5

    cs.SEcs.AIarXiv:2608.19854v12026
  14. A Comparative Study of Programming Languages in Rosetta Code

    Sebastian Nanz, Carlo A. Furia

    cs.SEarXiv:1409.0252v42014
  15. Flama: a Python framework for development and deployment of production-ready APIs, machine learning, and LLM services

    José A. Perdiguero López, Miguel A. Durán-Olivencia

    cs.SEcs.AIcs.LGarXiv:2608.18733v12026
  16. GraphCodeBERT: Pre-training Code Representations with Data Flow

    Daya Guo, Shuo Ren, Shuai Lu +15

    cs.SEcs.CLarXiv:2009.08366v42020
  17. The Quality of Claude AI-authored Python Tests Is Not Weaker Than Human-authored Tests

    Douglas J. Leith

    cs.SEcs.AIarXiv:2608.15188v12026
  18. SWE-Pruner Pro: The Coder LLM Already Knows What to Prune

    Yuhang Wang, Yuling Shi, Shaoqiu Zhang +6

    cs.CLcs.SEarXiv:2607.18213v12026
  19. SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering

    John Yang, Carlos E. Jimenez, Alexander Wettig +4

    cs.SEcs.AIcs.CLarXiv:2405.15793v32024
  20. SyGuS-Comp 2016: Results and Analysis

    Rajeev Alur, Dana Fisman, Rishabh Singh +1

    cs.SEcs.LGcs.LOarXiv:1611.07627v12016
  21. Proof-or-Stop: Don't Trust the Agent, Trust the Evidence -- Loop Engineering for Verifiable Evidence-Gated Lifecycle Control

    Jek Huang, Jeffery Hsia, Jiayi Sun +3

    cs.AIcs.SEarXiv:2607.14890v12026
  22. Don't Complete It! Preventing Unhelpful Code Completion for Productive and Sustainable Neural Code Completion Systems

    Zhensu Sun, Xiaoning Du, Fu Song +4

    cs.SEcs.AIarXiv:2209.05948v32022
  23. When to Show a Suggestion? Integrating Human Feedback in AI-Assisted Programming

    Hussein Mozannar, Gagan Bansal, Adam Fourney +1

    cs.HCcs.LGcs.SEarXiv:2306.04930v32023
    Summaries:한국어
  24. Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

    Evan Hubinger, Carson Denison, Jesse Mu +36

    cs.CRcs.AIcs.CLarXiv:2401.05566v32024
  25. An Analysis of the Automatic Bug Fixing Performance of ChatGPT

    Dominik Sobania, Martin Briesch, Carol Hanna +1

    cs.SEarXiv:2301.08653v12023
  26. Minimal Prompt Perturbations Lead to Code Vulnerabilities: Prompt Fragility and Hidden-State Signals in Coding LLMs

    Alexander Sternfeld, Andrei Kucharavy, Ljiljana Dolamic

    cs.CRcs.CLcs.SEarXiv:2605.29737v12026
  27. A Study on the Impact of Natural Language Differences in Prompts on Automatic Code Generation Using LLMs

    Haruka Tokumasu, Masanari Kondo, Alexander Serebrenik +5

    cs.SEarXiv:2609.18311v12026
  28. MVD: Memory-Related Vulnerability Detection Based on Flow-Sensitive Graph Neural Networks

    Sicong Cao, Xiaobing Sun, Lili Bo +3

    cs.CRcs.SEarXiv:2203.02660v12022
  29. Search-Based Metamorphic Testing of Vision-Language Models in Autonomous Underwater Robotic Software

    Muhammad Yousaf, Aitor Arrieta, Shaukat Ali +2

    cs.SEarXiv:2609.17007v12026
  30. A Case-Bundle Operating Model for Coding Agents in OpenFOAM-Based CFD

    Ke Xiao, Han Li, Teng Zhang +3

    cs.SEcs.CEcs.DCarXiv:2609.11941v12026
  31. Models as Governed Interfaces for AI-Native MBSE: Read-Side Adequacy and Write-Side Admissibility

    Jason Gower, Michael J. de C. Henshaw, Siyuan Ji

    cs.SEcs.AIarXiv:2609.16252v12026
  32. An Exploratory Study of Dependabot Cooldown Adoption in Open-Source GitHub Projects

    Hidetake Tanaka, Rikuto Tsuchida, Kazumasa Shimari +2

    cs.SEarXiv:2609.16605v12026
  33. Intelligent Semantic Matching (ISM) for Video Tutorial Search using Transformer Models

    Ahmad J. Tayeb, Sonia Haiduc

    cs.SEarXiv:2609.12921v12026
  34. AI Policies: Help or Hindrance? A Software Developer's Perspective

    Samuel Ferino, Rashina Hoda, John Grundy +2

    cs.SEcs.AIarXiv:2609.16496v12026
  35. Towards an Asset Administration Shell Maturity Model

    Carsten Ellwein, David Dietrich, Rozana Cvitkovic +1

    cs.SEeess.SYarXiv:2609.17084v12026
  36. MaRDMO: FAIR Documentation of In-Silico Research

    Marco Reidelbach, Marcus Weber

    cs.DLcs.SEarXiv:2609.11931v12026
  37. Natural-Language to SysMLv2 Translation via Conformance-Driven Iterative Refinement

    Chance LaVoie, Eladio Andujar Lugo, Taylan G. Topcu +1

    cs.SEcs.AIarXiv:2607.14162v12026
  38. FrontierCS: Evolving Challenges for Evolving Intelligence

    Qiuyang Mang, Wenhao Chai, Zhifei Li +48

    cs.LGcs.SEarXiv:2512.15699v12025
  39. Type-IV Code Clone Detection via Layer-Wise Non-Contrastive Representation Learning

    Luciano Marchezan, Kevin Delcourt, Eugene Syriani +1

    cs.SEcs.LGarXiv:2609.17338v12026
  40. Guidelines for the Search Strategy to Update Systematic Literature Reviews in Software Engineering

    Claes Wohlin, Emilia Mendes, Katia Romero Felizardo +1

    cs.SEarXiv:2006.05542v12020
  41. SkillMOO: Multi-Objective Optimization of Agent Skills for Software Engineering

    Jingzhi Gong, Ruizhen Gu, Zhiwei Fei +7

    cs.SEcs.AIarXiv:2604.09297v32026
  42. How Coding Agents Fail Their Users: A Large-Scale Analysis of Developer-Agent Misalignment in 20,574 Real-World Sessions

    Ningzhi Tang, Chaoran Chen, Gelei Xu +5

    cs.SEcs.AIcs.HCarXiv:2605.29442v22026
  43. Agentic Memory Enhanced Recursive Reasoning for Root Cause Localization in Microservices

    Lingzhe Zhang, Tong Jia, Yunpeng Zhai +5

    cs.SEcs.AIarXiv:2601.02732v12026
  44. FixMiner: Mining Relevant Fix Patterns for Automated Program Repair

    Anil Koyuncu, Kui Liu, Tegawendé F. Bissyandé +4

    cs.SEarXiv:1810.01791v22018
  45. Sorting and Transforming Program Repair Ingredients via Deep Learning Code Similarities

    Martin White, Michele Tufano, Matias Martinez +2

    cs.SEarXiv:1707.04742v22017
  46. Continuous Integration, Delivery and Deployment: A Systematic Review on Approaches, Tools, Challenges and Practices

    Mojtaba Shahin, Muhammad Ali Babar, Liming Zhu

    cs.SEarXiv:1703.07019v12017
  47. A Survey of Smart Contract Formal Specification and Verification

    Palina Tolmach, Yi Li, Shang-Wei Lin +2

    cs.SEarXiv:2008.02712v32020
  48. Learn&Fuzz: Machine Learning for Input Fuzzing

    Patrice Godefroid, Hila Peleg, Rishabh Singh

    cs.AIcs.CRcs.LGarXiv:1701.07232v12017
  49. The BrowserGym Ecosystem for Web Agent Research

    Thibault Le Sellier De Chezelles, Maxime Gasse, Alexandre Drouin +17

    cs.LGcs.AIcs.SEarXiv:2412.05467v42024
  50. Lie to Me: Finding Bugs in ZK DSL Toolchains with Adversarial Witness Injection

    Sebastian Watzinger, Christoph Hochrainer, Valentin Wüstholz +1

    cs.CRcs.SEarXiv:2608.30648v12026
  51. ChatDev: Communicative Agents for Software Development

    Chen Qian, Wei Liu, Hongzhang Liu +11

    cs.SEcs.CLcs.MAarXiv:2307.07924v52023
  52. FoldKit: A Python library for efficient storage and retrieval of co-folding predictions

    Jonathan A. Levine, Melissa Pathil, Samuel Nitz +2

    q-bio.QMcs.SEarXiv:2608.28788v12026
  53. Operationalizing Regulations into Code: A Model to Enhance Governance and Compliance in LLM Selection for Software Engineering

    Jonysberg Quintino, Hermano Moura, Filipe Calegário

    cs.SEarXiv:2608.27703v12026
  54. MemToC: Benchmarking Memory-Tool Conflict Resolution in Large Language Models

    Arseniy Varlamov, Rishat Zinnatullin, Elisei Rykov +2

    cs.CLcs.AIcs.MAarXiv:2608.26295v12026
  55. From Natural Language Requirements to Graphical User Interfaces: Automated Prototyping and Verification with Pretrained Language Models

    Kristian Kolthoff

    cs.SEarXiv:2608.24749v12026
  56. On the Use of Deep Learning in Software Defect Prediction

    Görkem Giray, Kwabena Ebo Bennin, Ömer Köksal +2

    cs.SEarXiv:2210.02236v12022
  57. Causal Inference-Based Root Cause Analysis for Online Service Systems with Intervention Recognition

    Mingjie Li, Zeyan Li, Kanglin Yin +4

    cs.SEcs.AIarXiv:2206.05871v12022
  58. Getafix: Learning to Fix Bugs Automatically

    Johannes Bader, Andrew Scott, Michael Pradel +1

    cs.SEarXiv:1902.06111v52019
  59. Leveraging Automated Unit Tests for Unsupervised Code Translation

    Baptiste Roziere, Jie M. Zhang, Francois Charton +3

    cs.SEcs.CLcs.LGarXiv:2110.06773v22021
  60. AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities

    Kai Chen, Zichen Ding, Jiaye Ge +20

    cs.AIcs.SEarXiv:2607.13705v32026