Software Engineering

Papers filed under cs.SE on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

121 to 180 of 1,389

  1. Model-based Exploration of the Frontier of Behaviours for Deep Learning System Testing

    Vincenzo Riccio, Paolo Tonella

    cs.SEcs.AIcs.LGarXiv:2007.02787v12020
  2. What Makes Good In-context Demonstrations for Code Intelligence Tasks with LLMs?

    Shuzheng Gao, Xin-Cheng Wen, Cuiyun Gao +3

    cs.SEarXiv:2304.07575v22023
  3. Detecting Optimization Bugs in Database Engines via Non-Optimizing Reference Engine Construction

    Manuel Rigger, Zhendong Su

    cs.SEcs.DBarXiv:2007.08292v12020
  4. Cross-Architecture Model Diffing with Crosscoders: Unsupervised Discovery of Differences Between LLMs

    Thomas Jiralerspong, Trenton Bricken

    cs.AIcs.LGcs.SEarXiv:2602.11729v12026
  5. Self-Improving AI Coding Agents Through Accumulated Behavioral Rules: A Closed-Loop Framework

    Aditya Aggarwal, Nahid Farhady Ghalaty

    cs.SEcs.AIarXiv:2607.13091v12026
  6. Wicked Problem, Parsimonious Solution: Securing Electric Vehicle Charging Station Software

    Emma Sheppard, Zachary Wadhams, Dalton Arford +2

    cs.CRcs.SEarXiv:2609.10502v12026
  7. TrajMark: Ownership Attribution and Segment-Level Tamper Localization for Coding-Agent Trajectories

    Bokang Zeng, Zheng Gao, Xiaoyu Li +2

    cs.CRcs.SEarXiv:2609.10416v12026
  8. dexamine: A Python package for Uniswap event data on Ethereum

    Magnus Hansson

    q-fin.TRcs.SEarXiv:2609.10407v12026
  9. Towards Scalable and Cost-Efficient Vulnerability Detection: A Study on Automatic Query Generation

    Ivana Clairine Irsan, Ratnadira Widyasari, Huihui Huang +6

    cs.SEcs.CRarXiv:2609.10412v12026
  10. Formal Verification of Autonomous Vehicle Platooning

    Maryam Kamali, Louise A. Dennis, Owen McAree +2

    cs.AIcs.SEarXiv:1602.01718v12016
  11. Ensembling LLMs for AI-Augmented Cybersecurity Software Requirements Generation

    Santiago Perez-Acuna, Yod-Samuel Martín, Juan C. Yelmo

    cs.SEarXiv:2609.10316v12026
  12. Assuring the Machine Learning Lifecycle: Desiderata, Methods, and Challenges

    Rob Ashmore, Radu Calinescu, Colin Paterson

    cs.LGcs.SEstat.MLarXiv:1905.04223v12019
  13. GraphDroid: Asynchronous LLM-Based Mobile App GUI Testing via History-Aware Exploration and Hybrid Intent Fulfillment

    Xiaolei Li, Jialun Cao, Zhijian Hou +3

    cs.SEarXiv:2609.10031v12026
  14. Beyond Repository Boundaries: Cross-Repository Graph Retrieval for Code Generation

    Minh Le-Anh, Nam Le Hai, Quyen Tran +4

    cs.SEarXiv:2609.09987v12026
  15. Socio-technical and Ethical Dimensions of Architecture Practices in FLOSS

    Sven Thielen

    cs.SEarXiv:2609.09975v12026
  16. HLSFactory-Agent: Large-Scale Agentic HLS Dataset Construction from Academic and Open-Source Projects

    Kaushik Chandana, Jay Imperatori, Tanmay Shukla +3

    cs.ARcs.SEarXiv:2609.09519v12026
  17. XAgent: eXecution-guided Agentic AI for Effective Localization and Resolution of GitHub Issues

    Hieu Huynh, Patanamon Thongtanunam, Michael Fu +2

    cs.SEarXiv:2609.09769v12026
  18. Keep Evaluation Fair: Detecting Data Leakage in Code Generation Benchmarks via Membership Inference Attacks

    Dongdong Zhao, Jian Chen, Guancheng Lin +3

    cs.SEarXiv:2609.09865v12026
  19. How effective are traditional test criteria at detecting bugs in large language models generated code?

    Asma Hamidi, Michael Konstantinou, Renzo Degiovanni +1

    cs.SEarXiv:2609.09315v12026
  20. The Double Measurement Confound in Agent Benchmarks: De-Scaffolding, Ground-Truth Scoring, and Reliability Beyond the Mean

    Yonghong Zhang, Shadi Motaali, Vu Phong Dinh +6

    cs.SEcond-mat.mtrl-sciarXiv:2609.09218v12026
  21. Evaluating Enterprise Analytics Agents: An End-to-End, Trace-Backed Methodology

    Teja Venkat Kolli, Sang Su Lee, Xueying Yan +5

    cs.SEarXiv:2609.09182v12026
  22. OASIS: A Rubric-Based Multimodal Assessment Platform Using Large Language Models

    Ameer H. Shakur, Shinyoung Kang, David Hein +8

    cs.SEarXiv:2609.09180v12026
  23. SPT-Code: Sequence-to-Sequence Pre-Training for Learning Source Code Representations

    Changan Niu, Chuanyi Li, Vincent Ng +3

    cs.SEarXiv:2201.01549v42022
  24. Automatic Repair of Buggy If Conditions and Missing Preconditions with SMT

    Favio Demarco, Jifeng Xuan, Daniel Le Berre +1

    cs.SEarXiv:1404.3186v12014
  25. DeepTriage: Exploring the Effectiveness of Deep Learning for Bug Triaging

    Senthil Mani, Anush Sankaran, Rahul Aralikatte

    cs.SEcs.LGarXiv:1801.01275v12018
  26. Automatic Bug Triage using Semi-Supervised Text Classification

    Jifeng Xuan, He Jiang, Zhilei Ren +2

    cs.SEarXiv:1704.04769v12017
  27. Automatically Securing Permission-Based Software by Reducing the Attack Surface: An Application to Android

    Alexandre Bartel, Jacques Klein, Martin Monperrus +1

    cs.CRcs.SEarXiv:1206.5829v22012
  28. Applied Metamodelling: A Foundation for Language Driven Development (Third Edition)

    Tony Clark, Paul Sammut, James Willans

    cs.SEarXiv:1505.00149v12015
  29. The Vision of Software Clone Management: Past, Present, and Future

    Chanchal K. Roy, Minhaz F. Zibran, Rainer Koschke

    cs.SEarXiv:2005.01005v12020
  30. Red teaming ChatGPT via Jailbreaking: Bias, Robustness, Reliability and Toxicity

    Terry Yue Zhuo, Yujin Huang, Chunyang Chen +1

    cs.CLcs.SEarXiv:2301.12867v42023
  31. Is Stack Overflow Obsolete? An Empirical Study of the Characteristics of ChatGPT Answers to Stack Overflow Questions

    Samia Kabir, David N. Udo-Imeh, Bonan Kou +1

    cs.SEcs.AIarXiv:2308.02312v42023
  32. RACK: Automatic API Recommendation using Crowdsourced Knowledge

    Mohammad Masudur Rahman, Chanchal K. Roy, David Lo

    cs.SEarXiv:1807.02953v12018
  33. Closing the gap between software engineering education and industrial needs

    Vahid Garousi, Görkem Giray, Eray Tüzün +2

    cs.SEarXiv:1812.01954v12018
  34. Cost-Aware Post-Hoc Deferral Under Calibration and Shift: An Environmental AI Case Study

    Haoran Yu, Lifei Liu, Danping Zhang

    cs.SEcs.LGarXiv:2609.09235v12026
  35. Query Expansion Based on Crowd Knowledge for Code Search

    Liming Nie, He Jiang, Zhilei Ren +2

    cs.SEarXiv:1703.01443v12017
  36. Adoption and Effects of Software Engineering Best Practices in Machine Learning

    Alex Serban, Koen van der Blom, Holger Hoos +1

    cs.SEarXiv:2007.14130v22020
  37. An Artificial Intelligence Life Cycle: From Conception to Production

    Daswin De Silva, Damminda Alahakoon

    cs.SEarXiv:2108.13861v12021
  38. Vul-RAG: Enhancing LLM-based Vulnerability Detection via Knowledge-level RAG

    Xueying Du, Geng Zheng, Kaixin Wang +9

    cs.SEcs.AIarXiv:2406.11147v32024
  39. NatGen: Generative pre-training by "Naturalizing" source code

    Saikat Chakraborty, Toufique Ahmed, Yangruibo Ding +2

    cs.PLcs.AIcs.LGarXiv:2206.07585v22022
  40. Retrofitting Code Using LLMs to Support Exceptional Behavior

    Linghan Zhong, Jiyang Zhang, Jayanth Srinivasa +2

    cs.SEcs.CLarXiv:2609.10397v12026
  41. If It's Not Buggy, Don't Fix It: On the Dynamics of Iterative Bug-fixing with LLMs

    Xietao Wang-Lin, Anton Isopoussu, Louis Mahon

    cs.SEcs.CLarXiv:2609.10123v12026
  42. Software Startups -- A Research Agenda

    Michael Unterkalmsteiner, Pekka Abrahamsson, Xiaofeng Wang +24

    cs.SEarXiv:2308.12816v12023
  43. Socio-Technical Grounded Theory for Software Engineering

    Rashina Hoda

    cs.SEarXiv:2103.14235v32021
  44. Fairness Testing: A Comprehensive Survey and Analysis of Trends

    Zhenpeng Chen, Jie M. Zhang, Max Hort +2

    cs.SEarXiv:2207.10223v42022
  45. SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code Agents

    Niels Mündler, Mark Niklas Müller, Jingxuan He +1

    cs.SEcs.AIcs.LGarXiv:2406.12952v32024
  46. Dissection of a Bug Dataset: Anatomy of 395 Patches from Defects4J

    Victor Sobreira, Thomas Durieux, Fernanda Madeiral +2

    cs.SEarXiv:1801.06393v32018
  47. StateAFL: Greybox Fuzzing for Stateful Network Servers

    Roberto Natella

    cs.CRcs.OScs.SEarXiv:2110.06253v22021
  48. Bears: An Extensible Java Bug Benchmark for Automatic Program Repair Studies

    Fernanda Madeiral, Simon Urli, Marcelo Maia +1

    cs.SEarXiv:1901.06024v12019
  49. Design, Monitoring, and Testing of Microservices Systems: The Practitioners' Perspective

    Muhammad Waseem, Peng Liang, Mojtaba Shahin +2

    cs.SEarXiv:2108.03384v22021
  50. Experience Report: Deep Learning-based System Log Analysis for Anomaly Detection

    Zhuangbin Chen, Jinyang Liu, Wenwei Gu +2

    cs.SEcs.LGarXiv:2107.05908v22021
  51. A Manually-Curated Dataset of Fixes to Vulnerabilities of Open-Source Software

    Serena E. Ponta, Henrik Plate, Antonino Sabetta +2

    cs.SEcs.CRcs.LGarXiv:1902.02595v32019
  52. CoSQA: 20,000+ Web Queries for Code Search and Question Answering

    Junjie Huang, Duyu Tang, Linjun Shou +5

    cs.CLcs.SEarXiv:2105.13239v12021
  53. Natural Language Generation and Understanding of Big Code for AI-Assisted Programming: A Review

    Man Fai Wong, Shangxin Guo, Ching Nam Hang +2

    cs.SEcs.AIcs.CLarXiv:2307.02503v12023
  54. Runtime Verification for Business Processes Utilizing the Bitcoin Blockchain

    Christoph Prybila, Stefan Schulte, Christoph Hochreiner +1

    cs.SEcs.DCarXiv:1706.04404v22017
  55. A-JIT: Agentic Just-In-Time Software Construction

    Mark Marron, Earl T. Barr

    cs.SEcs.AIarXiv:2609.10248v12026
  56. Building Program Vector Representations for Deep Learning

    Lili Mou, Ge Li, Yuxuan Liu +4

    cs.SEcs.LGcs.NEarXiv:1409.3358v12014
  57. Context operations to architecture modelling output from large language models and evaluation criteria for their use in systems engineering design

    Vinicius Kaster Marini, Petter Krus

    eess.SYcs.AIcs.SEarXiv:2609.10132v12026
  58. An Integrated Semantic Web Service Discovery and Composition Framework

    Pablo Rodriguez-Mier, Carlos Pedrinaci, Manuel Lama +1

    cs.AIcs.SEarXiv:1502.02840v12015
  59. What are Weak Links in the npm Supply Chain?

    Nusrat Zahan, Thomas Zimmermann, Patrice Godefroid +3

    cs.CRcs.CYcs.SEarXiv:2112.10165v22021
  60. Introducing Consort: A Spec-First Agent Framework for Enforced, Test-Driven Development on Live Database Branches

    Kevin Hartman

    cs.SEcs.AIcs.DBarXiv:2609.09671v12026