Software Engineering

Papers filed under cs.SE on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

421 to 480 of 1,389

  1. NVIDIA FLARE: Federated Learning from Simulation to Real-World

    Holger R. Roth, Yan Cheng, Yuhong Wen +20

    cs.LGcs.AIcs.CVarXiv:2210.13291v32022
  2. Detecting DBMS Bugs by Constructing Equivalent Representations of Intermediate Query Results

    Xiaoxu Niu, Gong Chen, Jinfu Chen +1

    cs.DBcs.SEarXiv:2608.30385v12026
  3. Astra: A Multi-Agent System for GPU Kernel Performance Optimization

    Anjiang Wei, Tianran Sun, Yogesh Seenichamy +5

    cs.DCcs.AIcs.CLarXiv:2509.07506v22025
  4. Agent Behavioral Contracts: Formal Specification and Runtime Enforcement for Reliable Autonomous AI Agents

    Varun Pratap Bhardwaj

    cs.AIcs.MAcs.SEarXiv:2602.22302v12026
  5. Getting pwn'd by AI: Penetration Testing with Large Language Models

    Andreas Happe, Jürgen Cito

    cs.CLcs.AIcs.CRarXiv:2308.00121v32023
  6. SWE-Lego: Pushing the Limits of Supervised Fine-tuning for Software Issue Resolving

    Chaofan Tao, Jierun Chen, Yuxin Jiang +11

    cs.SEcs.CLarXiv:2601.01426v22026
  7. Pandemic Programming: How COVID-19 affects software developers and how their organizations can help

    Paul Ralph, Sebastian Baltes, Gianisa Adisaputri +14

    cs.SEcs.CYarXiv:2005.01127v32020
  8. Contrastive Code Representation Learning

    Paras Jain, Ajay Jain, Tianjun Zhang +3

    cs.LGcs.AIcs.PLarXiv:2007.04973v42020
  9. Natural Language to Code Translation with Execution

    Freda Shi, Daniel Fried, Marjan Ghazvininejad +2

    cs.CLcs.SEarXiv:2204.11454v22022
  10. Who Needs MLOps: What Data Scientists Seek to Accomplish and How Can MLOps Help?

    Sasu Mäkinen, Henrik Skogström, Eero Laaksonen +1

    cs.SEarXiv:2103.08942v12021
  11. Legacy System Modernization with Coding Agents: A Case Study

    Iago da Silva Rodrigues Alves, Cristiano Politowski, João Eduardo Montandon

    cs.SEarXiv:2608.28972v12026
  12. What Is a System? An Interaction-Based Account of Structure-Behavior Coalescence in General Systems Theory

    William S. Chao

    cs.SEeess.SYarXiv:2609.00043v12026
  13. FEA-Bench: A Benchmark for Evaluating Repository-Level Code Generation for Feature Implementation

    Wei Li, Xin Zhang, Zhongxin Guo +6

    cs.SEcs.CLarXiv:2503.06680v22025
  14. ProgramBench: Can Language Models Rebuild Programs From Scratch?

    John Yang, Kilian Lieret, Jeffrey Ma +9

    cs.SEcs.AIarXiv:2605.03546v12026
  15. In-IDE Code Generation from Natural Language: Promise and Challenges

    Frank F. Xu, Bogdan Vasilescu, Graham Neubig

    cs.SEarXiv:2101.11149v32021
  16. Hints Help But Do They Teach? Evaluating Skills Transfer in Code Generation

    Will Badr

    cs.SEcs.AIcs.CLarXiv:2609.01106v12026
  17. Towards Robust Agentic CUDA Kernel Benchmarking, Verification, and Optimization

    Robert Tjarko Lange, Qi Sun, Aaditya Prasad +3

    cs.SEcs.AIcs.LGarXiv:2509.14279v12025
  18. Seed-Coder: Let the Code Model Curate Data for Itself

    ByteDance Seed, Yuyu Zhang, Jing Su +24

    cs.CLcs.SEarXiv:2506.03524v22025
  19. A Critical Review of "Automatic Patch Generation Learned from Human-Written Patches": Essay on the Problem Statement and the Evaluation of Automatic Software Repair

    Martin Monperrus

    cs.SEarXiv:1408.2103v12014
  20. UML Class Diagram Evaluation and Repair Strategies based on LLMs

    Jie Liang, Peng Liang, Chong Wang

    cs.SEarXiv:2608.28800v12026
  21. LiveCodeBench Pro: How Do Olympiad Medalists Judge LLMs in Competitive Programming?

    Zihan Zheng, Zerui Cheng, Zeyu Shen +16

    cs.SEcs.AIcs.CLarXiv:2506.11928v12025
  22. AIOpsLab: A Holistic Framework to Evaluate AI Agents for Enabling Autonomous Clouds

    Yinfang Chen, Manish Shetty, Gagan Somashekar +6

    cs.AIcs.DCcs.MAarXiv:2501.06706v12025
  23. Evaluating a 4B open-weights local LLM for agentic DFT workflows: a literature reproducibility audit

    Shambhu Bhandari Sharma

    cond-mat.mtrl-scics.SEarXiv:2608.29665v12026
  24. PostTrainBench: Can LLM Agents Automate LLM Post-Training?

    Ben Rank, Hardik Bhatnagar, Ameya Prabhu +4

    cs.SEcs.AIcs.LGarXiv:2603.08640v22026
  25. Write, Execute, Assess: Program Synthesis with a REPL

    Kevin Ellis, Maxwell Nye, Yewen Pu +3

    cs.PLcs.AIcs.LGarXiv:1906.04604v12019
  26. Self-Edit: Fault-Aware Code Editor for Code Generation

    Kechi Zhang, Zhuo Li, Jia Li +2

    cs.SEcs.CLarXiv:2305.04087v52023
  27. SWE-Debate: Competitive Multi-Agent Debate for Software Issue Resolution

    Han Li, Yuling Shi, Shaoxin Lin +6

    cs.SEcs.CLcs.LGarXiv:2507.23348v12025
  28. Collaboration Challenges in Building ML-Enabled Systems: Communication, Documentation, Engineering, and Process

    Nadia Nahar, Shurui Zhou, Grace Lewis +1

    cs.SEcs.LGarXiv:2110.10234v42021
  29. Towards Agentic Cloud Engineering: Graph and Loop Engineering with a Zero-Trust Agent Harness

    Sagar Srinivas Sakhinana, Venkataramana Runkana

    cs.SEcs.AIcs.LGarXiv:2609.00050v12026
  30. A Survey of LLM-based Automated Program Repair: Taxonomies, Design Paradigms, and Applications

    Boyang Yang, Zijian Cai, Fengling Liu +5

    cs.SEarXiv:2506.23749v32025
  31. Assistance or Disruption? Exploring and Evaluating the Design and Trade-offs of Proactive AI Programming Support

    Kevin Pu, Daniel Lazaro, Ian Arawjo +4

    cs.HCcs.AIcs.SEarXiv:2502.18658v42025
  32. AgentLogs: A Dataset for Opening the Black Box of GitHub's Cloud Agent

    Jonan Richards, Kosei Horikawa, Youmei Fan +2

    cs.SEcs.AIarXiv:2608.29204v12026
  33. SWE-Exp: Experience-Driven Software Issue Resolution

    Silin Chen, Shaoxin Lin, Yuling Shi +8

    cs.SEcs.CLcs.LGarXiv:2507.23361v22025
  34. OpenCodeInstruct: A Large-scale Instruction Tuning Dataset for Code LLMs

    Wasi Uddin Ahmad, Aleksander Ficek, Mehrzad Samadi +4

    cs.SEcs.CLarXiv:2504.04030v22025
  35. On the Effectiveness of LLM-as-a-judge for Code Generation and Summarization

    Giuseppe Crupi, Rosalia Tufano, Alejandro Velasco +3

    cs.SEarXiv:2507.16587v12025
  36. VulnBot: Autonomous Penetration Testing for A Multi-Agent Collaborative Framework

    He Kong, Die Hu, Jingguo Ge +3

    cs.SEarXiv:2501.13411v12025
  37. OrcaLoca: An LLM Agent Framework for Software Issue Localization

    Zhongming Yu, Hejia Zhang, Yujie Zhao +4

    cs.SEcs.AIarXiv:2502.00350v22025
  38. SkillForge: Forging Domain-Specific, Self-Evolving Agent Skills in Cloud Technical Support

    Xingyan Liu, Xiyue Luo, Linyu Li +3

    cs.IRcs.AIcs.SEarXiv:2604.08618v22026
  39. Fine-Tuning Large Language Models to Classify Pull Request-Issue Alignments: Going Beyond Prompting

    Mustafa Yasir Altunhan, Hüseyin Özgür Kamalı, Eray Tüzün

    cs.SEarXiv:2609.01087v12026
  40. From Failed Trajectories to Reliable LLM Agents: Diagnosing and Repairing Harness Flaws

    Mengzhuo Chen, Junjie Wang, Zhe Liu +3

    cs.SEcs.MAarXiv:2606.06324v22026
  41. Agentic Much? Adoption of Coding Agents on GitHub

    Romain Robbes, Théo Matricon, Thomas Degueule +2

    cs.SEarXiv:2601.18341v22026
  42. Vibe Coding vs. Agentic Coding: Fundamentals and Practical Implications of Agentic AI

    Ranjan Sapkota, Konstantinos I. Roumeliotis, Manoj Karkee

    cs.SEcs.AIcs.CLarXiv:2505.19443v12025
  43. SAVIOR: Towards Bug-Driven Hybrid Testing

    Yaohui Chen, Peng Li, Jun Xu +5

    cs.SEarXiv:1906.07327v12019
  44. SWE-PolyBench: A multi-language benchmark for repository level evaluation of coding agents

    Muhammad Shihab Rashid, Christian Bock, Yuan Zhuang +10

    cs.SEarXiv:2504.08703v32025
  45. A Survey of AIOps in the Era of Large Language Models

    Lingzhe Zhang, Tong Jia, Mengxi Jia +7

    cs.SEcs.CLarXiv:2507.12472v12025
  46. Deep Learning for Source Code Modeling and Generation: Models, Applications and Challenges

    Triet H. M. Le, Hao Chen, M. Ali Babar

    cs.SEcs.AIcs.LGarXiv:2002.05442v12020
  47. Empirical Software Engineering in Practice: Insights from Google

    Roberto Verdecchia, Justus Bogner

    cs.SEarXiv:2609.00247v12026
  48. CUDA-Harness: Harnessing Agentic CUDA Kernel Generation and Optimization from Natural Language

    Qi Fan, An Zou, Yehan Ma

    cs.CLcs.AIcs.MAarXiv:2609.00058v12026
  49. sbom-unifier: Integration Framework for Heterogeneous SBOMs

    Yusuke Moriwaki, Tetsuya Kanda, Yuki Manabe +5

    cs.SEarXiv:2608.30708v12026
  50. Exploring the Potential of ChatGPT in Automated Code Refinement: An Empirical Study

    Qi Guo, Junming Cao, Xiaofei Xie +4

    cs.SEarXiv:2309.08221v12023
  51. How Do AI Agents Spend Your Money? Analyzing and Predicting Token Consumption in Agentic Coding Tasks

    Longju Bai, Zhemin Huang, Xingyao Wang +5

    cs.CLcs.AIcs.CYarXiv:2604.22750v22026
  52. Continuous Autonomous Refactoring: A Research Roadmap for AI-Driven Code Quality Maintenance

    Xin Sun, Daniel Ståhl, Kristian Sandahl +1

    cs.SEarXiv:2609.01236v12026
  53. TOGA: A Neural Method for Test Oracle Generation

    Elizabeth Dinella, Gabriel Ryan, Todd Mytkowicz +1

    cs.SEarXiv:2109.09262v22021
  54. Code-Aware Prompting: A study of Coverage Guided Test Generation in Regression Setting using LLM

    Gabriel Ryan, Siddhartha Jain, Mingyue Shang +4

    cs.SEcs.LGarXiv:2402.00097v22024
  55. MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers

    Chaithanya Bandi, Razvan-Gabriel Dumitru, Ben Hertzberg +20

    cs.SEcs.AIarXiv:2602.00933v32026
  56. BugsInPy: A Database of Existing Bugs in Python Programs to Enable Controlled Testing and Debugging Studies

    Ratnadira Widyasari, Sheng Qin Sim, Camellia Lok +13

    cs.SEarXiv:2401.15481v12024
  57. LeetCodeDataset: A Temporal Dataset for Robust Evaluation and Efficient Training of Code LLMs

    Yunhui Xia, Wei Shen, Yan Wang +5

    cs.LGcs.CLcs.SEarXiv:2504.14655v12025
  58. ACECODER: Acing Coder RL via Automated Test-Case Synthesis

    Huaye Zeng, Dongfu Jiang, Haozhe Wang +3

    cs.SEcs.AIcs.CLarXiv:2502.01718v42025
  59. A Tale of Two Cities: Software Developers Working from Home During the COVID-19 Pandemic

    Denae Ford, Margaret-Anne Storey, Thomas Zimmermann +6

    cs.SEcs.CYcs.HCarXiv:2008.11147v32020
  60. Montage: a grid portal and software toolkit for science-grade astronomical image mosaicking

    Joseph C. Jacob, Daniel S. Katz, G. Bruce Berriman +8

    astro-ph.IMcs.DCcs.SEarXiv:1005.4454v12010