Software Engineering
Papers filed under cs.SE on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
421 to 480 of 1,389
NVIDIA FLARE: Federated Learning from Simulation to Real-World
Holger R. Roth, Yan Cheng, Yuhong Wen +20
cs.LGcs.AIcs.CVarXiv:2210.13291v32022Detecting DBMS Bugs by Constructing Equivalent Representations of Intermediate Query Results
Xiaoxu Niu, Gong Chen, Jinfu Chen +1
cs.DBcs.SEarXiv:2608.30385v12026Astra: A Multi-Agent System for GPU Kernel Performance Optimization
Anjiang Wei, Tianran Sun, Yogesh Seenichamy +5
cs.DCcs.AIcs.CLarXiv:2509.07506v22025Agent Behavioral Contracts: Formal Specification and Runtime Enforcement for Reliable Autonomous AI Agents
Varun Pratap Bhardwaj
cs.AIcs.MAcs.SEarXiv:2602.22302v12026Getting pwn'd by AI: Penetration Testing with Large Language Models
Andreas Happe, Jürgen Cito
cs.CLcs.AIcs.CRarXiv:2308.00121v32023SWE-Lego: Pushing the Limits of Supervised Fine-tuning for Software Issue Resolving
Chaofan Tao, Jierun Chen, Yuxin Jiang +11
cs.SEcs.CLarXiv:2601.01426v22026Pandemic Programming: How COVID-19 affects software developers and how their organizations can help
Paul Ralph, Sebastian Baltes, Gianisa Adisaputri +14
cs.SEcs.CYarXiv:2005.01127v32020Contrastive Code Representation Learning
Paras Jain, Ajay Jain, Tianjun Zhang +3
cs.LGcs.AIcs.PLarXiv:2007.04973v42020Natural Language to Code Translation with Execution
Freda Shi, Daniel Fried, Marjan Ghazvininejad +2
cs.CLcs.SEarXiv:2204.11454v22022Who Needs MLOps: What Data Scientists Seek to Accomplish and How Can MLOps Help?
Sasu Mäkinen, Henrik Skogström, Eero Laaksonen +1
cs.SEarXiv:2103.08942v12021Legacy System Modernization with Coding Agents: A Case Study
Iago da Silva Rodrigues Alves, Cristiano Politowski, João Eduardo Montandon
cs.SEarXiv:2608.28972v12026What Is a System? An Interaction-Based Account of Structure-Behavior Coalescence in General Systems Theory
William S. Chao
cs.SEeess.SYarXiv:2609.00043v12026FEA-Bench: A Benchmark for Evaluating Repository-Level Code Generation for Feature Implementation
Wei Li, Xin Zhang, Zhongxin Guo +6
cs.SEcs.CLarXiv:2503.06680v22025ProgramBench: Can Language Models Rebuild Programs From Scratch?
John Yang, Kilian Lieret, Jeffrey Ma +9
cs.SEcs.AIarXiv:2605.03546v12026In-IDE Code Generation from Natural Language: Promise and Challenges
Frank F. Xu, Bogdan Vasilescu, Graham Neubig
cs.SEarXiv:2101.11149v32021Hints Help But Do They Teach? Evaluating Skills Transfer in Code Generation
Will Badr
cs.SEcs.AIcs.CLarXiv:2609.01106v12026Towards Robust Agentic CUDA Kernel Benchmarking, Verification, and Optimization
Robert Tjarko Lange, Qi Sun, Aaditya Prasad +3
cs.SEcs.AIcs.LGarXiv:2509.14279v12025Seed-Coder: Let the Code Model Curate Data for Itself
ByteDance Seed, Yuyu Zhang, Jing Su +24
cs.CLcs.SEarXiv:2506.03524v22025A Critical Review of "Automatic Patch Generation Learned from Human-Written Patches": Essay on the Problem Statement and the Evaluation of Automatic Software Repair
Martin Monperrus
cs.SEarXiv:1408.2103v12014UML Class Diagram Evaluation and Repair Strategies based on LLMs
Jie Liang, Peng Liang, Chong Wang
cs.SEarXiv:2608.28800v12026LiveCodeBench Pro: How Do Olympiad Medalists Judge LLMs in Competitive Programming?
Zihan Zheng, Zerui Cheng, Zeyu Shen +16
cs.SEcs.AIcs.CLarXiv:2506.11928v12025AIOpsLab: A Holistic Framework to Evaluate AI Agents for Enabling Autonomous Clouds
Yinfang Chen, Manish Shetty, Gagan Somashekar +6
cs.AIcs.DCcs.MAarXiv:2501.06706v12025Evaluating a 4B open-weights local LLM for agentic DFT workflows: a literature reproducibility audit
Shambhu Bhandari Sharma
cond-mat.mtrl-scics.SEarXiv:2608.29665v12026PostTrainBench: Can LLM Agents Automate LLM Post-Training?
Ben Rank, Hardik Bhatnagar, Ameya Prabhu +4
cs.SEcs.AIcs.LGarXiv:2603.08640v22026Write, Execute, Assess: Program Synthesis with a REPL
Kevin Ellis, Maxwell Nye, Yewen Pu +3
cs.PLcs.AIcs.LGarXiv:1906.04604v12019Self-Edit: Fault-Aware Code Editor for Code Generation
Kechi Zhang, Zhuo Li, Jia Li +2
cs.SEcs.CLarXiv:2305.04087v52023SWE-Debate: Competitive Multi-Agent Debate for Software Issue Resolution
Han Li, Yuling Shi, Shaoxin Lin +6
cs.SEcs.CLcs.LGarXiv:2507.23348v12025Collaboration Challenges in Building ML-Enabled Systems: Communication, Documentation, Engineering, and Process
Nadia Nahar, Shurui Zhou, Grace Lewis +1
cs.SEcs.LGarXiv:2110.10234v42021Towards Agentic Cloud Engineering: Graph and Loop Engineering with a Zero-Trust Agent Harness
Sagar Srinivas Sakhinana, Venkataramana Runkana
cs.SEcs.AIcs.LGarXiv:2609.00050v12026A Survey of LLM-based Automated Program Repair: Taxonomies, Design Paradigms, and Applications
Boyang Yang, Zijian Cai, Fengling Liu +5
cs.SEarXiv:2506.23749v32025Assistance or Disruption? Exploring and Evaluating the Design and Trade-offs of Proactive AI Programming Support
Kevin Pu, Daniel Lazaro, Ian Arawjo +4
cs.HCcs.AIcs.SEarXiv:2502.18658v42025AgentLogs: A Dataset for Opening the Black Box of GitHub's Cloud Agent
Jonan Richards, Kosei Horikawa, Youmei Fan +2
cs.SEcs.AIarXiv:2608.29204v12026SWE-Exp: Experience-Driven Software Issue Resolution
Silin Chen, Shaoxin Lin, Yuling Shi +8
cs.SEcs.CLcs.LGarXiv:2507.23361v22025OpenCodeInstruct: A Large-scale Instruction Tuning Dataset for Code LLMs
Wasi Uddin Ahmad, Aleksander Ficek, Mehrzad Samadi +4
cs.SEcs.CLarXiv:2504.04030v22025On the Effectiveness of LLM-as-a-judge for Code Generation and Summarization
Giuseppe Crupi, Rosalia Tufano, Alejandro Velasco +3
cs.SEarXiv:2507.16587v12025VulnBot: Autonomous Penetration Testing for A Multi-Agent Collaborative Framework
He Kong, Die Hu, Jingguo Ge +3
cs.SEarXiv:2501.13411v12025OrcaLoca: An LLM Agent Framework for Software Issue Localization
Zhongming Yu, Hejia Zhang, Yujie Zhao +4
cs.SEcs.AIarXiv:2502.00350v22025SkillForge: Forging Domain-Specific, Self-Evolving Agent Skills in Cloud Technical Support
Xingyan Liu, Xiyue Luo, Linyu Li +3
cs.IRcs.AIcs.SEarXiv:2604.08618v22026Fine-Tuning Large Language Models to Classify Pull Request-Issue Alignments: Going Beyond Prompting
Mustafa Yasir Altunhan, Hüseyin Özgür Kamalı, Eray Tüzün
cs.SEarXiv:2609.01087v12026From Failed Trajectories to Reliable LLM Agents: Diagnosing and Repairing Harness Flaws
Mengzhuo Chen, Junjie Wang, Zhe Liu +3
cs.SEcs.MAarXiv:2606.06324v22026Agentic Much? Adoption of Coding Agents on GitHub
Romain Robbes, Théo Matricon, Thomas Degueule +2
cs.SEarXiv:2601.18341v22026Vibe Coding vs. Agentic Coding: Fundamentals and Practical Implications of Agentic AI
Ranjan Sapkota, Konstantinos I. Roumeliotis, Manoj Karkee
cs.SEcs.AIcs.CLarXiv:2505.19443v12025SAVIOR: Towards Bug-Driven Hybrid Testing
Yaohui Chen, Peng Li, Jun Xu +5
cs.SEarXiv:1906.07327v12019SWE-PolyBench: A multi-language benchmark for repository level evaluation of coding agents
Muhammad Shihab Rashid, Christian Bock, Yuan Zhuang +10
cs.SEarXiv:2504.08703v32025A Survey of AIOps in the Era of Large Language Models
Lingzhe Zhang, Tong Jia, Mengxi Jia +7
cs.SEcs.CLarXiv:2507.12472v12025Deep Learning for Source Code Modeling and Generation: Models, Applications and Challenges
Triet H. M. Le, Hao Chen, M. Ali Babar
cs.SEcs.AIcs.LGarXiv:2002.05442v12020Empirical Software Engineering in Practice: Insights from Google
Roberto Verdecchia, Justus Bogner
cs.SEarXiv:2609.00247v12026CUDA-Harness: Harnessing Agentic CUDA Kernel Generation and Optimization from Natural Language
Qi Fan, An Zou, Yehan Ma
cs.CLcs.AIcs.MAarXiv:2609.00058v12026sbom-unifier: Integration Framework for Heterogeneous SBOMs
Yusuke Moriwaki, Tetsuya Kanda, Yuki Manabe +5
cs.SEarXiv:2608.30708v12026Exploring the Potential of ChatGPT in Automated Code Refinement: An Empirical Study
Qi Guo, Junming Cao, Xiaofei Xie +4
cs.SEarXiv:2309.08221v12023How Do AI Agents Spend Your Money? Analyzing and Predicting Token Consumption in Agentic Coding Tasks
Longju Bai, Zhemin Huang, Xingyao Wang +5
cs.CLcs.AIcs.CYarXiv:2604.22750v22026Continuous Autonomous Refactoring: A Research Roadmap for AI-Driven Code Quality Maintenance
Xin Sun, Daniel Ståhl, Kristian Sandahl +1
cs.SEarXiv:2609.01236v12026TOGA: A Neural Method for Test Oracle Generation
Elizabeth Dinella, Gabriel Ryan, Todd Mytkowicz +1
cs.SEarXiv:2109.09262v22021Code-Aware Prompting: A study of Coverage Guided Test Generation in Regression Setting using LLM
Gabriel Ryan, Siddhartha Jain, Mingyue Shang +4
cs.SEcs.LGarXiv:2402.00097v22024MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers
Chaithanya Bandi, Razvan-Gabriel Dumitru, Ben Hertzberg +20
cs.SEcs.AIarXiv:2602.00933v32026BugsInPy: A Database of Existing Bugs in Python Programs to Enable Controlled Testing and Debugging Studies
Ratnadira Widyasari, Sheng Qin Sim, Camellia Lok +13
cs.SEarXiv:2401.15481v12024LeetCodeDataset: A Temporal Dataset for Robust Evaluation and Efficient Training of Code LLMs
Yunhui Xia, Wei Shen, Yan Wang +5
cs.LGcs.CLcs.SEarXiv:2504.14655v12025ACECODER: Acing Coder RL via Automated Test-Case Synthesis
Huaye Zeng, Dongfu Jiang, Haozhe Wang +3
cs.SEcs.AIcs.CLarXiv:2502.01718v42025A Tale of Two Cities: Software Developers Working from Home During the COVID-19 Pandemic
Denae Ford, Margaret-Anne Storey, Thomas Zimmermann +6
cs.SEcs.CYcs.HCarXiv:2008.11147v32020Montage: a grid portal and software toolkit for science-grade astronomical image mosaicking
Joseph C. Jacob, Daniel S. Katz, G. Bruce Berriman +8
astro-ph.IMcs.DCcs.SEarXiv:1005.4454v12010