Software Engineering
Papers filed under cs.SE on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
601 to 660 of 1,389
SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution
Yuxiang Wei, Olivier Duchenne, Jade Copet +6
cs.SEcs.AIcs.CLarXiv:2502.18449v22025A Programming Paradigm for Spatiotemporal Composability
Yifan Shi, Wei Zhang, Tianyi Cui
cs.PLcs.SEarXiv:2608.25512v12026Testing and Evaluation of Agentic AI Systems In Military Command and Control
Ulysse Richard, Heather Frase, Sarah Cao +3
cs.SEcs.AIcs.CYarXiv:2608.20597v12026Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution
Anton Razzhigaev, Andrei Gritsaev, Andrei Kaznacheev +3
cs.SEcs.AIarXiv:2608.08311v22026SWE-Review: Closing the Loop on Issue Resolution with Agentic Code Review
Ruoyu Wang, Jierun Chen, Shaowei Wang +7
cs.SEarXiv:2607.06065v12026RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources
Yijia Fan, Zonglin Di, Zimo Wen +8
cs.SEcs.AIarXiv:2606.29538v42026Summaries:한국어Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses
Jiahang Lin, Shichun Liu, Chengjun Pan +8
cs.CLcs.SEarXiv:2604.25850v42026Shepherd: Enabling Programmable Meta-Agents via Reversible Agentic Execution Traces
Simon Yu, Derek Chong, Ananjan Nandi +4
cs.AIcs.PLcs.SEarXiv:2605.10913v32026ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents
Fanqing Meng, Lingxiao Du, Zijian Wu +46
cs.CVcs.SEarXiv:2604.23781v22026The Faiss library
Matthijs Douze, Alexandr Guzhva, Chengqi Deng +6
cs.LGcs.CVcs.SEarXiv:2401.08281v42024KernelBench: Can LLMs Write Efficient GPU Kernels?
Anne Ouyang, Simon Guo, Simran Arora +4
cs.LGcs.AIcs.PFarXiv:2502.10517v12025SWE-smith: Scaling Data for Software Engineering Agents
John Yang, Kilian Lieret, Carlos E. Jimenez +7
cs.SEcs.AIcs.CLarXiv:2504.21798v22025Automated Directed Fairness Testing
Sakshi Udeshi, Pryanshu Arora, Sudipta Chattopadhyay
cs.LGcs.AIcs.SEarXiv:1807.00468v22018SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Xiang Deng, Jeff Da, Edwin Pan +19
cs.SEcs.CLarXiv:2509.16941v22025DRACO: Fine-Grained Credit Assignment with Dynamic Rubrics for Long-Horizon Agent Training
Shubham Gandhi, Saurabh Goyal, Kiran Kate +1
cs.AIcs.LGcs.SEarXiv:2609.04094v12026Towards Behavior Tree-Guided Vulnerability Detection with Lightweight LLMs
Enna Basic, Alberto Giaretta
cs.CRcs.SEarXiv:2609.01758v12026iStar 2.0 Language Guide
Fabiano Dalpiaz, Xavier Franch, Jennifer Horkoff
cs.SEarXiv:1605.07767v32016Migrating to Cloud-Native Architectures Using Microservices: An Experience Report
Armin Balalaie, Abbas Heydarnoori, Pooyan Jamshidi
cs.SEcs.DCarXiv:1507.08217v12015Baldur: Whole-Proof Generation and Repair with Large Language Models
Emily First, Markus N. Rabe, Talia Ringer +1
cs.LGcs.LOcs.SEarXiv:2303.04910v22023Learning Program Embeddings to Propagate Feedback on Student Code
Chris Piech, Jonathan Huang, Andy Nguyen +3
cs.LGcs.NEcs.SEarXiv:1505.05969v12015Exploring and Unleashing the Power of Large Language Models in Automated Code Translation
Zhen Yang, Fang Liu, Zhongxing Yu +7
cs.SEcs.AIarXiv:2404.14646v22024Debug like a Human: A Large Language Model Debugger via Verifying Runtime Execution Step-by-step
Li Zhong, Zilong Wang, Jingbo Shang
cs.SEcs.AIcs.CLarXiv:2402.16906v62024GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning
Lakshya A Agrawal, Shangyin Tan, Dilara Soylu +14
cs.CLcs.AIcs.LGarXiv:2507.19457v22025Smart Contracts Claimed Vulnerable by the CVE Database, with Labels and Source Locations
Monika di Angelo, Gernot Salzer
cs.CRcs.SEarXiv:2609.01186v12026Code Transformation Rule Synthesis using LLMs: Potential and Limits
Axel Allain, Aymeric Blot, Djamel Eddine Khelladi +1
cs.SEarXiv:2609.03592v12026Graph-based, Self-Supervised Program Repair from Diagnostic Feedback
Michihiro Yasunaga, Percy Liang
cs.SEcs.CLcs.LGarXiv:2005.10636v22020Efficient SWE Agent Benchmarking via Trajectory-Aware Evaluation
Kefeng Duan, Dewu Zheng, Yanlin Wang +7
cs.SEcs.AIcs.CLarXiv:2609.01603v12026Adaptive Critical Token-Aware Retrieval for Repository-Level Code Generation
Kefeng Duan, Dewu Zheng, Yanlin Wang +8
cs.SEcs.AIcs.CLarXiv:2609.01601v12026Predicting Program Exit Code with LLMs and Programming Language Semantics
Lara Marinov, Aditya Thimmaiah, Jayanth Srinivasa +2
cs.PLcs.AIcs.CLarXiv:2609.00579v12026Simulation-based Adversarial Test Generation for Autonomous Vehicles with Machine Learning Components
Cumhur Erkan Tuncali, Georgios Fainekos, Hisahiro Ito +1
eess.SYcs.AIcs.SEarXiv:1804.06760v42018RosettaBitcoin: An Artifact-Backed Experience Report on Verification Infrastructure for Agent-Assisted Consensus Validators
Donavon Guyot
cs.SEarXiv:2609.01702v12026TreeGen: A Tree-Based Transformer Architecture for Code Generation
Zeyu Sun, Qihao Zhu, Yingfei Xiong +3
cs.LGcs.SEarXiv:1911.09983v22019Requirements Engineering for Machine Learning: Perspectives from Data Scientists
Andreas Vogelsang, Markus Borg
cs.LGcs.SEarXiv:1908.04674v12019Quantum Software Engineering: Landscapes and Horizons
Jianjun Zhao
cs.SEcs.PLquant-pharXiv:2007.07047v22020ToolGate: An Executable Acceptance Pipeline for Tool-Dependent Scientific Benchmark Construction
Ke Zhang, Yankang Liu, Roya Zandi +1
cs.AIcs.MScs.SEarXiv:2609.02067v12026On Testing Machine Learning Programs
Houssem Ben Braiek, Foutse Khomh
cs.SEarXiv:1812.02257v12018Summaries:한국어Automated Repair of Programs from Large Language Models
Zhiyu Fan, Xiang Gao, Martin Mirchev +2
cs.SEarXiv:2205.10583v42022MicroHECL: High-Efficient Root Cause Localization in Large-Scale Microservice Systems
Dewei Liu, Chuan He, Xin Peng +6
cs.SEarXiv:2103.01782v12021Runtime-Independent Persistent Agents: Preserving Identity, Memory, and Code Across Models, Harnesses, and Servers
Zhenyu Zhao, Roy Zhao
cs.SEcs.AIarXiv:2609.00546v12026Data Quality for Software Vulnerability Datasets
Roland Croft, M. Ali Babar, Mehdi Kholoosi
cs.SEarXiv:2301.05456v12023On Learning Meaningful Assert Statements for Unit Test Cases
Cody Watson, Michele Tufano, Kevin Moran +2
cs.SEarXiv:2002.05800v22020APIGen: Automated Pipeline for Generating Verifiable and Diverse Function-Calling Datasets
Zuxin Liu, Thai Hoang, Jianguo Zhang +14
cs.CLcs.AIcs.LGarXiv:2406.18518v12024Replacing Training with Memory: Listwise Selection for Text-to-SQL
Yeonseok Jeong, Soyoung Yoon, Seongjun Lee +1
cs.SEcs.AIcs.CLarXiv:2609.00834v12026ShikumiMiner: Mining Recurring Implementation Patterns in AI Codebases
Afsana Tasnim, Sheikh Motahar Naim
cs.SEarXiv:2609.02789v12026Easy over Hard: A Case Study on Deep Learning
Wei Fu, Tim Menzies
cs.SEcs.LGarXiv:1703.00133v22017LLM Hallucinations in Practical Code Generation: Phenomena, Mechanism, and Mitigation
Ziyao Zhang, Yanlin Wang, Chong Wang +2
cs.SEcs.AIcs.CLarXiv:2409.20550v22024On the Robustness of Code Generation Techniques: An Empirical Study on GitHub Copilot
Antonio Mastropaolo, Luca Pascarella, Emanuela Guglielmi +4
cs.SEarXiv:2302.00438v12023Towards Automating Code Review Activities
Rosalia Tufano, Luca Pascarella, Michele Tufano +2
cs.SEarXiv:2101.02518v42021Barriers to Using Static Application Security Testing (SAST) Tools: A Literature Review
Zachary Wadhams, Clemente Izurieta, Ann Marie Reinhold
cs.SEcs.CRarXiv:2609.01669v12026AVATAR : Fixing Semantic Bugs with Fix Patterns of Static Analysis Violations
Kui Liu, Anil Koyuncu, Dongsun Kim +1
cs.SEarXiv:1812.07270v32018Work-From-Home is Here to Stay: Call for Flexibility in Post-Pandemic Work Policies
Darja Smite, Nils Brede Moe, Jarle Hildrum +2
cs.SEcs.CYarXiv:2203.11136v12022From Prompting to Engineering: A Research Agenda for Prompt Engineering in Software Engineering
Vincenzo De Martino, Giovanna Broccia, Fabiano Pecorelli +6
cs.SEarXiv:2609.02248v12026Type Hints in Python Libraries and Frameworks: An Empirical Analysis of Adoption and Maintenance
Thiago Roberto Magalhães, Fabio Petrillo, João Eduardo Montandon
cs.SEarXiv:2609.02782v12026SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
John Yang, Carlos E. Jimenez, Alex L. Zhang +10
cs.CLcs.AIcs.SEarXiv:2410.03859v12024A Review of Software Quality Models for the Evaluation of Software Products
Jose P. Miguel, David Mauricio, Glen Rodriguez
cs.SEarXiv:1412.2977v12014Mining Idioms from Source Code
Miltiadis Allamanis, Charles Sutton
cs.SEarXiv:1404.0417v32014Happy software developers solve problems better: psychological measurements in empirical software engineering
Daniel Graziotin, Xiaofeng Wang, Pekka Abrahamsson
cs.SEcs.HCarXiv:1505.00922v12015Adversarial Sample Detection for Deep Neural Network through Model Mutation Testing
Jingyi Wang, Guoliang Dong, Jun Sun +2
cs.LGcs.SEstat.MLarXiv:1812.05793v22018Limitations of Agile Software Processes
Dan Turk, Robert France, Bernhard Rumpe
cs.SEarXiv:1409.6600v12014Automatic Semantic Augmentation of Language Model Prompts (for Code Summarization)
Toufique Ahmed, Kunal Suresh Pai, Premkumar Devanbu +1
cs.SEcs.LGarXiv:2304.06815v32023