Software Engineering
Papers filed under cs.SE on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
1,261 to 1,320 of 1,389
SWE-WebDevBench: Evaluating Coding Agent Application Platforms as Virtual Software Agencies
Siddhant Saxena, Nilesh Trivedi, Vinayaka Jyothi
cs.MAcs.SEarXiv:2605.04637v12026Model-Driven Development of Complex Software: A Research Roadmap
Robert France, Bernhard Rumpe
cs.SEarXiv:1409.6620v12014Do not copy and paste! Rewriting strategies for code retrieval
Andrea Gurioli, Federico Pennino, Maurizio Gabbrielli
cs.SEcs.AIarXiv:2605.08299v12026DeepTest: Automated Testing of Deep-Neural-Network-driven Autonomous Cars
Yuchi Tian, Kexin Pei, Suman Jana +1
cs.SEcs.AIcs.LGarXiv:1708.08559v22017DeepXplore: Automated Whitebox Testing of Deep Learning Systems
Kexin Pei, Yinzhi Cao, Junfeng Yang +1
cs.LGcs.CRcs.SEarXiv:1705.06640v42017MOCHA: Multi-Objective Chebyshev Annealing for Agent Skill Optimization
Md Mehrab Tanjim, Jayakumar Subramanian, Xiang Chen +6
cs.AIcs.LGcs.SEarXiv:2605.19330v12026CodeXGLUE: A Machine Learning Benchmark Dataset for Code Understanding and Generation
Shuai Lu, Daya Guo, Shuo Ren +19
cs.SEcs.CLarXiv:2102.04664v22021Towards Evaluation Engineering: An Empirical Study of ML Evaluation Harnesses in the Wild
Zhimin Zhao, Zehao Wang, Abdul Ali Bangash +2
cs.SEcs.AIcs.LGarXiv:2605.24213v12026Efficient and Scalable Provenance Tracking for LLM-Generated Code Snippets
Andrea Gurioli, Davide D'Ascenzo, Federico Pennino +2
cs.SEcs.AIcs.IRarXiv:2605.28510v22026Benchmarks are Not Enough: RAMP for Runtime Assessing of Agentic Models in Production Systems
Yipeng Ouyang, Xin Huang, Bingjie Liu +3
cs.SEcs.AIarXiv:2605.27492v12026REPOT: Recoverable Program-of-Thought via Checkpoint Repair
Parsa Mazaheri
cs.SEcs.AIcs.CLarXiv:2605.30052v12026DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence
Daya Guo, Qihao Zhu, Dejian Yang +10
cs.SEcs.CLcs.LGarXiv:2401.14196v22024A Prompt Pattern Catalog to Enhance Prompt Engineering with ChatGPT
Jules White, Quchen Fu, Sam Hays +6
cs.SEcs.AIarXiv:2302.11382v12023LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
Naman Jain, King Han, Alex Gu +7
cs.SEcs.CLcs.LGarXiv:2403.07974v22024Combinatorial Synthesis: Scaling Code RLVR via Atomic Decomposition and Recombination
Jiasheng Zheng, Boxi Cao, Boxi Yu +6
cs.CLcs.SEarXiv:2605.31058v12026FVSpec: Real-World Property-Based Tests as Lean Challenges
Quinn Dougherty, Max von Hippel, Simon Henniger +2
cs.SEcs.AIarXiv:2606.01008v22026Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation
Jiawei Liu, Chunqiu Steven Xia, Yuyao Wang +1
cs.SEcs.CLcs.LGarXiv:2305.01210v32023Lean4Agent: Formal Modeling and Verification for Agent Workflow and Trajectory
Ruida Wang, Jerry Huang, Pengcheng Wang +3
cs.AIcs.LGcs.LOarXiv:2606.06523v22026Empirical Study on the Characteristics and Evolution of AI-usage in GitHub Repositories: Evidence from Code Comments
Abdullah Al Mujahid, Preetha Chatterjee, Mia Mohammad Imran
cs.SEarXiv:2606.06843v12026PIPE-Cypher: Automatic Enterprise Benchmark Generation for Text-to-Cypher Systems
Suraj Ranganath, Anish Raghavendra
cs.LGcs.AIcs.DBarXiv:2606.08481v12026Context Aware Computing for The Internet of Things: A Survey
Charith Perera, Arkady Zaslavsky, Peter Christen +1
cs.SEcs.HCarXiv:1305.0982v12013Lean Refactor: Multi-Objective Controllable Proof Optimization via Agentic Strategy Search
Jialin Lu, Soonho Kong, Rodrigo Stehling +4
cs.LOcs.AIcs.CLarXiv:2605.20244v22026SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?
Zhipeng Xu, Jiahao Lu, Yining Zheng +2
cs.CLcs.SEarXiv:2608.19799v12026What Makes Software Issue Resolution Tasks Difficult for Agents?
Ebtesam Al-Haque, Brittany Johnson
cs.SEcs.AIcs.CLarXiv:2608.18280v12026ARIS: Autonomous Research via Adversarial Multi-Agent Collaboration
Ruofeng Yang, Yongcan Li, Shuai Li
cs.SEcs.AIarXiv:2605.03042v12026BenchEvolver: Frontier Task Synthesis via Solution-Centric Evolution
Yangzhen Wu, Aaron J. Li, Wenjie Ma +10
cs.SEcs.AIcs.CLarXiv:2606.01286v12026Programming with Data: Test-Driven Data Engineering for Self-Improving LLMs from Raw Corpora
Chenkai Pan, Xinglong Xu, Yuhang Xu +6
cs.SEcs.AIarXiv:2604.24819v12026DarwinX: Evolving Agent Harnesses Through Natural Selection
Yifan Zhang, Yutong Dai, Juntao Tan +9
cs.NEcs.AIcs.LGarXiv:2608.07545v12026SemPlan: Benchmarking Structured Semantic Planning for LLM-Based Queries over Enterprise Data
Bruno Santos Teixeira
cs.AIcs.SEarXiv:2608.13612v12026SkillForge: Self-Distilling Agents for Project-Specific Issue Resolution
Silin Chen, Han Li, Xiaodong Gu +2
cs.SEcs.AIarXiv:2608.18933v12026SemaPLC: A Project-Grounded, Verification-Gated Agent Harness for PLC Code Generation
Yanlun Tu, Huacan Wang, Ziyue Zhou +10
cs.SEarXiv:2608.18565v12026Adversarial Review: Structured Disagreement for Grounded Agentic Code Review
Eric S. Qiu, Joyce Gill
cs.AIcs.SEarXiv:2608.18167v12026When Agents Coordinate: Measuring Coordination in Multi-Agent AI Coding
Giuseppe Destefanis, Tomaso Aste
cs.AIcs.SEarXiv:2608.16801v12026The Recall Trap: A Recall-Maximizing Retriever Configuration Reduces Issue Resolution in Fixed-Budget Code Context
Alexander Adkins, Teimuraz Trapaidze
cs.SEcs.CLcs.IRarXiv:2608.14838v12026Second Thought: Reasoning in Parallel as LLM Agents Act and Observe
Zhensu Sun, Chengran Yang, Yunbo Lyu +2
cs.AIcs.SEarXiv:2608.13667v12026Grammar-Constrained Decoding Can Jailbreak LLMs into Generating Malicious Code
Yitong Zhang, Shiteng Lu, Jia Li
cs.CRcs.AIcs.CLarXiv:2606.11817v12026When Gradients Collide: Failure Modes of Multi-Objective Prompt Optimization for LLM Judges
Parth Darshan, Abhishek Divekar
cs.CLcs.AIcs.LGarXiv:2605.26046v22026SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
Carlos E. Jimenez, John Yang, Alexander Wettig +4
cs.CLcs.AIcs.SEarXiv:2310.06770v32023MicroPython and CircuitPython: Pythons Quiet Takeover of IoT and Robotics
Sayed Mahbub Hasan Amiri, Atiar Zahan
cs.PLcs.CLcs.SEarXiv:2608.18160v12026Pre-Compiled Pipeline Shards for Distributed LLM Inference on Intel AI PC Fleets
Tate Berenbaum, Muthaiah Venkatachalam
cs.DCcs.AIcs.SEarXiv:2608.19147v12026One Gate Is Not Enough: Composing Stateful Pre-Action Controls for Agentic AI
Gaston Besanson
cs.SEcs.AIcs.CYarXiv:2608.18360v12026ImageJ2: ImageJ for the next generation of scientific image data
Curtis T. Rueden, Johannes Schindelin, Mark C. Hiner +4
cs.SEq-bio.QMarXiv:1701.05940v42017Grouping the Stochastic Machine: Precision, Not Capability, as the Frontier Metric for AI Systems
George Andrikopoulos
cs.AIcs.CYcs.LGarXiv:2608.19140v12026Tuning the Stochastic Machine: A Systems Engineer's Operating Model for Human-AI Engineering
George Andrikopoulos
cs.AIcs.SEarXiv:2608.19125v12026Metrics That Write Themselves: Evolving an Evaluator from Its Own Blind Spots
Xing Zhang, Yanwei Cui, Guanghui Wang +2
cs.AIcs.CLcs.SEarXiv:2608.18744v12026Securing AI-Generated Code: A Just-in-Time Vulnerability Detection and Remediation Pipeline
Mikhail Surikov
cs.CRcs.AIcs.SEarXiv:2608.16187v12026Too Sure to Be Safe: Model Calibration for Reliable Log Anomaly Detection
Bin Li, Dongdong Wang, Siyang Lu
cs.LGcs.AIcs.SEarXiv:2608.17965v12026What Aggregate Scores Miss: Measuring Item-Level Regressions in Commercial LLM API Migrations
Xiaonan Xu, Wenjing Wu
cs.SEcs.AIcs.CLarXiv:2608.17719v12026No Resource, No Benchmarks, No Problem? Evaluating and Improving LLMs for Code Generation in No-Resource Languages
Alessandro Giagnorio, Alberto Martin-Lopez, Gabriele Bavota
cs.SEarXiv:2606.16827v12026A Benchmark and Framework for Evaluating Next Action Predictions in Spreadsheets
Tejas Agrawal, Vu Le, Sumit Gulwani +1
cs.SEcs.AIcs.HCarXiv:2606.13802v12026When Uncertainty Isn't Enough: An Empirical Study of Self-Correction in Code Generation
Pranav Rakasi, Maanas Lalwani, Arnav Srivastava +4
cs.AIcs.LGcs.SEarXiv:2608.14659v12026When Agentic Executions Fail: Detecting and Localizing Runtime Faults from Telemetry
Chenkai Zhang, Yiran Li, Yifang Tian +2
cs.AIcs.SEarXiv:2608.14680v12026Summaries:한국어Evaluating Agentic Code Repair Capabilities in Distributed Systems
Yibo Yan, Huijuan Wang, Junzhou He +4
cs.SEcs.AIcs.DCarXiv:2608.14863v12026PandasCorpus: A Resource of Real-World Pandas Workflows and Usage Patterns
Syrym Abdikhan, Mazhar Hameed
cs.SEcs.LGarXiv:2608.14742v12026WeSCE: A Benchmark for Measuring Security Drift in LLM-Driven Code Editing
Zhiyu Zhang, Tingyue Wen, Senke Sun +2
cs.CRcs.AIcs.SEarXiv:2608.15092v12026FAPO: Fully Automated Prompt Optimization of Multi-Step LLM Pipelines
Paul Kassianik, Baturay Saglam, Huaibo Zhao +4
cs.SEcs.AIarXiv:2606.19605v22026OpenRath: Session-Centered Runtime State for Agent Systems
Fukang Wen, Zhijie Wang, Ruilin Xu
cs.SEcs.PLarXiv:2606.19409v12026An Exploratory Case Study of LLM-Assisted Refactoring and Gameplay Feature Generation in an Endless Runner Game
Jan Wunderlich, Markus Kleffmann, Sebastian Lempert
cs.SEcs.AIarXiv:2606.21171v12026Causal Discovery in the Era of Agents
Yujia Zheng, Vishal Verma, Mantej Gill +3
cs.AIcs.LGcs.SEarXiv:2606.23608v12026EnterpriseClawBench: Benchmarking Agents from Real Workplace Sessions
Jincheng Zhong, Weizhi Wang, Che Jiang +5
cs.CLcs.SEarXiv:2606.23654v12026