Software Engineering

Papers filed under cs.SE on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

1,141 to 1,200 of 1,389

  1. MEnvAgent: Scalable Polyglot Environment Construction for Verifiable Software Engineering

    Chuanzhe Guo, Jingjing Wu, Sijun He +10

    cs.SEcs.AIarXiv:2601.22859v32026
  2. AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces

    Sungho Park, Wonjoong Kim, Rongyuan Tan +10

    cs.AIcs.CLcs.LGarXiv:2608.23041v12026
  3. Formalizing and Automating Fine-Grained Move Refactorings Across Methods

    Kota Yasuhara, Shinpei Hayashi

    cs.SEarXiv:2608.23377v12026
  4. Execution-Anchored Hallucination Calibration Reranking for Verilog Code Generation

    Guang Yang, Xing Hu, Xiang Chen +2

    cs.SEcs.ARarXiv:2608.22938v12026
  5. DIVE: Scaling Diversity in Agentic Task Synthesis for Generalizable Tool Use

    Aili Chen, Chi Zhang, Junteng Liu +11

    cs.AIcs.SEarXiv:2603.11076v12026
  6. An Empirical Study of the TianoCore Community

    Nazanin Siavash, Connor Glosner, Ayushi Sharma +4

    cs.SEarXiv:2608.23280v12026
  7. IQuest-Coder-V1 Technical Report

    Jian Yang, Wei Zhang, Shawn Guo +35

    cs.AIcs.CLcs.SEarXiv:2603.16733v12026
  8. ContractFuzzer: Fuzzing Smart Contracts for Vulnerability Detection

    Bo Jiang, Ye Liu, W. K. Chan

    cs.SEcs.CRarXiv:1807.03932v22018
  9. Disagree to Explore, Agree to Commit: Routing-Guided Test-Time Scaling for Software Agents

    Kang Chen, Junjie Nian, Yixin Cao +1

    cs.AIcs.SEarXiv:2608.22191v12026
  10. Prompt Injection attack against LLM-integrated Applications

    Yi Liu, Gelei Deng, Yuekang Li +9

    cs.CRcs.AIcs.CLarXiv:2306.05499v32023
  11. Repo2Skill-Evo: Repository Skills Go Stale in Silence

    Chenyuan Duan, Ge Shi, Zineng Mao +10

    cs.AIcs.SEarXiv:2608.21964v12026
  12. graph2vec: Learning Distributed Representations of Graphs

    Annamalai Narayanan, Mahinthan Chandramohan, Rajasekar Venkatesan +3

    cs.AIcs.CLcs.CRarXiv:1707.05005v12017
  13. On Randomness in Agentic Evals

    Bjarni Haukur Bjarnason, André Silva, Martin Monperrus

    cs.LGcs.AIcs.SEarXiv:2602.07150v32026
  14. GameDevBench: Evaluating Agentic Capabilities Through Game Development

    Wayne Chi, Yixiong Fang, Arnav Yayavaram +8

    cs.AIcs.CLcs.SEarXiv:2602.11103v22026
  15. FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance

    Lingjiao Chen, Matei Zaharia, James Zou

    cs.LGcs.AIcs.CLarXiv:2305.05176v12023
  16. Guidelines to Prompt Large Language Models for Code Generation: An Empirical Characterization

    Alessandro Midolo, Alessandro Giagnorio, Fiorella Zampetti +3

    cs.SEarXiv:2601.13118v12026
  17. ToolPRMBench: Evaluating and Advancing Process Reward Models for Tool-using Agents

    Dawei Li, Yuguang Yao, Zhen Tan +2

    cs.AIcs.SEarXiv:2601.12294v12026
  18. Advances and Frontiers of LLM-based Issue Resolution in Software Engineering: A Comprehensive Survey

    Caihua Li, Lianghong Guo, Yanlin Wang +9

    cs.SEcs.CLarXiv:2601.11655v12026
  19. AACR-Bench: Evaluating Automatic Code Review with Holistic Repository-Level Context

    Lei Zhang, Yongda Yu, Minghui Yu +11

    cs.SEcs.AIarXiv:2601.19494v32026
  20. ReflexiCoder: Teaching Large Language Models to Self-Reflect on Generated Code and Self-Correct It via Reinforcement Learning

    Juyong Jiang, Jiasi Shen, Sunghun Kim +3

    cs.CLcs.LGcs.SEarXiv:2603.05863v22026
  21. ASA: Backbone-Training-Free Representation Engineering for Tool-Calling Agents

    Youjin Wang, Run Zhou, Yingjie Ma +6

    cs.SEcs.AIarXiv:2602.04935v32026
  22. Composable Trust Infrastructure for Manufacturing Knowledge Graphs: Cross-System Provenance, Temporal Reasoning, and Decision Traceability

    Grama Chethan

    cs.AIcs.SEarXiv:2608.21418v12026
  23. Confident and Wrong: Silent Semantic Failures in Coding Agents

    Aman Mehta

    cs.SEcs.AIarXiv:2603.25764v32026
  24. InCoder-32B: Code Foundation Model for Industrial Scenarios

    Jian Yang, Wei Zhang, Jiajun Wu +25

    cs.SEcs.AIarXiv:2603.16790v32026
  25. Machine Learning Testing: Survey, Landscapes and Horizons

    Jie M. Zhang, Mark Harman, Lei Ma +1

    cs.LGcs.AIcs.SEarXiv:1906.10742v22019
  26. KAPSO: A Knowledge-grounded framework for Autonomous Program Synthesis and Optimization

    Alireza Nadafian, Alireza Mohammadshahi, Majid Yazdani

    cs.AIcs.CLcs.SEarXiv:2601.21526v22026
  27. A Comprehensive Survey on Fog Computing: State-of-the-art and Research Challenges

    Carla Mouradian, Diala Naboulsi, Sami Yangui +3

    cs.DCcs.SEarXiv:1710.11001v32017
  28. MegaFlow: Large-Scale Distributed Orchestration System for the Agentic Era

    Lei Zhang, Mouxiang Chen, Ruisheng Cao +16

    cs.DCcs.SEarXiv:2601.07526v22026
  29. Prime Agent: A Self-Improving RLM Harness

    Seth Karten, Alex L. Zhang, Kevin Thomas +8

    cs.AIcs.CLcs.SEarXiv:2608.23552v12026
  30. Test-Driven AI Agent Definition (TDAD): Compiling Tool-Using Agents from Behavioral Specifications

    Tzafrir Rehan

    cs.SEcs.AIarXiv:2603.08806v12026
  31. Rethinking the Value of Agent-Generated Tests for LLM-Based Software Engineering Agents

    Zhi Chen, Zhensu Sun, Yuling Shi +4

    cs.SEcs.AIarXiv:2602.07900v22026
  32. InCoder: A Generative Model for Code Infilling and Synthesis

    Daniel Fried, Armen Aghajanyan, Jessy Lin +7

    cs.SEcs.CLcs.LGarXiv:2204.05999v32022
  33. CodeBLEU: a Method for Automatic Evaluation of Code Synthesis

    Shuo Ren, Daya Guo, Shuai Lu +7

    cs.SEcs.CLarXiv:2009.10297v22020
  34. SWE-Universe: Scale Real-World Verifiable Environments to Millions

    Mouxiang Chen, Lei Zhang, Yunlong Feng +15

    cs.SEcs.AIarXiv:2602.02361v12026
  35. Learning to Represent Programs with Graphs

    Miltiadis Allamanis, Marc Brockschmidt, Mahmoud Khademi

    cs.LGcs.AIcs.PLarXiv:1711.00740v32017
  36. OpenHands: An Open Platform for AI Software Developers as Generalist Agents

    Xingyao Wang, Boxuan Li, Yufan Song +21

    cs.SEcs.AIcs.CLarXiv:2407.16741v32024
  37. Agentic Code Reasoning

    Shubham Ugare, Satish Chandra

    cs.SEcs.AIcs.PLarXiv:2603.01896v22026
  38. Computer-Using World Model

    Yiming Guan, Rui Yu, John Zhang +15

    cs.SEarXiv:2602.17365v12026
  39. Terminal Agents Suffice for Enterprise Automation

    Patrice Bechard, Orlando Marquez Ayala, Emily Chen +5

    cs.SEcs.AIcs.CLarXiv:2604.00073v32026
  40. Vision2Web: A Hierarchical Benchmark for Visual Website Development with Agent Verification

    Zehai He, Wenyi Hong, Zhen Yang +4

    cs.SEcs.AIarXiv:2603.26648v32026
  41. Slither: A Static Analysis Framework For Smart Contracts

    Josselin Feist, Gustavo Grieco, Alex Groce

    cs.SEcs.CRarXiv:1908.09878v12019
  42. An Overview on Smart Contracts: Challenges, Advances and Platforms

    Zibin Zheng, Shaoan Xie, Hong-Ning Dai +4

    cs.SEcs.DCarXiv:1912.10370v12019
  43. Human-AI Synergy in Agentic Code Review

    Suzhen Zhong, Shayan Noei, Ying Zou +1

    cs.SEarXiv:2603.15911v12026
  44. daVinci-Dev: Agent-native Mid-training for Software Engineering

    Ji Zeng, Dayuan Fu, Tiantian Mi +14

    cs.SEcs.AIarXiv:2601.18418v22026
  45. Blockchain for Internet of Things: A Survey

    Hong-Ning Dai, Zibin Zheng, Yan Zhang

    cs.NIcs.DCcs.SEarXiv:1906.00245v52019
  46. Questioning the AI: Informing Design Practices for Explainable AI User Experiences

    Q. Vera Liao, Daniel Gruen, Sarah Miller

    cs.HCcs.AIcs.LGarXiv:2001.02478v32020
  47. Spike-Killer: Evidence-Gated LLM Assistance for Safe Performance Diagnosis on a Real Windows Workstation

    Baocheng Zeng, Jinhao Yang

    cs.SEarXiv:2608.21069v12026
  48. Toward Understanding Operating System Defects

    Hongyao Zuo, Jiali Li, Jiajun Jiang

    cs.SEarXiv:2608.20643v12026
  49. An Extensive Empirical Study on Code Translation Technique

    Ruihang Fan, Jiajun Jiang, Xinpeng Wang +3

    cs.SEcs.PLarXiv:2608.20776v12026
  50. BeyondSWE: Can Current Code Agent Survive Beyond Single-Repo Bug Fixing?

    Guoxin Chen, Fanzhe Meng, Jiale Zhao +12

    cs.CLcs.SEarXiv:2603.03194v22026
  51. SERA: Soft-Verified Efficient Repository Agents

    Ethan Shen, Daniel Tormoen, Saurabh Shah +2

    cs.CLcs.LGcs.SEarXiv:2601.20789v32026
  52. Benchmarking Reward Hack Detection in Code Environments via Contrastive Analysis

    Darshan Deshpande, Anand Kannappan, Rebecca Qian

    cs.SEcs.AIcs.LGarXiv:2601.20103v12026
  53. The Substitution Escrow Threshold: When "Compatible With" Becomes Safe Enough to Buy

    Amadeus Brandes

    cs.SEcs.CYarXiv:2608.21221v12026
  54. SWE-Master: Unleashing the Potential of Software Engineering Agents via Post-Training

    Huatong Song, Lisheng Huang, Shuang Sun +11

    cs.SEcs.CLarXiv:2602.03411v22026
  55. An introduction to Docker for reproducible research, with examples from the R environment

    Carl Boettiger

    cs.SEarXiv:1410.0846v12014
  56. SWE-rebench V2: Language-Agnostic SWE Task Collection at Scale

    Ibragim Badertdinov, Maksim Nekrashevich, Anton Shevtsov +1

    cs.SEcs.CLarXiv:2602.23866v22026
  57. Beyond Fault Localization: A Trajectory-Level Study of LLM Agents for Microservice Root Cause Analysis

    Qisheng Lu, Aoyang Fang, Junjielong Xu +5

    cs.SEarXiv:2608.21310v12026
  58. LongCLI-Bench: A Preliminary Benchmark and Study for Long-horizon Agentic Programming in Command-Line Interfaces

    Yukang Feng, Jianwen Sun, Zelai Yang +16

    cs.SEcs.MAarXiv:2602.14337v22026
  59. AutoWebWorld: Synthesizing Infinite Verifiable Web Environments via Finite State Machines

    Yifan Wu, Yiran Peng, Yiyu Chen +12

    cs.AIcs.SEarXiv:2602.14296v12026
  60. SWE-World: Building Software Engineering Agents in Docker-Free Environments

    Shuang Sun, Huatong Song, Lisheng Huang +11

    cs.SEcs.CLarXiv:2602.03419v12026