Software Engineering

Papers filed under cs.SE on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

481 to 540 of 1,389

  1. SkillReducer: Optimizing LLM Agent Skills for Token Efficiency

    Yudong Gao, Zongjie Li, Yuanyuan Yuan +3

    cs.SEarXiv:2603.29919v22026
  2. An Analysis of the Search Spaces for Generate and Validate Patch Generation Systems

    Fan Long, Martin Rinard

    cs.SEarXiv:1602.05643v12016
  3. Harness Engineering in LLM Tool Use via Agent-Native Reusable Tool Primitives

    Haibo Jin, Suijin Wang, Xucheng Yu +2

    cs.SEcs.AIcs.CLarXiv:2609.01736v12026
  4. Understanding Software Engineering Agents: A Study of Thought-Action-Result Trajectories

    Islem Bouzenia, Michael Pradel

    cs.SEcs.AIarXiv:2506.18824v22025
  5. RepoAudit: An Autonomous LLM-Agent for Repository-Level Code Auditing

    Jinyao Guo, Chengpeng Wang, Xiangzhe Xu +2

    cs.SEcs.PLarXiv:2501.18160v32025
  6. StarCoder: may the source be with you!

    Raymond Li, Loubna Ben Allal, Yangtian Zi +64

    cs.CLcs.AIcs.PLarXiv:2305.06161v22023
  7. SWE-Milestone: Evaluating AI Agents on Continuous Software Evolution

    Gangda Deng, Zhaoling Chen, Zhongming Yu +11

    cs.SEcs.AIarXiv:2603.13428v42026
  8. Conformance Checking Based on Multi-Perspective Declarative Process Models

    Andrea Burattin, Fabrizio Maria Maggi, Alessandro Sperduti

    cs.SEcs.DBcs.LOarXiv:1503.04957v12015
  9. SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?

    Samuel Miserendino, Michele Wang, Tejal Patwardhan +1

    cs.LGcs.SEarXiv:2502.12115v42025
  10. Agentic Software Engineering: Foundational Pillars and a Research Roadmap

    Ahmed E. Hassan, Hao Li, Dayi Lin +4

    cs.SEcs.AIarXiv:2509.06216v32025
  11. Multi-Modal Attention Network Learning for Semantic Source Code Retrieval

    Yao Wan, Jingdong Shu, Yulei Sui +4

    cs.SEcs.PLarXiv:1909.13516v12019
  12. CATERPILLAR: A Business Process Execution Engine on the Ethereum Blockchain

    Orlenys López-Pintado, Luciano García-Bañuelos, Marlon Dumas +2

    cs.SEarXiv:1808.03517v32018
  13. Vibe Coding in Practice: Motivations, Challenges, and a Future Outlook -- a Grey Literature Review

    Ahmed Fawzy, Amjed Tahir, Kelly Blincoe

    cs.SEarXiv:2510.00328v12025
  14. Framework and Benchmark for Code-Driven Agentic Testing in Web Development

    Bin Hong, Zhenchao Zhang, Jiyuan He +2

    cs.SEarXiv:2609.00081v12026
  15. Beneath the Diff: Diagnosing and Mitigating Algorithmic Mode Collapse in Code-Level Autonomous Research Loops

    Bowei He, Weixu Zhang, Yili Jin +1

    cs.CLcs.SEarXiv:2609.00077v12026
  16. Probabilistic Model Checking of Autoregressive Neural Sequence Models

    Helge Spieker, Dennis Gross, Arnaud Gotlieb

    cs.SEcs.AIarXiv:2609.00838v12026
  17. Schwarz: Solver-Aware Agentic Program Verification

    Jingyu Ke, Ling-I Wu, Guoqiang Li

    cs.LOcs.SEarXiv:2608.30803v12026
  18. Developer Prioritization in Bug Repositories

    Jifeng Xuan, He Jiang, Zhilei Ren +1

    cs.SEarXiv:1704.04764v12017
  19. Make LLM a Testing Expert: Bringing Human-like Interaction to Mobile GUI Testing via Functionality-aware Decisions

    Zhe Liu, Chunyang Chen, Junjie Wang +5

    cs.SEarXiv:2310.15780v12023
  20. WiseSpec: Requirements-Driven Agents for Code Generation

    Zhao Tian

    cs.SEcs.AIarXiv:2609.00568v12026
  21. Does Fault Localization Beat a Fresh Attempt? A Placebo-Controlled Study of Test-Guided Code Repair

    Anik Jha

    cs.SEcs.AIcs.LGarXiv:2609.00854v12026
  22. On the Prospects of Dynamic LLM Conversations in Software Development

    Annemarie Wittig, Alina Mailach, Janet Siegmund +1

    cs.SEcs.AIarXiv:2608.30756v12026
  23. CWEval: Outcome-driven Evaluation on Functionality and Security of LLM Code Generation

    Jinjun Peng, Leyi Cui, Kele Huang +2

    cs.SEcs.CLcs.LGarXiv:2501.08200v12025
  24. Designing an Auditable LLM-Supported Workflow for Qualitative Thematic Analysis

    Nadia Jul Jeldtoft, Tariq Yousef

    cs.AIcs.SEarXiv:2608.30543v12026
  25. Towards a Reliable and Practical Eval Pipeline

    Emma Thuong Nguyen, Abhishek Ghose

    cs.AIcs.SEarXiv:2609.00805v12026
  26. MAGIS: LLM-Based Multi-Agent Framework for GitHub Issue Resolution

    Wei Tao, Yucheng Zhou, Yanlin Wang +3

    cs.SEcs.AIarXiv:2403.17927v22024
  27. Kevin: Multi-Turn RL for Generating CUDA Kernels

    Carlo Baronio, Pietro Marsella, Ben Pan +2

    cs.LGcs.AIcs.PFarXiv:2507.11948v12025
  28. Spec-Driven Development for Agentic Software Engineering: Harnessing Human-Agent Teamwork

    Jessica Diaz, Joaquin Gayoso, Andrea Cimminio +1

    cs.SEarXiv:2609.00252v12026
  29. PentestGPT: An LLM-empowered Automatic Penetration Testing Tool

    Gelei Deng, Yi Liu, Víctor Mayoral-Vilches +7

    cs.SEcs.CRarXiv:2308.06782v22023
  30. Oreo: Detection of Clones in the Twilight Zone

    Vaibhav Saini, Farima Farmahinifarahani, Yadong Lu +2

    cs.SEarXiv:1806.05837v12018
  31. Neural Transfer Learning for Repairing Security Vulnerabilities in C Code

    Zimin Chen, Steve Kommrusch, Martin Monperrus

    cs.SEcs.CRcs.LGarXiv:2104.08308v32021
  32. A Systematic Review of Unsupervised Learning Techniques for Software Defect Prediction

    Ning Li, Martin Shepperd, Yuchen Guo

    cs.SEarXiv:1907.12027v42019
  33. A Survey on Code Generation with LLM-based Agents

    Yihong Dong, Xue Jiang, Jiaru Qian +4

    cs.SEcs.AIcs.CLarXiv:2508.00083v22025
  34. Model Context Protocol (MCP) at First Glance: Studying the Security and Maintainability of MCP Servers

    Mohammed Mehedi Hasan, Hao Li, Emad Fallahzadeh +3

    cs.SEcs.ETarXiv:2506.13538v52025
  35. Software Engineering Challenges of Deep Learning

    Anders Arpteg, Björn Brinne, Luka Crnkovic-Friis +1

    cs.SEcs.AIcs.LGarXiv:1810.12034v12018
  36. Repository-Level Prompt Generation for Large Language Models of Code

    Disha Shrivastava, Hugo Larochelle, Daniel Tarlow

    cs.LGcs.AIcs.PLarXiv:2206.12839v32022
  37. On the Use of Agentic Coding: An Empirical Study of Pull Requests on GitHub

    Miku Watanabe, Hao Li, Yutaro Kashiwa +3

    cs.SEarXiv:2509.14745v32025
  38. Reliable LLM-Generated Programs for High-Energy Physics Experiments through Graph-Grounded Software Knowledge

    Yue Sun, Tong Liu, Yipu Liao +2

    cs.SEhep-exarXiv:2609.01095v12026
  39. CWM: An Open-Weights LLM for Research on Code Generation with World Models

    FAIR CodeGen team, Jade Copet, Quentin Carbonneaux +48

    cs.SEcs.AIcs.LGarXiv:2510.02387v12025
  40. Identifying Patch Correctness in Test-Based Program Repair

    Yingfei Xiong, Xinyuan Liu, Muhan Zeng +2

    cs.SEarXiv:1706.09120v32017
  41. SWE-rebench: An Automated Pipeline for Task Collection and Decontaminated Evaluation of Software Engineering Agents

    Ibragim Badertdinov, Alexander Golubev, Maksim Nekrashevich +6

    cs.SEcs.CLarXiv:2505.20411v22025
  42. OS-Harm: A Benchmark for Measuring Safety of Computer Use Agents

    Thomas Kuntz, Agatha Duzan, Hao Zhao +4

    cs.SEcs.LGarXiv:2506.14866v22025
  43. Can LLMs Replace Human Evaluators? An Empirical Study of LLM-as-a-Judge in Software Engineering

    Ruiqi Wang, Jiyu Guo, Cuiyun Gao +3

    cs.SEcs.AIarXiv:2502.06193v32025
  44. Technology Readiness Levels for AI & ML

    Alexander Lavin, Gregory Renard

    cs.SEcs.AIcs.LGarXiv:2006.12497v32020
  45. Beyond Locks and Thread IDs: Static Data Race Detection Off The Beaten Path (Extended Version)

    Daniel Bund, Julian Erhard, Michael Petter +1

    cs.PLcs.SEarXiv:2609.00246v12026
  46. AgOSS: A Dataset and Multi-Layer Characterization of Open-Source Agricultural Software

    Vatsal Dudhaiya, Mikhail Golovenchits, Aryan Banerjee +1

    cs.SEarXiv:2609.02591v12026
  47. Natural Emergent Misalignment from Reward Hacking in Production RL

    Monte MacDiarmid, Benjamin Wright, Jonathan Uesato +19

    cs.AIcs.SEarXiv:2511.18397v12025
  48. Engineering Reliable Coding Agents: Evaluating and Operating the System Around the Model

    Stephanie Jarmak

    cs.SEcs.AIarXiv:2608.13867v12026
  49. Automated Vulnerability Injection in Smart Contracts Using Large Language Models

    Luca Migliaccio, Roberto Natella, Naghmeh Ivaki +2

    cs.SEcs.AIcs.CRarXiv:2609.02624v12026
    Summaries:한국어
  50. POLYFLOW: A Neuro-Symbolic Framework for Static Cross-Language Information Flow Analysis

    Haoran Yang, Zhixuan Zhong, Jiawei Guo +1

    cs.CRcs.PLcs.SEarXiv:2608.29808v12026
  51. Repair Is Nearly Generation: Multilingual Program Repair with LLMs

    Harshit Joshi, José Cambronero, Sumit Gulwani +3

    cs.SEcs.AIcs.PLarXiv:2208.11640v32022
  52. Towards an Understanding of Large Language Models in Software Engineering Tasks

    Zibin Zheng, Kaiwen Ning, Qingyuan Zhong +5

    cs.SEarXiv:2308.11396v32023
  53. ReproAgent: Contract-Guided Paper-to-Code Reproduction

    Xue Hu, Zewei Pan, Zhongyuan Wang +3

    cs.AIcs.SEarXiv:2608.24291v12026
  54. REVERE: Reflective Evolving Research Engineer

    Balaji Dinesh Gangireddi, Aniketh Garikaparthi, Manasi Patwardhan +1

    cs.SEcs.AIarXiv:2603.20667v22026
  55. Seeing is Coding: On the Effectiveness of Vision Language Models in Code Understanding

    Yuling Shi, Chaoxiang Xie, Zhensu Sun +7

    cs.CLcs.SEarXiv:2602.01785v32026
  56. Is Self-Repair a Silver Bullet for Code Generation?

    Theo X. Olausson, Jeevana Priya Inala, Chenglong Wang +2

    cs.CLcs.AIcs.PLarXiv:2306.09896v52023
  57. The Empire, Long Divided, Must Unite: Architectural Convergence in Three LLM Agent Harnesses

    Dai Jiahong

    cs.SEcs.AIcs.CEarXiv:2608.23953v12026
  58. DreamBench-SWE: A Multi-Session Memory-Hygiene Benchmark for Software Agents

    Sarthak Singh

    cs.AIcs.SEarXiv:2608.20664v12026
  59. FlavourBench: Ranking Frontier Language Models with Executable Culinary Ground Truth

    Josef Chen, Erim Hayretci

    cs.AIcs.CYcs.LGarXiv:2608.20574v12026
  60. Resume Means Resume: A Machine-Checked Conformance Contract for Checkpoint, Interrupt, and Resume Semantics in Workflow Persistence Layers

    Sajjad Khan

    cs.LGcs.DCcs.LOarXiv:2608.03836v32026