Software Engineering

Papers filed under cs.SE on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.

Search paper metadata (including unsummarized papers)

61 to 120 of 1,389

  1. Identifying Implementation Bugs in Machine Learning based Image Classifiers using Metamorphic Testing

    Anurag Dwarakanath, Manish Ahuja, Samarth Sikand +4

    cs.SEcs.LGarXiv:1808.05353v12018
  2. Practical Program Repair via Bytecode Mutation

    Ali Ghanbari, Lingming Zhang

    cs.SEarXiv:1807.03512v12018
  3. Measuring Coding Challenge Competence With APPS

    Dan Hendrycks, Steven Basart, Saurav Kadavath +8

    cs.SEcs.CLcs.LGarXiv:2105.09938v32021
  4. ATIBA: Grounded Integrity and Quality Checking for Research Papers

    Veli Karakaya, Semih Çağlar, Yusuf Yiğit Korkmaz +1

    cs.SEarXiv:2609.04123v12026
  5. SpikingJelly: An open-source machine learning infrastructure platform for spike-based intelligence

    Wei Fang, Yanqi Chen, Jianhao Ding +7

    cs.NEcs.LGcs.SEarXiv:2310.16620v12023
  6. Virtual Testing of Automated Driving Systems through Credible Simulations

    Riccardo Dona, Espedito Rusciano, Biagio Ciuffo

    cs.ROcs.SEarXiv:2609.03760v12026
  7. AI Harness Engineering: A Runtime Substrate for Foundation-Model Software Agents

    Hailin Zhong, Shengxin Zhu

    cs.SEcs.AIarXiv:2605.13357v12026
  8. Agentic Refactoring: An Empirical Study of AI Coding Agents

    Kosei Horikawa, Hao Li, Yutaro Kashiwa +3

    cs.SEarXiv:2511.04824v12025
  9. RefDiff: Detecting Refactorings in Version Histories

    Danilo Silva, Marco Tulio Valente

    cs.SEarXiv:1704.01544v12017
  10. CodeScore: Evaluating Code Generation by Learning Code Execution

    Yihong Dong, Jiazheng Ding, Xue Jiang +3

    cs.SEarXiv:2301.09043v42023
  11. Can AI Remediate Backend Failures Safely? GuardedAct with Blast-Radius-Aware Sandboxing

    Wanrong Cai, Tianyu Yu, Shaorui Pi +2

    cs.DCcs.SEarXiv:2609.11264v12026
  12. SemVerBench: Benchmarking LLM Comprehension of Version-Constraint Resolution Semantics

    Qibai Chen, Zeming Liu

    cs.AIcs.SEarXiv:2609.11180v12026
  13. BenchShield: Formal Model-Backed Instrumentation for Reward Integrity in LLM-Agent Evaluation Infrastructure

    Shenghan Zheng, Zonglin Di, Yimin Liu +19

    cs.CRcs.AIcs.SEarXiv:2609.11028v12026
    Summaries:한국어
  14. Grounding Agent Memory: Environment-Probing Curation for Enterprise Agents

    Susheel Suresh, Hazel Mak, Sahil Bhatnagar +2

    cs.AIcs.SEarXiv:2609.11060v12026
  15. Decoupling Readiness from Release for Tail-Aware Scheduling of Agentic LLM Workflows

    Bochao Feng, Jianjiang Li, Haojie Wang +4

    cs.AIcs.SEarXiv:2609.10964v12026
  16. DeFiFusion: Combining Transaction Events with Smart Contracts to Detect Price Manipulation Attacks

    Rui Cao, Shaojing Fan, Liming Fang +3

    cs.CRcs.AIcs.SEarXiv:2609.11008v12026
  17. AspisAI: A Canonical, Machine-Interpretable Governance Framework for Automated Multi-Standard Compliance Monitoring

    Tsafac Nkombong Regine Cyrille, Hasan Dag, Reiner Creutzburg +1

    cs.CRcs.CYcs.SEarXiv:2609.10881v12026
  18. An analysis of the relationship of input metrics

    Addison Crump

    cs.SEcs.FLarXiv:2609.11824v12026
  19. Beyond Static Guarantees: Measuring the Static-Pass Dynamic-Fail Gap in Security-Sensitive and LLM-Generated Python Code

    Jessica Pourleyli, Maitreyee Das Urmi, Glaucia Melo

    cs.CRcs.AIcs.SEarXiv:2609.10762v12026
  20. A2ABreak: Systematic Security Analysis of the A2A Protocol

    Alireza Lotfi, Mirza Masfiqur Rahman, Imtiaz Karim +1

    cs.CRcs.SEarXiv:2609.10871v12026
  21. Towards a Deterministic Math Solver for Clinical Language Models

    Felipe Ocampo Osorio, Sebastián Andrés Cajas Ordoñez, Maximin Lange +5

    cs.AIcs.SEarXiv:2609.10728v12026
  22. Reproducibility in the Age of Agentic AI: Context Engineering at the Timescale of a Codebase

    Lorena A. Barba

    cs.SEcs.CYarXiv:2609.11728v12026
  23. PRISMA-LLM: An Empirical Reporting Framework for AI-Assisted Systematic Reviews

    Miguel Zabaleta, Baihan Lin

    cs.SEarXiv:2609.11559v12026
  24. Agent-Integrated Software: Interaction Contracts and Continuous Assurance

    Shengcheng Yu, Chunrong Fang, Zhenyu Chen

    cs.SEcs.AIarXiv:2609.11381v12026
  25. Deep Learning-based Bug Triage System

    Sourabh Pal

    cs.SEarXiv:2609.11420v12026
  26. ChurnBench: A Drift-Aware Benchmark Demonstrating That Refresh Scheduling, Not Cache Age, Governs Staleness in Agentic AI

    Vivek Kumar Singh, Preeti Priyam

    cs.SEarXiv:2609.11515v12026
  27. Ecdysis: Efficient and Effective Training of Runtime Harnesses for LLM Agents

    Ruiqing Yue, Yu Cui, Zhuoyu Sun +11

    cs.SEcs.AIarXiv:2609.11677v12026
  28. TripleBound: Triplet-Guided Heterogeneous Graph Learning for Microservice Decomposition

    Mineth Weerasinghe, Himindu Kularathne, Methmini Madhushika +4

    cs.SEarXiv:2609.11212v12026
  29. A Model-Centric DevOps Architecture for DEVS-Based Digital Twin Simulation Services

    Arnis Lektauers, Gusts Linkevičs, Guntis Mosāns +2

    cs.SEcs.CEcs.DCarXiv:2609.11122v12026
  30. CoSTAR: Data Synthesis-Driven Constraint-Aware COBOL Section Summarization for Legacy System Modernization

    Hao Lin, He Jiang, Xiaochen Li +4

    cs.SEarXiv:2609.11332v12026
  31. FST Pay: Deterministic Safety-Gated Architecture for Youth Digital Payments

    Shaikh Mohammed Burhan, Syed Farhaan Quadri, Tabassum Nahid Sultana

    cs.SEcs.CRarXiv:2609.11195v12026
  32. Exploring the Role of Security Experience and ChatGPT Usage Strategies on Secure Software Engineering Education

    Alessio Ferrari, Minh An Nguyen, Kushal Ramkumar +1

    cs.SEarXiv:2609.11303v12026
  33. SaltBench: A Referee-Gated Protocol for Measuring Method Effects in Machine-Checked Software Work

    Jason Hickey

    cs.SEcs.LOarXiv:2609.11076v12026
  34. RCL: A Retrieval-Confidence Layer for Detecting Insufficient Context in Enterprise Retrieval-Augmented Code Generation

    Chandra Mohan Ravuri

    cs.SEarXiv:2609.11023v12026
  35. Engineering Reliable Commit Gates for Agentic AI: Cost-Aware Verification Portfolios under Common-Mode Data Failures

    Zihao Zheng, Baichuan Li, Junyi Yao +1

    cs.SEarXiv:2609.10969v12026
  36. What a Random Draw from the MCP Registry Contains, and What Tool-Use Benchmarks Contain Instead

    Haseeb Mohammed Afsar

    cs.SEcs.AIarXiv:2609.10962v12026
  37. LLMVul: A Vulnerability-Labeled Dataset of LLM-Generated C/C++ Functions from Real Production Repositories

    Mohammad Farhad, Shuvalaxmi Dass

    cs.SEarXiv:2609.10945v12026
  38. When Passing Tests Hides Vulnerabilities: An Empirical Study of Silent Failures in Agentic Systems

    Wenji Bai, Muhammad Waseem, Zeeshan Rasheed +2

    cs.SEcs.CRarXiv:2609.10548v12026
  39. Generative AI for trustworthy systems - Towards a health check model

    Jan Bosch, Rick Kazman, Henry Muccini +1

    cs.SEarXiv:2609.10595v12026
  40. ReqEvolve: User-Oriented Software Self-Evolution through Automatic Requirement Interpretation

    Md Asif Iqbal Fahim, Alessio Ferrari

    cs.SEarXiv:2609.10590v12026
  41. Governed Human-AI Prioritization Under Uncertainty: Adaptive Estimation and Dependency-Constrained Portfolio Selection

    Azzeddine Ihsine, Sara Ihsine

    cs.SEarXiv:2609.10648v12026
  42. AI Safety: Not Optional, Not Later

    Qinghua Lu, Yoshua Bengio

    cs.SEarXiv:2609.10630v12026
  43. Agent READMEs: An Empirical Study of Context Files for Agentic Coding

    Worawalan Chatlatanagulchai, Hao Li, Yutaro Kashiwa +8

    cs.SEarXiv:2511.12884v22025
  44. VeriGuard: Enhancing LLM Agent Safety via Verified Code Generation

    Lesly Miculicich, Mihir Parmar, Hamid Palangi +4

    cs.SEcs.AIcs.CRarXiv:2510.05156v12025
  45. SWE-QA: Can Language Models Answer Repository-level Code Questions?

    Weihan Peng, Yuling Shi, Yuhang Wang +3

    cs.CLcs.PLcs.SEarXiv:2509.14635v22025
  46. SWE-EVO: Benchmarking Coding Agents in Long-Horizon Software Evolution Scenarios

    Tue Le, Minh V. T. Thai, Dung Nguyen Manh +2

    cs.SEcs.AIcs.MAarXiv:2512.18470v62025
  47. LongCodeZip: Compress Long Context for Code Language Models

    Yuling Shi, Yichun Qian, Hongyu Zhang +2

    cs.CLcs.SEarXiv:2510.00446v12025
  48. Numbat: Building and Verifying a Self-Contained Machine-Learning Stack

    Thang Tran, Lan Dang

    cs.SEcs.LGarXiv:2609.10632v12026
  49. Finding Faster Configurations using FLASH

    Vivek Nair, Zhe Yu, Tim Menzies +2

    cs.SEarXiv:1801.02175v22018
  50. On the Relation between Code Quality and Machine Learning Performance: A Large-scale Empirical Study

    Marius Mignard, Steven Costiou, Anne Etien

    cs.SEcs.LGarXiv:2609.10610v12026
  51. Optimizing AI Inference Across the Deployment Stack

    Tejinder Singh, John Pflueger, Jeebak Mitra +3

    cs.SEcs.LGarXiv:2609.10550v12026
  52. What Makes a Good LLM Agent for Real-world Penetration Testing?

    Gelei Deng, Yi Liu, Yuekang Li +5

    cs.CRcs.SEarXiv:2602.17622v12026
  53. DeFiFlowBench: Benchmarking and Improving Safe Executability in Natural-Language DeFi Workflow Synthesis

    Abhinav Rajeev Kumar, Harshit Arora, Varun Singh +1

    cs.LGcs.SEarXiv:2609.11504v12026
  54. Estimating Inconsistency Response Surfaces under Uncertainty in Cyber-Physical System Development

    Johannes Mäkelburg, Tim Schwabe, Maribel Acosta

    cs.LGcs.SEeess.SYarXiv:2609.11331v12026
  55. Exploring LLM-based Agents for Root Cause Analysis

    Devjeet Roy, Xuchao Zhang, Rashi Bhave +4

    cs.SEcs.CLcs.LGarXiv:2403.04123v12024
  56. DR-LabStack: Design and Implementation of a Clinician-Facing Web System for Diabetic Retinopathy Prediction

    Yingfan Xu, Tieming Liu, Ye Liang

    cs.LGcs.SEarXiv:2609.10796v12026
  57. DIRE: A Neural Approach to Decompiled Identifier Naming

    Jeremy Lacomis, Pengcheng Yin, Edward J. Schwartz +4

    cs.SEarXiv:1909.09029v22019
  58. LLM-Based Test-Driven Interactive Code Generation: User Study and Empirical Evaluation

    Sarah Fakhoury, Aaditya Naik, Georgios Sakkas +2

    cs.SEarXiv:2404.10100v22024
  59. Predictors of Well-being and Productivity among Software Professionals during the COVID-19 Pandemic -- A Longitudinal Study

    Daniel Russo, Paul H. P. Hanel, Seraphina Altnickel +1

    cs.CYcs.SEarXiv:2007.12580v42020
  60. An Analysis of ISO 26262: Using Machine Learning Safely in Automotive Software

    Rick Salay, Rodrigo Queiroz, Krzysztof Czarnecki

    cs.AIcs.LGcs.SEarXiv:1709.02435v12017