Software Engineering
Papers filed under cs.SE on arXiv, each one already summarized by Paperlayer. Open any of them to read the summary beside the original PDF, with every point linked to the line, figure, or table it came from.
Search paper metadata (including unsummarized papers)
61 to 120 of 1,389
Identifying Implementation Bugs in Machine Learning based Image Classifiers using Metamorphic Testing
Anurag Dwarakanath, Manish Ahuja, Samarth Sikand +4
cs.SEcs.LGarXiv:1808.05353v12018Practical Program Repair via Bytecode Mutation
Ali Ghanbari, Lingming Zhang
cs.SEarXiv:1807.03512v12018Measuring Coding Challenge Competence With APPS
Dan Hendrycks, Steven Basart, Saurav Kadavath +8
cs.SEcs.CLcs.LGarXiv:2105.09938v32021ATIBA: Grounded Integrity and Quality Checking for Research Papers
Veli Karakaya, Semih Çağlar, Yusuf Yiğit Korkmaz +1
cs.SEarXiv:2609.04123v12026SpikingJelly: An open-source machine learning infrastructure platform for spike-based intelligence
Wei Fang, Yanqi Chen, Jianhao Ding +7
cs.NEcs.LGcs.SEarXiv:2310.16620v12023Virtual Testing of Automated Driving Systems through Credible Simulations
Riccardo Dona, Espedito Rusciano, Biagio Ciuffo
cs.ROcs.SEarXiv:2609.03760v12026AI Harness Engineering: A Runtime Substrate for Foundation-Model Software Agents
Hailin Zhong, Shengxin Zhu
cs.SEcs.AIarXiv:2605.13357v12026Agentic Refactoring: An Empirical Study of AI Coding Agents
Kosei Horikawa, Hao Li, Yutaro Kashiwa +3
cs.SEarXiv:2511.04824v12025RefDiff: Detecting Refactorings in Version Histories
Danilo Silva, Marco Tulio Valente
cs.SEarXiv:1704.01544v12017CodeScore: Evaluating Code Generation by Learning Code Execution
Yihong Dong, Jiazheng Ding, Xue Jiang +3
cs.SEarXiv:2301.09043v42023Can AI Remediate Backend Failures Safely? GuardedAct with Blast-Radius-Aware Sandboxing
Wanrong Cai, Tianyu Yu, Shaorui Pi +2
cs.DCcs.SEarXiv:2609.11264v12026SemVerBench: Benchmarking LLM Comprehension of Version-Constraint Resolution Semantics
Qibai Chen, Zeming Liu
cs.AIcs.SEarXiv:2609.11180v12026BenchShield: Formal Model-Backed Instrumentation for Reward Integrity in LLM-Agent Evaluation Infrastructure
Shenghan Zheng, Zonglin Di, Yimin Liu +19
cs.CRcs.AIcs.SEarXiv:2609.11028v12026Summaries:한국어Grounding Agent Memory: Environment-Probing Curation for Enterprise Agents
Susheel Suresh, Hazel Mak, Sahil Bhatnagar +2
cs.AIcs.SEarXiv:2609.11060v12026Decoupling Readiness from Release for Tail-Aware Scheduling of Agentic LLM Workflows
Bochao Feng, Jianjiang Li, Haojie Wang +4
cs.AIcs.SEarXiv:2609.10964v12026DeFiFusion: Combining Transaction Events with Smart Contracts to Detect Price Manipulation Attacks
Rui Cao, Shaojing Fan, Liming Fang +3
cs.CRcs.AIcs.SEarXiv:2609.11008v12026AspisAI: A Canonical, Machine-Interpretable Governance Framework for Automated Multi-Standard Compliance Monitoring
Tsafac Nkombong Regine Cyrille, Hasan Dag, Reiner Creutzburg +1
cs.CRcs.CYcs.SEarXiv:2609.10881v12026An analysis of the relationship of input metrics
Addison Crump
cs.SEcs.FLarXiv:2609.11824v12026Beyond Static Guarantees: Measuring the Static-Pass Dynamic-Fail Gap in Security-Sensitive and LLM-Generated Python Code
Jessica Pourleyli, Maitreyee Das Urmi, Glaucia Melo
cs.CRcs.AIcs.SEarXiv:2609.10762v12026A2ABreak: Systematic Security Analysis of the A2A Protocol
Alireza Lotfi, Mirza Masfiqur Rahman, Imtiaz Karim +1
cs.CRcs.SEarXiv:2609.10871v12026Towards a Deterministic Math Solver for Clinical Language Models
Felipe Ocampo Osorio, Sebastián Andrés Cajas Ordoñez, Maximin Lange +5
cs.AIcs.SEarXiv:2609.10728v12026Reproducibility in the Age of Agentic AI: Context Engineering at the Timescale of a Codebase
Lorena A. Barba
cs.SEcs.CYarXiv:2609.11728v12026PRISMA-LLM: An Empirical Reporting Framework for AI-Assisted Systematic Reviews
Miguel Zabaleta, Baihan Lin
cs.SEarXiv:2609.11559v12026Agent-Integrated Software: Interaction Contracts and Continuous Assurance
Shengcheng Yu, Chunrong Fang, Zhenyu Chen
cs.SEcs.AIarXiv:2609.11381v12026Deep Learning-based Bug Triage System
Sourabh Pal
cs.SEarXiv:2609.11420v12026ChurnBench: A Drift-Aware Benchmark Demonstrating That Refresh Scheduling, Not Cache Age, Governs Staleness in Agentic AI
Vivek Kumar Singh, Preeti Priyam
cs.SEarXiv:2609.11515v12026Ecdysis: Efficient and Effective Training of Runtime Harnesses for LLM Agents
Ruiqing Yue, Yu Cui, Zhuoyu Sun +11
cs.SEcs.AIarXiv:2609.11677v12026TripleBound: Triplet-Guided Heterogeneous Graph Learning for Microservice Decomposition
Mineth Weerasinghe, Himindu Kularathne, Methmini Madhushika +4
cs.SEarXiv:2609.11212v12026A Model-Centric DevOps Architecture for DEVS-Based Digital Twin Simulation Services
Arnis Lektauers, Gusts Linkevičs, Guntis Mosāns +2
cs.SEcs.CEcs.DCarXiv:2609.11122v12026CoSTAR: Data Synthesis-Driven Constraint-Aware COBOL Section Summarization for Legacy System Modernization
Hao Lin, He Jiang, Xiaochen Li +4
cs.SEarXiv:2609.11332v12026FST Pay: Deterministic Safety-Gated Architecture for Youth Digital Payments
Shaikh Mohammed Burhan, Syed Farhaan Quadri, Tabassum Nahid Sultana
cs.SEcs.CRarXiv:2609.11195v12026Exploring the Role of Security Experience and ChatGPT Usage Strategies on Secure Software Engineering Education
Alessio Ferrari, Minh An Nguyen, Kushal Ramkumar +1
cs.SEarXiv:2609.11303v12026SaltBench: A Referee-Gated Protocol for Measuring Method Effects in Machine-Checked Software Work
Jason Hickey
cs.SEcs.LOarXiv:2609.11076v12026RCL: A Retrieval-Confidence Layer for Detecting Insufficient Context in Enterprise Retrieval-Augmented Code Generation
Chandra Mohan Ravuri
cs.SEarXiv:2609.11023v12026Engineering Reliable Commit Gates for Agentic AI: Cost-Aware Verification Portfolios under Common-Mode Data Failures
Zihao Zheng, Baichuan Li, Junyi Yao +1
cs.SEarXiv:2609.10969v12026What a Random Draw from the MCP Registry Contains, and What Tool-Use Benchmarks Contain Instead
Haseeb Mohammed Afsar
cs.SEcs.AIarXiv:2609.10962v12026LLMVul: A Vulnerability-Labeled Dataset of LLM-Generated C/C++ Functions from Real Production Repositories
Mohammad Farhad, Shuvalaxmi Dass
cs.SEarXiv:2609.10945v12026When Passing Tests Hides Vulnerabilities: An Empirical Study of Silent Failures in Agentic Systems
Wenji Bai, Muhammad Waseem, Zeeshan Rasheed +2
cs.SEcs.CRarXiv:2609.10548v12026Generative AI for trustworthy systems - Towards a health check model
Jan Bosch, Rick Kazman, Henry Muccini +1
cs.SEarXiv:2609.10595v12026ReqEvolve: User-Oriented Software Self-Evolution through Automatic Requirement Interpretation
Md Asif Iqbal Fahim, Alessio Ferrari
cs.SEarXiv:2609.10590v12026Governed Human-AI Prioritization Under Uncertainty: Adaptive Estimation and Dependency-Constrained Portfolio Selection
Azzeddine Ihsine, Sara Ihsine
cs.SEarXiv:2609.10648v12026AI Safety: Not Optional, Not Later
Qinghua Lu, Yoshua Bengio
cs.SEarXiv:2609.10630v12026Agent READMEs: An Empirical Study of Context Files for Agentic Coding
Worawalan Chatlatanagulchai, Hao Li, Yutaro Kashiwa +8
cs.SEarXiv:2511.12884v22025VeriGuard: Enhancing LLM Agent Safety via Verified Code Generation
Lesly Miculicich, Mihir Parmar, Hamid Palangi +4
cs.SEcs.AIcs.CRarXiv:2510.05156v12025SWE-QA: Can Language Models Answer Repository-level Code Questions?
Weihan Peng, Yuling Shi, Yuhang Wang +3
cs.CLcs.PLcs.SEarXiv:2509.14635v22025SWE-EVO: Benchmarking Coding Agents in Long-Horizon Software Evolution Scenarios
Tue Le, Minh V. T. Thai, Dung Nguyen Manh +2
cs.SEcs.AIcs.MAarXiv:2512.18470v62025LongCodeZip: Compress Long Context for Code Language Models
Yuling Shi, Yichun Qian, Hongyu Zhang +2
cs.CLcs.SEarXiv:2510.00446v12025Numbat: Building and Verifying a Self-Contained Machine-Learning Stack
Thang Tran, Lan Dang
cs.SEcs.LGarXiv:2609.10632v12026Finding Faster Configurations using FLASH
Vivek Nair, Zhe Yu, Tim Menzies +2
cs.SEarXiv:1801.02175v22018On the Relation between Code Quality and Machine Learning Performance: A Large-scale Empirical Study
Marius Mignard, Steven Costiou, Anne Etien
cs.SEcs.LGarXiv:2609.10610v12026Optimizing AI Inference Across the Deployment Stack
Tejinder Singh, John Pflueger, Jeebak Mitra +3
cs.SEcs.LGarXiv:2609.10550v12026What Makes a Good LLM Agent for Real-world Penetration Testing?
Gelei Deng, Yi Liu, Yuekang Li +5
cs.CRcs.SEarXiv:2602.17622v12026DeFiFlowBench: Benchmarking and Improving Safe Executability in Natural-Language DeFi Workflow Synthesis
Abhinav Rajeev Kumar, Harshit Arora, Varun Singh +1
cs.LGcs.SEarXiv:2609.11504v12026Estimating Inconsistency Response Surfaces under Uncertainty in Cyber-Physical System Development
Johannes Mäkelburg, Tim Schwabe, Maribel Acosta
cs.LGcs.SEeess.SYarXiv:2609.11331v12026Exploring LLM-based Agents for Root Cause Analysis
Devjeet Roy, Xuchao Zhang, Rashi Bhave +4
cs.SEcs.CLcs.LGarXiv:2403.04123v12024DR-LabStack: Design and Implementation of a Clinician-Facing Web System for Diabetic Retinopathy Prediction
Yingfan Xu, Tieming Liu, Ye Liang
cs.LGcs.SEarXiv:2609.10796v12026DIRE: A Neural Approach to Decompiled Identifier Naming
Jeremy Lacomis, Pengcheng Yin, Edward J. Schwartz +4
cs.SEarXiv:1909.09029v22019LLM-Based Test-Driven Interactive Code Generation: User Study and Empirical Evaluation
Sarah Fakhoury, Aaditya Naik, Georgios Sakkas +2
cs.SEarXiv:2404.10100v22024Predictors of Well-being and Productivity among Software Professionals during the COVID-19 Pandemic -- A Longitudinal Study
Daniel Russo, Paul H. P. Hanel, Seraphina Altnickel +1
cs.CYcs.SEarXiv:2007.12580v42020An Analysis of ISO 26262: Using Machine Learning Safely in Automotive Software
Rick Salay, Rodrigo Queiroz, Krzysztof Czarnecki
cs.AIcs.LGcs.SEarXiv:1709.02435v12017