Publications

Papers and preprints on LLM post-training for software engineering, AI systems, and empirical software engineering. The full, always-current list is on Google Scholar.

2026 12

  1. TOSEM 2026Journal

    Towards AI-Native Software Engineering (SE 3.0): A Vision and a Challenge Roadmap

    Ahmed E. Hassan, Gustavo A. Oliva, Dayi Lin, Boyuan Chen, Zhen Ming (Jack) Jiang

    A vision of AI-native, intent-first software engineering with AI teammates, and the technology roadmap to get there.

  2. EMNLP 2026Industry Track, accepted

    LLM Post-Training as Brownfield Maintenance: An Industrial Perspective on Dataware Engineering

    Gopi Krishnan Rajbahadur, Amir M. Ebrahimi, Boyuan Chen, Ahmed E. Hassan

    Treats industrial post-training as maintaining a deployed checkpoint through budgeted data-mixture patches, and distills what makes that hard.

  3. ASE 2026Industry Showcase

    DCAS: Decoupling CLI Agent Scaffolding to Internalize Planning across Scaffolds

    Kishanthan Thangarajah, Boyuan Chen, Ahmed E. Hassan

    An interception layer that decouples CLI coding-agent scaffolds from models; planning learned from Claude Code trajectories transfers to OpenCode and mini-swe-agent.

  4. arXiv 2026Preprint

    SWE-NFI: Studying and Benchmarking Coding Agents for Non-Functional Improvements

    Pengyu Xue, He Yang Yuan, Xin Wang, Junkai Chen, Haonan Zhang, Boyuan Chen, Zishuo Ding, Zhenhao Li, Weiyi Shang

    A benchmark of 188 tasks from merged pull requests for evaluating coding agents on behavior-preserving, non-functional improvements.

  5. arXiv 2026Preprint

    SpecFirst: Behavioral Specification Elicitation as a First-Class Step in Agent-Based Program Synthesis from Scratch

    Yihao Chen, Shi Chang, Feng Lin, Khaled Chawa, Boyuan Chen, Shaowei Wang, Ahmed E. Hassan

    Makes behavioral-specification elicitation an explicit first step before code synthesis when agents build programs from scratch.

  6. arXiv 2026Preprint

    MindForge: Teaching Small Language Models Whole-Life-Cycle Software Engineering via Source-Free Program Synthesis

    Yihao Chen, Shi Chang, Khaled Chawa, Feng Lin, Boyuan Chen, Shaowei Wang, Ahmed E. Hassan

    Turns open-source CLI programs into source-free environments for whole-life-cycle program synthesis; fine-tuning Qwen3.6-27B lifts the ProgramBench pass rate from 37.98% to 49.51%.

  7. AACL-IJCNLP 2026Main conference

    Beyond Correctness: Enhancing Architectural Reasoning in Code LLMs via Scalable Labeling with Agentic Judgment

    Kirill Vasilevski, Ximing Dong, Benjamin Rombaut, Milad Soltany, Ruochen Deng, Jiahuei Lin, Arthur Leung, Dayi Lin, Boyuan Chen, Shaowei Wang, Ahmed E. Hassan

    Agentic LLM judges curate data for architecture-aware fine-tuning; Qwen3 models trained on 3,360 instances reach up to 27.2% on SWE-bench Verified.

  8. arXiv 2026Preprint

    SynConfRoute: Syntax-Aware Routing for Efficient Code Completion with Small CodeLLMs

    Kishanthan Thangarajah, Boyuan Chen, Ahmed E. Hassan

    Training-free routing that combines token confidence with syntax checks to decide when a small local code model is good enough, cutting accelerator use by 58%.

  9. ICSE 2026Technical briefing

    Software Engineering for Foundation Models (SE4FM)

    Boyuan Chen, Dayi Lin, Arthur Leung, Zhilong Chen, Gopi Krishnan Rajbahadur, Gustavo A. Oliva, Yihao Chen, Xiaoshuang Liu, Chun Yong Chong, Ahmed E. Hassan

    Companion paper for the ICSE 2026 technical briefing on software engineering for foundation models.

  10. TOSEM 2026Journal

    Software Performance Engineering for Foundation Model-Powered Software

    Haoxiang Zhang, Shi Chang, Arthur Leung, Kishanthan Thangarajah, Boyuan Chen, Hanan Lutfiyya, Ahmed E. Hassan

    Four performance-engineering challenges for software built on foundation models, from cognitive architecture design to deployment, with research directions.

  11. ASE 2026Industry Showcase

    When Elo Lies: Hidden Biases in Codeforces-Based Evaluation of Large Language Models in Practice

    Shenyu Zheng, Ximing Dong, Xiaoshuang Liu, Gustavo A. Oliva, Chun Yong Chong, Dayi Lin, Boyuan Chen, Shaowei Wang, Ahmed E. Hassan

    Shows how submission order, contest selection and run-to-run variance bias Codeforces Elo ratings of LLMs; submission order alone shifts scores by 394 points.

  12. EMNLP 2026Main conference

    Beyond Tokens: Semantic-Aware Speculative Decoding for Efficient Inference by Probing Internal States

    Ximing Dong, Shaowei Wang, Dayi Lin, Boyuan Chen, Ahmed E. Hassan

    SemanticSpec verifies whole semantic steps instead of single tokens by probing hidden states, for up to 2.7× faster decoding on DeepSeek-R1-32B.

2025 7

  1. ASE 2025Research track

    SPICE: An Automated SWE-Bench Labeling Pipeline for Issue Clarity, Test Coverage, and Effort Estimation

    Gustavo A. Oliva, Gopi Krishnan Rajbahadur, Aaditya Bhatia, Haoxiang Zhang, Yihao Chen, Zhilong Chen, Arthur Leung, Dayi Lin, Boyuan Chen, Ahmed E. Hassan

    Labels SWE-bench-style instances for issue clarity, test coverage and effort, cutting the cost for 1,000 instances from about $100,000 to $5.10.

  2. ASE 2025Industry Showcase

    Context-Aware CodeLLM Eviction for AI-assisted Coding

    Kishanthan Thangarajah, Boyuan Chen, Shi Chang, Ahmed E. Hassan

    Context-aware model eviction for self-hosted, multi-model code-LLM serving under limited accelerator memory.

  3. arXiv 2025Preprint

    SWE-Effi: Re-Evaluating Software AI Agent System Effectiveness Under Resource Constraints

    Zhiyu Fan, Kirill Vasilevski, Dayi Lin, Boyuan Chen, Yihao Chen, Zhiqing Zhong, Jie M. Zhang, Pinjia He, Ahmed E. Hassan

    Effectiveness metrics that weigh SWE-agent resolve rates against the tokens and time they consume, used to re-rank AI systems on SWE-bench.

  4. arXiv 2025Technical report

    RepoForge: Training a SOTA Fast-thinking SWE Agent with an End-to-End Data Curation Pipeline Synergizing SFT and RL at Scale

    Zhilong Chen, Chengzong Zhao, Boyuan Chen, Dayi Lin, Yihao Chen, Arthur Leung, Gopi Krishnan Rajbahadur, Gustavo A. Oliva, Haoxiang Zhang, Aaditya Bhatia, Chong Chun Yong, Ahmed E. Hassan

    Generates 7,304 executable environments from GitHub commits and trains with SFT and RL; RepoForge-8B-Agent reaches 17.4% on SWE-bench Verified.

  5. KDD 2025Tutorial

    The Hitchhikers Guide to Production-ready Trustworthy Foundation Model Powered Software (FMware)

    Kirill Vasilevski, Gopi Krishnan Rajbahadur, Gustavo A. Oliva, Benjamin Rombaut, Keheliya Gallaba, Filipe R. Cogo, Jiahuei (Justina) Lin, Dayi Lin, Haoxiang Zhang, Boyuan Chen, Kishanthan Thangarajah, Ahmed E. Hassan, Zhen Ming (Jack) Jiang

    Tutorial on the challenges and technology roadmap for production-ready, trustworthy software built on foundation models.

  6. arXiv 2025Preprint

    Adaptive Request Scheduling for CodeLLM Serving with SLA Guarantees

    Shi Chang, Boyuan Chen, Kishanthan Thangarajah, Hanan Lutfiyya, Ahmed E. Hassan

    SLA-aware dynamic batching for self-hosted code-LLM serving, with up to 26% higher goodput and 45% lower latency variability.

  7. arXiv 2025Preprint

    SLA-Awareness for AI-assisted coding

    Kishanthan Thangarajah, Arthur Leung, Boyuan Chen, Ahmed E. Hassan

    A runtime that serves mixed coding tasks with different latency targets while keeping cluster utilization high.

2024 2

  1. FSE 2024Industry track

    Rethinking Software Engineering in the Era of Foundation Models: A Curated Catalogue of Challenges in the Development of Trustworthy FMware

    Ahmed E. Hassan, Dayi Lin, Gopi Krishnan Rajbahadur, Keheliya Gallaba, Filipe R. Cogo, Boyuan Chen, Haoxiang Zhang, Kishanthan Thangarajah, Gustavo A. Oliva, Jiahuei (Justina) Lin, Wali Mohammad Abdullah, Zhen Ming (Jack) Jiang

    Ten challenges that make enterprise FMware development risky, and FMArts, Huawei Canada's platform for engineering trustworthy FMware.

  2. arXiv 2024Preprint

    Rethinking Software Engineering in the Foundation Model Era: From Task-Driven AI Copilots to Goal-Driven AI Pair Programmers

    Ahmed E. Hassan, Gustavo A. Oliva, Dayi Lin, Boyuan Chen, Zhen Ming (Jack) Jiang

    Argues for goal-driven AI pair programmers instead of task-driven copilots.

2023 1

  1. ICSE-SEIP 2023Software Engineering in Practice

    An Empirical Comparison on the Results of Different Clone Detection Setups for C-based Projects

    Yan Zhou, Jinfu Chen, Yong Shi, Boyuan Chen, Zhen Ming (Jack) Jiang

    Compares how different clone-detection setups change the results on C-based projects.

2022 4

  1. TSE 2022Journal

    An Experience Report on Producing Verifiable Builds for Large-Scale Commercial Systems

    Yong Shi, Mingzhi Wen, Filipe R. Cogo, Boyuan Chen, Zhen Ming (Jack) Jiang

    A process and toolkit for producing verifiable builds, evaluated on three large commercial systems at Huawei.

  2. ICSE 2022Technical track

    Towards Training Reproducible Deep Learning Models

    Boyuan Chen, Mingzhi Wen, Yong Shi, Dayi Lin, Gopi Krishnan Rajbahadur, Zhen Ming (Jack) Jiang

    Reproducibility criteria plus record-and-replay and profile-and-patch techniques that make deep-learning training reproducible despite software and hardware non-determinism.

  3. ICSE-SEIP 2022Software Engineering in Practice

    Towards Build Verifiability for Java-based Systems

    Jiawen Xiong, Yong Shi, Boyuan Chen, Filipe R. Cogo, Zhen Ming (Jack) Jiang

    A systematic approach to verifiable builds for Java, applied to 46 Reproducible Central projects and 13 open-source projects used in Huawei products.

  4. TOSEM 2022Journal

    Towards a Consistent Interpretation of AIOps Models

    Yingzhe Lyu, Gopi Krishnan Rajbahadur, Dayi Lin, Boyuan Chen, Zhen Ming (Jack) Jiang

    Studies whether interpretations of AIOps models stay consistent across models, data and time.

2021 2

  1. arXiv 2021Preprint

    Can I use this publicly available dataset to build commercial AI software? – A Case Study on Publicly Available Image Datasets

    Gopi Krishnan Rajbahadur, Erika Tuck, Li Zi, Dayi Lin, Boyuan Chen, Zhen Ming (Jack) Jiang, Daniel M. German

    Assesses potential license violations when commercial AI software is built from publicly available datasets.

  2. CSUR 2021Journal

    A Survey of Software Log Instrumentation

    Boyuan Chen, Zhen Ming (Jack) Jiang

    A survey of research on software log instrumentation: logging approaches, logging utilities and logging-code quality.

2020 2

  1. Ph.D. thesisYork University

    Improving the Logging Practices in DevOps

    Boyuan Chen

    Automated approaches to improve logging practices on both the development and operations sides of DevOps.

  2. ICSE 2020Technical track

    Studying the Use of Java Logging Utilities in the Wild

    Boyuan Chen, Zhen Ming (Jack) Jiang

    A study of 3,856 logging utilities across 11,194 Java projects on GitHub: why projects use multiple and custom loggers.

2019 3

  1. ASE 2019Industry experience report

    An Industrial Experience Report on Performance-Aware Refactoring on a Database-Centric Web Application

    Boyuan Chen, Zhen Ming (Jack) Jiang, Paul Matos, Michael Lacaria

    Seventeen performance anti-patterns and refactorings applied to an industrial, database-centric web application.

  2. EMSE 2019Journal

    Extracting and studying the Logging-Code-Issue-Introducing changes in Java-based large-scale open source software systems

    Boyuan Chen, Zhen Ming (Jack) Jiang

    Extracts and studies logging-code-issue-introducing changes in six large Java systems; existing detectors catch only 3% of them.

  3. ICSE 2019Doctoral symposium

    Improving the Software Logging Practices in DevOps

    Boyuan Chen

    Outlines doctoral research on improving logging practices across development and operations.

2018 1

  1. ASE 2018Research track

    An Automated Approach to Estimating Code Coverage Measures via Execution Logs

    Boyuan Chen, Jian Song, Peng Xu, Xing Hu, Zhen Ming (Jack) Jiang

    LogCoCo estimates code-coverage measures from execution logs, evaluated on industrial projects at Baidu and on open-source projects.

2017 3

  1. M.A.Sc. thesisYork University

    Characterizing and Improving Logging Practices in Java-based Open Source Software Projects - A Large-scale Case Study in Apache Software Foundation

    Boyuan Chen

    Characterizes and improves logging practices in Java-based Apache projects; the basis of the EMSE 2017 paper.

  2. ICSE 2017Technical track

    Characterizing and Detecting Anti-Patterns in the Logging Code

    Boyuan Chen, Zhen Ming (Jack) Jiang

    Six logging-code anti-patterns derived from 352 changes in ActiveMQ, Hadoop and Maven, detected with the LCAnalyzer tool.

  3. EMSE 2017Journal

    Characterizing logging practices in Java-based open source software projects - a replication study in Apache Software Foundation

    Boyuan Chen, Zhen Ming (Jack) Jiang

    Replicates Yuan et al.'s logging-practice study on 21 Java projects from the Apache Software Foundation.

Patents 6

  1. US 2026/0133843 A1Application, published 2026

    Systems and Methods for Service Level Agreements for Foundation Model Applications (SLA-aware scheduler)

  2. US 2026/0133848 A1Application, published 2026

    Systems and Methods for Service Level Agreements for Foundation Model Applications (SLA-aware resource provisioner)

  3. US 12,541,446 B2Granted 2026

    Method and Apparatus to Trace and Visualize Data Movement

  4. US 12,468,637 B2Granted 2025

    Method and Apparatus for Providing Artificial Intelligence Model Swapping to Support Foundation Models

  5. WO 2023/060525 A1PCT application, 2023

    Methods and Systems for Generating Verifiable Software Releases

  6. WO 2023/028996 A1PCT application, 2023

    Methods and Devices for Ensuring the Reproducibility of Software Systems