EffiPair: Improving the Efficiency of LLM-generated Code with Differential Execution Feedback

  • 2026-09-26 17:37:54
  • Samira Hajizadeh, Suman Jana
  • 0

Abstract

Large language models (LLMs) can generate functionally correct programs that differ substantially in execution efficiency. Existing inference-time optimization methods typically refine each candidate using pointwise runtime or profiling feedback, which identifies how costly an implementation is or where the cost arises, but offers limited guidance on how the computation should change. We introduce DIFFERENTIAL EXECUTION FEEDBACK (DEF), which compares implementations that are nearby in structural program space but separated in performance, turning their execution and implementation differences into directional optimization evidence. We demonstrate DEF in EffiPair, a training-free, test-time framework that pairs structurally similar programs with different efficiencies, distills their relative execution behavior into compact feedback, and iteratively refines a candidate pool. Across EvalPerf, Mercury, and ENAMEL, using GPT-4o mini, DeepSeek-V4.1 Flash, and GPT-5 mini, EffiPair achieves the highest value on each benchmark's official efficiency metric in all nine model-benchmark settings under matched evaluation conditions. Moreover, two contrastive refinement rounds improve efficiency over the selected initial draft in every setting while preserving or improving Pass@1. These results demonstrate the effectiveness of relational execution feedback as a lightweight signal for test-time code optimization.

 

Quick Read (beta)

loading the full paper ...