No Loss, No Gain: Gated Refinement and Adaptive Compression for Prompt Optimization

Wenhang Shi (Renmin University of China) · Yiren Chen (Coupang) · Shuqing Bian (Tencent) · Xinyi Zhang (Renmin University of China) · Kai Tang (tencent ailab) · Pengfei Hu (Tencent ) · Zhe Zhao (University of Science and Technology of China) · WEI LU (Renmin University of China) · Xiaoyong Du (Renmin University of China)
adaptive compressionautomatic prompt optimizationcomputational overheadcore concepts distillationefficient prompt optimizationfeedback regulation gategated refinementinformation losslocal optimaoptimization stagnationoptimization trace restructuringperformance improvementsprompt engineeringupdate rejection gate

Prompt engineering is crucial for leveraging the full potential of large language models (LLMs). While automatic prompt optimization offers a scalable alternative to costly manual design, generating effective prompts remains challenging. Existing methods often struggle to stably generate improved prompts, leading to low efficiency, and overlook that prompt optimization easily gets trapped in local optima. Addressing this, we propose GRACE, a framework that integrates two synergistic strategies: Gated Refinement and Adaptive Compression, achieving Efficient prompt optimization. The gated refinement strategy introduces a feedback regulation gate and an update rejection gate, which refine update signals to produce stable and effective prompt improvements. When optimization stagnates, the adaptive compression strategy distills the prompt’s core concepts, restructuring the optimization trace and opening new paths. By strategically introducing information loss through refinement and compression, GRACE delivers substantial gains in performance and efficiency. In extensive experiments on 11 tasks across three practical domains, including BIG-Bench Hard (BBH), domain-specific, and general NLP tasks, GRACE achieves significant average relative performance improvements of 4.7\%, 4.4\% and 2.7\% over state-of-the-art methods, respectively. Further analysis shows that GRACE achieves these gains using only 25\% of the prompt generation budget required by prior methods, highlighting its high optimization efficiency and low computational overhead. Our code is available at https://github.com/Eric8932/GRACE.