Although the capabilities of large language models (LLMs) keep expanding, prompt design is still largely heuristic and ad‑hoc. This work focuses on $\textit{prompt\ minimization}$—reducing prompts to their smallest, most information‑dense form while preserving output fidelity. Shorter prompts cut computational cost and inference latency, especially when large contexts such as whole documents or codebases are unnecessarily included; moreover, overly long prompts can impair the model’s reasoning and accuracy. Theoretically, the existence of many different prompts that yield identical outputs indicates substantial redundancy in the input space, raising fundamental questions about which information is essential to trigger specific model behaviors. We introduce three variant frameworks to identify and evaluate minimal prompts and show experimentally that minimal prompts often produce outputs comparable to their longer counterparts. These findings point to new directions for efficient prompt engineering and deepen our understanding of input compression in LLMs.
Review: Prompt minimization reveals a promising path to improve efficiency without sacrificing performance, offering practical benefits for deployment and insights for future research.