You Don't Need Prompt Engineering Anymore: The Prompting Inversion

2510.22251v1 cs.CL, cs.AI, cs.LG 2025-10-29

Авторы:

Imran Khan

Abstract

Prompt engineering, particularly Chain-of-Thought (CoT) prompting, significantly enhances LLM reasoning capabilities. We introduce "Sculpting," a constrained, rule-based prompting method designed to improve upon standard CoT by reducing errors from semantic ambiguity and flawed common sense. We evaluate three prompting strategies (Zero Shot, standard CoT, and Sculpting) across three OpenAI model generations (gpt-4o-mini, gpt-4o, gpt-5) using the GSM8K mathematical reasoning benchmark (1,317 problems). Our findings reveal a "Prompting Inversion": Sculpting provides advantages on gpt-4o (97% vs. 93% for standard CoT), but becomes detrimental on gpt-5 (94.00% vs. 96.36% for CoT on full benchmark). We trace this to a "Guardrail-to-Handcuff" transition where constraints preventing common-sense errors in mid-tier models induce hyper-literalism in advanced models. Our detailed error analysis demonstrates that optimal prompting strategies must co-evolve with model capabilities, suggesting simpler prompts for more capable models.

Ссылки и действия

Читать на arXiv Скачать PDF

Дополнительные ресурсы:

You Don't Need Prompt Engineering Anymore: The Prompting Inversion

Авторы:

Abstract

Ссылки и действия

Связанные статьи

Becoming Experienced Judges: Selective Test-Time Learning for Evaluators

LYNX: Learning Dynamic Exits for Confidence-Controlled Reasoning

To Think or Not to Think: The Hidden Cost of Meta-Training with Excessive CoT Ex...

Arbitrage: Efficient Reasoning via Advantage-Aware Speculation

Structured Document Translation via Format Reinforcement Learning

Навигация