Prompt advice is cheap. A new model arrives, a prompt guide lands, and the obvious move is to rewrite every instruction file before lunch.

We did the rewrite. Then we tested it.

Across 240 scored GPT-6 Astra runs, the new instructions produced no measurable correctness gain. They showed a small favourable speed estimate, but not one strong enough to call a real improvement. That result is less exciting than “10× better.” It is also more useful.

The 2×2 experiment

We crossed two versions of the global instructions with two versions of a merge-conflict skill. Twenty synthetic local tasks ran three times under each variant: Git integration, debugging, operations diagnosis, source-grounded research, data transformation, migration safety and formula generation.

VariantGlobal instructionsMerge skill
AOldOld
BRewrittenOld
COldRewritten
DRewrittenRewritten

Every task had an external objective grader: behaviour, files, JSON and CSV output, Git history, unresolved conflicts, protected tests, and unrelated staged work. No model graded another model’s prose.

Correctness hit the ceiling

Every variant passed every corrected objective check: 60 out of 60, four times over. That does not prove the prompts are equivalent. It means these 20 task types were too easy to expose a difference. A benchmark ceiling is still a result: it tells us not to claim quality gains from this evidence.

Speed moved a little

ComparisonTime estimate95% interval
Rewritten global instructions−3.8%−8.3% to +0.8%
Rewritten skill+0.3%−3.8% to +5.4%
Both rewritten−3.5%−9.9% to +3.4%

Negative means faster, and each interval crosses zero. So the honest conclusion is not “the rewrite improved speed.” It is that the point estimate favoured the new global instructions, but this experiment did not establish a reliable gain.

What we shipped

We kept the rewritten global instructions and retained the original merge skill. The skill rewrite was slower on the Git cases and used more output tokens; with correctness tied, that was enough to withhold promotion.

Prompt engineering is still engineering: frozen candidates, objective checks, paired comparisons, failure accounting, and a willingness to keep the old version. If every candidate scores perfectly, make the next suite harder.