Prompt advice is cheap. A new model arrives, a prompt guide lands, and the obvious move is to rewrite every instruction file before lunch.
We did the rewrite. Then we tested it.
Across 240 scored GPT-6 Astra runs, the new instructions produced no measurable correctness gain. They showed a small favourable speed estimate, but not one strong enough to call a real improvement. That result is less exciting than “10× better.” It is also more useful.
The 2×2 experiment
We crossed two versions of the global instructions with two versions of a merge-conflict skill. Twenty synthetic local tasks ran three times under each variant: Git integration, debugging, operations diagnosis, source-grounded research, data transformation, migration safety and formula generation.
| Variant | Global instructions | Merge skill |
|---|---|---|
| A | Old | Old |
| B | Rewritten | Old |
| C | Old | Rewritten |
| D | Rewritten | Rewritten |
Every task had an external objective grader: behaviour, files, JSON and CSV output, Git history, unresolved conflicts, protected tests, and unrelated staged work. No model graded another model’s prose.
Correctness hit the ceiling
Every variant passed every corrected objective check: 60 out of 60, four times over. That does not prove the prompts are equivalent. It means these 20 task types were too easy to expose a difference. A benchmark ceiling is still a result: it tells us not to claim quality gains from this evidence.
Speed moved a little
| Comparison | Time estimate | 95% interval |
|---|---|---|
| Rewritten global instructions | −3.8% | −8.3% to +0.8% |
| Rewritten skill | +0.3% | −3.8% to +5.4% |
| Both rewritten | −3.5% | −9.9% to +3.4% |
Negative means faster, and each interval crosses zero. So the honest conclusion is not “the rewrite improved speed.” It is that the point estimate favoured the new global instructions, but this experiment did not establish a reliable gain.
What we shipped
We kept the rewritten global instructions and retained the original merge skill. The skill rewrite was slower on the Git cases and used more output tokens; with correctness tied, that was enough to withhold promotion.
Prompt engineering is still engineering: frozen candidates, objective checks, paired comparisons, failure accounting, and a willingness to keep the old version. If every candidate scores perfectly, make the next suite harder.