Teaching AI coding agents to be pragmatic
AI coding agents can write a lot of code very quickly. That makes engineering judgment more important, not less.
I build most of my client work hands-on with agents, and I kept seeing the same habits. Left to themselves, agents will:
- rewrite working code instead of making a small, careful change
- add a new dependency or another layer of abstraction too eagerly
- guess at files, schemas, versions and APIs instead of reading them
- patch the symptom without understanding the cause
- widen the scope of a task in the name of “cleanup”
- declare success because the code compiles, without checking the product
None of these is new. They’re the mistakes The Pragmatic Programmer by Andy Hunt and Dave Thomas warned human developers about in 1999. It’s one of my favourite books, and I wanted its principles in every project I work on. So I turned them into a skill: operational guidance an agent follows while it builds.
Behaviour, not philosophy
The goal isn’t an agent that quotes software philosophy at you. It’s an agent that behaves like a pragmatic engineer. The short version of the rules:
- Inspect before assuming.
- Preserve before replacing.
- Make the smallest correct change.
- Verify; don’t hallucinate.
- Fix root causes.
- Don’t program by coincidence.
- Keep knowledge DRY.
- Keep components orthogonal.
- Prefer reversible decisions.
- Use the existing stack first.
- Match verification to the task.
- Don’t declare success from plausibility.
And one question that matters more for agents than for people:
Did I change more than necessary?
A few design choices keep the skill practical rather than dogmatic:
- It works silently. Routine compliance shouldn’t turn every reply into a lecture. The agent only raises what affects risk, architecture, scope or a decision you need to make.
- Testing is proportional. A behavioural bug usually deserves a failing regression test first. A CSS tweak doesn’t need a new test framework.
- Dependencies depend on context. For security-sensitive or standards-heavy problems, a mature library beats a few hundred lines the agent wrote itself. For a genuinely simple problem, in-house code can be the better choice.
- Abstractions have to earn their keep. No wrapping every library just in case it gets replaced one day.
Light enough for small models
The skill comes in two levels. CORE.md is a compact set of rules that’s cheap enough to keep in context all the time, including for smaller local models. Focused playbooks for debugging, refactoring, architecture, testing, integrations, security and web development load only when a task needs them. A copy change doesn’t need to carry the full engineering manifesto.
It’s model-agnostic. It works with cloud agents such as Claude Code and Codex, and with open models such as Qwen, Gemma, DeepSeek and Llama. What changes between them is how the instructions are delivered, not the standard.
Does it make a difference?
I wanted evidence, not a feeling, so I built two small test codebases and gave an agent the same tasks twice: once with its default behaviour and once with the skill. Everything is in the repository, including the challenges, the code and the results, so you can rerun it.
A backend event pipeline, with three tasks:
- A deduplication bug. Without the skill, the agent added a second cache that was never cleared, which is a memory leak, and touched 25 lines. With the skill, it wrote a failing test first, traced the bug to cleanup code that used the wall clock instead of the event timeline, and fixed it in 8 lines.
- Webhook signature checks. Without the skill, the agent installed a new package (
crypto-js) and compared signatures with a plain===, which is open to timing attacks. With the skill, it used Node’s built-incryptowith a constant-time comparison and added no dependencies. - New export formats. Without the skill, the agent changed a function signature in a way that broke existing callers, then edited an existing test to make it pass. With the skill, it added a default parameter, so every existing caller and test stayed untouched.
A web app (a drag-and-drop task board), with three tasks:
- Mobile layout. Without the skill, the agent hard-coded six colours and ignored the design tokens. With the skill, it reused the existing tokens and left the desktop layout as it was.
- Search and filters. Without the skill, the agent rebuilt every card on each keystroke. That broke drag-and-drop after the first search and stripped out accessibility attributes. With the skill, it hid and showed the existing cards instead, so drag-and-drop and input focus kept working.
- A task detail popup. Without the skill, the agent built a
<div>that keyboard users couldn’t reach or close properly. With the skill, it used the native<dialog>element, with keyboard access, Escape to close, and focus returned to the card you came from.
The pattern across all six was the same: the agent with the skill read before it wrote, changed less, and kept what already worked.
These are small codebases and a handful of tasks, so I’d treat the results as an illustration rather than proof. But the failures without the skill are exactly the ones I’ve had to catch in real projects.
Try it
Copy the skill into your project, or into your agent’s global skills folder, then add a short rule to your project instructions (AGENTS.md, CLAUDE.md or similar):
## Engineering philosophy
For non-trivial software changes, default to the pragmatic-programmer skill.
Prefer evidence over assumptions, surgical changes over rewrites, existing
conventions, and verification before declaring success. Apply silently.
The how-to guide has ready-made prompts for testing and fixing an existing codebase, plus the full benchmark write-up.
And if you find these ideas useful, read the book. The skill is an adaptation of its principles, not a replacement for it.