The uncomfortable failure mode for a coding AI is not that it cannot write code. It is that it can produce convincing code while quietly ignoring an instruction that mattered to the project. A tool may complete the visible task and still use the wrong dependency, modify a protected file, skip a required check or disregard a repository rule. For teams evaluating coding agents, that distinction changes what reliability should mean.
The headline is a reliability problem, not a single verified anecdote
No accessible primary source located for this article established one specific incident that can responsibly be presented as the definitive story behind the title. The stronger evidence is broader: recent research is explicitly testing whether coding agents follow operational instructions across real repository environments rather than judging them only by whether the final code appears to work.
That matters because an impressive coding result can coexist with process failure. In production software work, instructions often encode architecture, security, compatibility and deployment assumptions that are not visible in a superficial demonstration.
Coding success and instruction compliance are different tests
A conventional evaluation may ask whether tests pass. An instruction-following evaluation asks additional questions: did the agent respect project rules, use the permitted tools and avoid actions it was told not to take?
Harness-IF, a recent research benchmark for coding agents, evaluates operational rules across several instruction surfaces used by deployed agents. Its premise is revealing: an agent can appear compliant simply because the requested behaviour matches what it would have done anyway. The benchmark therefore examines rules that push against default behaviour as a more demanding test of obedience.
Small wording changes can expose fragile reliability
The problem extends beyond coding. ACL research on instruction-following reliability examines whether models remain competent across closely related prompts that express analogous intentions with subtle differences. The researchers report substantial inconsistency across the models they evaluated.
For software teams, the practical lesson is not to search for one magical prompt. Important rules should be represented in ways the surrounding development process can verify. If compliance depends on exactly how a developer happened to phrase an instruction, the system is too fragile for consequential unattended work.
Repository instructions are part of the engineering system
Coding agents may receive direction from user prompts, project files, tool descriptions and other configuration. Those sources can overlap or conflict. A team therefore needs to know which rules are authoritative and whether its agent environment presents them consistently.
Keep critical instructions concise and testable where possible. Instead of relying only on a sentence saying not to alter a particular interface, automated checks can detect an incompatible change. Instead of merely requesting a formatting or dependency rule, the repository can enforce it through established development tooling.
Verification should target the failure you care about
A green unit-test suite cannot prove that every project requirement was respected. Verification should reflect the risk. That can include reviewing changed files, inspecting dependency modifications, running static checks, testing interfaces and requiring human approval before sensitive actions.
Recent work on coding-agent monitoring also treats instruction-following failure as a distinct category worth detecting. Apollo Research's coding-agent monitoring evaluation includes instruction-following failures alongside other undesirable trajectory behaviours, reinforcing the point that the path an agent takes can matter as much as its final output.
Use autonomy in proportion to reversibility
An agent preparing a draft change on an isolated branch is different from an agent with authority to modify production infrastructure. The harder an action is to inspect or reverse, the stronger the approval and verification around it should be.
This does not make coding AI unhelpful. It makes deployment an engineering decision. Teams can use agents aggressively where automated tests and review provide a dependable safety net, while retaining tighter human control over migrations, permissions, security-sensitive code and changes with a wide operational blast radius.
The best coding model is still one component
Model rankings are tempting because they turn a complicated buying decision into a score. Yet the most capable model cannot know every local rule merely from its general coding ability. Reliability emerges from the model, repository instructions, tool permissions, test environment, review process and the team's ability to observe what happened.
So the useful response to a coding AI that does not follow instructions is not simply to write a longer prompt or switch immediately to whichever model leads a benchmark. Identify which rule was missed, make important requirements verifiable, reduce unnecessary authority and test the complete workflow. In professional software delivery, “best” is not good enough if the system cannot demonstrate that the instructions which protect the project were actually followed.