What I changed in my AI workflowsafter trying Claude Opus 5.5
What changed when I audited my coding-agent skills with Opus 5.5: fewer procedural rules, clearer ownership, and stronger verification.
My prompting habits were starting to lag behind Claude Opus 5.5. After trying it on my own work, I had it audit the reusable instructions I give my coding agents, and we removed old procedural rules while making the checks on finished work more demanding.
How little direction it needed
I was already test-driving Opus 5.5 before I read the positive reviews. In my own work, I'd put it on par with GPT-6 Astra, and sometimes ahead. That's my experience using them, rather than a controlled comparison between the models.
Creating a video was the example that made the difference clear to me. Previously, I would give Claude Code the tools and walk it through how to put everything together. With Opus 5.5, I've been able to get it done in one shot from a simple prompt.
In my Claude Code setup, it figures out the steps and acts on them without coming back for approval on every routine decision. That changes how much work I can start with a few sentences. It also changes what I need to spend those sentences on.
Anthropic announced Opus 5.5 on September 22. By the time I was reading about it, I already had a practical question from using it: how much of my existing setup was still helping?
I'd been telling it to follow the old process
This weekend I had it audit agent-team-execute, my reusable instructions for getting several coding agents to work through a plan. We also looked at the surrounding skills that decide how to approach a task and which execution method to use.
A skill, in this context, is a saved set of instructions the agent can use again. Mine described things such as how to divide work, when to delegate, and what to check before calling a task finished. The audit gave me a chance to separate the rules that still served those purposes from the ones carrying older assumptions.
One of the older orchestrator instructions told the lead agent never to suggest doing the work itself. The team-execution skill prescribed engineers, a tester and a reviewer as separate roles. That arrangement was specified before considering whether the particular task benefited from it.
There were also pinned model versions and mandatory research-tool sequences. The planning instructions even carried a fixed year-based search rule. Those details could remain in the skill long after the reason for putting them there had passed.
The cleanup removed the forced research sequence and old model assignments. The revised guidance kept the requirement to check version-sensitive APIs against the installed version. That was the useful part: establish that the implementation fits the environment it actually runs in.
I'd been telling a more capable model to keep following the old process, and some of that process no longer described the work I wanted it to do.
A task doesn't need a team just because I can launch one
We changed long-running plans to execute in the current session by default. Previously, that workflow handed the work to a separate command-line runner. The current agent now works through the plan, with delegation where the tasks can safely run in parallel.
The separate runner still has a purpose for overnight work or work that needs to outlive the current session. We kept it available. The change was making it a deliberate choice instead of the starting point for every long plan.
Parallel work also needed a more concrete test than whether a task sounded large. If the pieces own different files and their dependencies are ready, they can be candidates for delegation. If they overlap, the current agent can work through them in order.
That distinction matters to how I write the request. I can describe the result and the constraints without deciding in advance that the job needs a miniature engineering department. A task doesn't need a team just because I have a skill that can launch one.
The definition of finished became more demanding
The audit didn't lead me to throw away agent-team-execute. We rewrote it around the boundaries that still matter when several agents work on the same application.
If two agents touch the same files, being smarter doesn't resolve who owns them. If they build different parts, they still need to agree on how those parts connect. The revised workflow defines shared interfaces before parallel implementation starts, so each agent has the same agreement to build against.
The checks happen along the way. Each piece gets an independent verification step, and the combined code gets tested as the pieces come together. A final review looks at the whole change against the requirements. The current agent also runs the task's verification command rather than accepting a subagent's report that the tests passed.
An agent saying "done" isn't the evidence that closes the task. That remained true even as I became more comfortable giving the model fewer directions about how to get there.
I had Opus update the skills, and the checks we ran passed. I cancelled the full old-versus-new comparison, so I don't have a measured speedup to claim from the rewrite. Passing those checks supports the changes we examined; it doesn't establish that the new workflow will outperform the old one on every real project.
What I do have is a different starting point for my next prompt: the outcome I want, the constraints that matter, and how we'll know it worked. I'm leaving more of the route to the model, while keeping the evidence of completion explicit.
What's an instruction you still give AI that it may no longer need?
How I Automated Myself Out of a Job (In a Good Way)
I built an AI webmaster so my client could stop waiting on me. Here's how it works, what broke, and what I'd do again.
ReadWhy So Much AI Writing Sounds the Same, and What I Am Doing About Mine
A model can copy how I sound. It cannot decide what I think, which is why I'm spending more time writing, not less.
ReadWhy I Ordered a $10,000 Mac Studio for Local AI
A practical look at the rate limits, data constraints, and hardware tradeoffs behind one agency's move toward local AI.
ReadWorking on something like this?
Bring the app or the process to a free 15-minute call. I will tell you what I would look at first, and whether I am the right person for it.