A while back I wrote a series of blog posts about my first real experiences with AI-assisted coding — what I called “Vibe Coding Reflections.” If you haven’t read them, the short version is: I used Claude Code, GitHub Copilot, and AWS Q to modernize my hobby project LTHOI.com, and then spent about ten posts unpacking what that experience meant for enterprise software development. The two biggest limitations I flagged were that the AI behaved like a really knowledgeable but occasionally misguided intern, and that it was fundamentally trapped inside my IDE — it could only see what was in the repo, not the broader environment the code ran in. Since then the tools have evolved in ways worth talking about. Neither problem is fully solved, but both are materially better.
The Intern Now Raises Its Hand Before Acting
The intern problem, as I described it, wasn’t about intelligence — the AI knows an enormous amount. The problem was behavior. Like a new hire who’s read the entire internet but doesn’t understand the business world yet, AI coding assistants would confidently dive into an implementation that was technically coherent but practically wrong. They’d build something that was too clever, too complicated, or just not what you actually needed. By the time you noticed, it had already built a small city in the wrong location.
Anthropic has introduced a feature called Planning Mode that addresses this in a straightforward way: before Claude writes a single line of code, it describes everything it intends to do. The whole approach, laid out for your review. Once the plan is described, you can easily iterate on it. You can push back, ask Claude to do more research on a particular approach, or redirect entirely. Then, once there’s a good plan in the room, you execute with much more confidence.
I want to be clear about what Planning Mode is and isn’t. It doesn’t fix the intern. The intern still has the same instincts — including the occasional instinct to over-engineer things. What it does is move the human checkpoint to the right place: before the work starts rather than after you’ve untangled three hours of overcomplicated code. I recently implemented a Wine Sommelier Agent to add to “my team”. I call her Sam the Som. The planning took over an hour of back and forth with Claude but when I pushed “Go” we got to a working first version of Sam in a hurry and the whole project only took a few hours. If you’re curious, I exported plan and uploaded it with Sam to my Git repo, here’s the first version.
Breaking Out of the Box: Claude Hits the CLI
The IDE limitation was, in my original series, the complaint I felt most strongly about. Claude could see the code. It could not see the pipeline, the infrastructure, the runtime logs, or anything else that exists outside the repository. This meant that debugging a deployment failure involved me manually copying error messages out of one tool and pasting them into another — using myself as a relay between Claude and the real world.
That limitation is largely gone. With Claude Code running in the CLI and access to cloud credentials, Claude can now operate across your entire environment. I don’t mean it can look things up. I mean it investigates, builds hypotheses, tests them, watches what happens, and adjusts. I’ve sat and watched in genuine awe as Claude interrogated the logs of live GCP services, built small prototype scripts to test a concept, learned from the failure, and tried a different approach — all without me doing anything except occasionally redirecting.
Here’s a real example. I was working on a scheduled job feature in a middleware service I’ve built for managing agent communications. The job was supposed to send a message to a Vertex AI agent and, only if the agent replied, forward that reply to the user in Slack. After implementing a fix and deploying to Cloud Run, Claude checked the live logs and saw the fix hadn’t worked. What followed was a genuine debugging loop — entirely driven by Claude:
[Checks Cloud Run logs]
Error acquiring lock for job [job-id]: 'AsyncTransaction' object is not callable
The error is still occurring after deployment. Let me verify the deployed
code and investigate the transaction API more closely.
[Runs local Python script to inspect the Firestore library's AsyncTransaction class]
Supports async context manager: True
Now I understand. AsyncTransaction is an async context manager, not a
decorator. Let me fix the code properly.
[Amends commit, force-pushes, redeploys to Cloud Run]
[Checks logs again — different error now]
Error: Transaction not in progress, cannot be used in API requests.
The error changed. Let me research the correct async transaction pattern.
[Runs another local test script against the live Firestore service]
[After further testing, recommends simplifying away from transactions entirely]
What you’re seeing in that exchange isn’t a code editor with a chat window. It’s a developer who has access to the whole stack — reading logs, probing library internals, deploying, watching what happens, and learning from it. The fact that it chose an overly complicated approach initially is a reminder that the intern is still an intern. But the fact that it could investigate and resolve a multi-layer deployment problem on its own, across GCP services, with nothing from me but a redirect — that’s something genuinely new.
So What Does This Mean?
When I ended the Vibe Coding Reflections series, I pointed to MCP servers and agentic frameworks as the likely solutions to both the IDE limitation and the intern problem. That’s still true at the enterprise level, where you want purpose-built agents tuned for specific workflows. But at the level of a developer working in Claude Code, those capabilities are arriving faster than I expected. The IDE-bound code-suggestion assistant is becoming something much closer to a junior engineer who owns the problem end to end.
The intern metaphor still applies — you need a competent person in the room who can evaluate plans and redirect when the approach gets too complicated. But it’s a better intern than it was six months ago. And the room just got a lot bigger.
