August 19, 20268 min read

AI-assisted coding speeds up what you already know (and disguises what you don't)

Cover image about AI-assisted programming, featuring code, development interfaces and technology graphics, alongside the title “Programming with AI accelerates what you already know how to do (and disguises what you don’t)”

I started using AI as support while programming in 2022, when ChatGPT opened to the public. Codex came in 2025, and I didn’t use it the way people usually describe: I didn’t ask it to build me an app from scratch.

I gave it a project I had already built, with my components, my styles and my decisions already made, and asked for something specific: to build new components that respected the logic and style of mine.

Looking back, I think that framing is what made it work, and also what taught me where it stops working.

Since then I’ve used Codex and Claude. Here’s what I’ve taken away: where they’ve saved me hours, where they made me lose them, and what I had to do with each.

Why it helps to respect what’s already there

If you ask a model for a component from scratch, it has to invent the conventions: how to name things, how to organize the styles, where to put the logic. It’ll invent something reasonable, but it won’t be yours.

When you give it a project where the decisions are already made, what it does is follow an existing pattern, and it does that very well.

It saved me a huge amount of time. What it wrote wasn’t code I couldn’t have written; it was the code I was going to write anyway, and it left the decisions to me. I come back to this at the end, because it explains almost everything else.

The best use turned out not to be writing components

What has helped me most isn’t what I expected: it’s been working with data.

On real projects, you often don’t start from a database but from a .csv someone exported from another system, or an .xlsx with several sheets, colors and merged cells that made sense at some point.

Before migrating that, you have to understand it: which columns are really there, which ones are the same thing written two ways, where the duplicates are, which fields are really a relationship hidden in a string.

AI has been very useful for that. It helps me analyze the files, understand the data and convert it to .json so I can migrate it into real databases, improving the models and normalizing along the way.

I think it works so well there because the work is tedious, mechanical and verifiable. The records either add up or they don’t, the keys either resolve or they don’t, so it’s hard for an error to slip through unnoticed.

200 GB nobody had read

The second use that’s given me the most value has nothing to do with writing code. It’s about reading it.

I had to review servers with around 200 GB of files, looking for obsolete or dangerous code. Done by hand it isn’t difficult, it’s just impossible to cover. Nobody reads 200 GB; in practice you search for the patterns you already suspect, so you only find what you were already looking for.

With AI I found important vulnerabilities that would have taken me hours to locate. And, above all, I was able to fix them without breaking anything else, which in a legacy system is half the problem.

I think it’s one of the most underrated uses. It barely gets talked about because it doesn’t produce anything new and doesn’t look good in a screenshot; what it produces is an understanding of something that already existed and that nobody had time to look at, which in legacy code is exactly what’s missing.

The risk here is real, though. What you get back is a list of suspicions, not a verdict. Every finding has to be confirmed by looking at the code, and every fix has to be understood before you apply it. Otherwise you can end up trading one security problem for another.

I tried pure vibe coding, and no

I also ran the opposite experiment, and I’m telling it because it’s what gets published least.

I went all in on vibe coding: ask for things and see what came out, without steering, without reviewing closely, trusting the whole thing would hold itself up.

Nothing good or functional came out of it for a real use case. And the worst part is that it wasn’t producing obvious garbage.

Without clear, specific instructions, the model imagines things that work in theory but don’t hold up in production. It compiles, it looks fine and it seems to make sense if you skim it. The failure is in the decisions it made for you to fill in what you didn’t tell it.

An obvious error you see and fix in two minutes. One that looks correct gets into the repository and stays there.

Even so, I didn’t conclude that the tool is useless. Used carefully, keeping strict track of what gets done and what doesn’t, you can build really good things. Between the disaster and what worked, the model was the same; what changed was how much judgment I put in the middle.

Placing components isn’t designing a system

This is what I learned most, and what I’d like someone starting out to take away.

Dropping components (charts, buttons, menus) onto a page and arranging them any old way has nothing to do with designing a real dashboard that works and can grow.

At first glance they look like the same task. Both end up as a screen with cards, charts and a side menu, you can ask for either one with almost the same words, and in both cases the model will hand you something that looks good.

But the second one involves decisions that don’t show on screen:

  • Which data loads up front and which on demand, for when the table reaches ten thousand rows.
  • What gets recalculated when a filter changes, and what shouldn’t be.
  • Where the state shared by three widgets lives, and who wins when two of them touch it.
  • What happens when a request fails and half the dashboard is left without data.
  • How you add widget number twelve without rewriting the other eleven.

None of that shows up in the screenshot, and all of it shows up six months later.

That’s where what I consider the rule comes from:

If you haven’t had experience building or using the kind of thing you’re asking for, you risk creating a “product” that starts dragging errors along over time.

The problem isn’t that the AI gets things wrong too often, it’s that you can’t evaluate what it gives you. And if you don’t evaluate it, you’re really just accepting proposals blindly, even if it feels fast.

The technical debt this creates doesn’t show up on day one. It shows up when you need to add something nobody planned for and discover that the structure you’re building on was never chosen by anyone.

What it costs

There’s another consequence I think is one of the most practical, and I hardly see it mentioned: it doesn’t cost everyone the same.

Someone with no programming background who wants to build their website or app is going to spend a lot more than someone with experience. Partly because they write worse prompts, but mostly because every answer you can’t evaluate is one you’ll have to ask for again.

Without judgment, the cycle is: ask, look, not know if it’s right, ask again. With judgment, you ask for something scoped, review it, correct the course in one sentence and move on. One costs five attempts and the other costs one.

That’s why what makes it pay off isn’t just knowing how to write good prompts. It’s knowing how to program, design and ground ideas: knowing what you’re asking for, recognizing when what comes back is good enough, and noticing when it has drifted before you spend three more attempts.

Because you don’t depend entirely on the tool, it ends up cheaper. It’s a complement to what you already know how to do, and what each credit is worth depends on what you bring.

It also answers the question of whether this will replace anyone: the person who gets the most out of it is the one who needs it least.

How I use them today

Between Codex and Claude I don’t have a strict split: I use both to refactor and write code, backend and frontend. What I am clear on is the method, which is the same with either one:

  • I give it context instead of asking it for judgment: my components, my conventions, the code that already exists. The more scoped the instruction, the better the result. Scoped doesn’t mean short; it means it doesn’t leave important gaps for the model to fill on its own.
  • Where the result can be verified, I let it work on its own: data migrations, transformations, going through a whole server looking for dead code. That’s where I save time without taking on risk.
  • Everything that goes to production I review, keeping track of what changed and what didn’t. That was exactly the difference between my failed experiment and the projects that went well.
  • Architecture decisions I make myself. Not out of purism, but because they’re exactly the ones I can’t review by reading the result. A badly written component is visible; a badly chosen structure doesn’t show until it’s expensive to change.

In the end, AI speeds up the judgment you already have, but it doesn’t lend it to you. If you wouldn’t know how to do what you’re asking for, you won’t be able to judge what you get back, and on top of that it’ll cost you more.

Share