AI can write code, but who checks what it wrote?

Today an AI agent can take a task, read an existing project, build a new feature, run the tests and suggest a fix. The question is no longer whether developers will use AI. It is how much we hand over to it.

By Nemanja Pavlović  ·  CEO / Founder

5 min read

AI can write code, but who checks what it wrote?

A few years ago, AI mostly finished a few lines of code for a developer. That change is huge for the industry. Development gets faster, repetitive work gets automated, and developers spend more time on architecture, product logic and problems that need real thought. But one thing is easy to forget when a feature that took hours appears in minutes: code that works is not always good code.

AI can write a feature, but it does not always know the full context

Picture an AI agent that gets a task to build a login, connect a new payment method or add a data export. The agent builds a solution in minutes, and it works at first glance. It does not always know why the system was built a certain way three years ago. It does not know which other features depend on that code, or which business rules sit behind it. It does not know what happens when ten thousand users arrive instead of ten.

In a serious project, code never exists alone. One change can touch the database, performance, security, integrations with other systems, or a feature that seems unrelated. Understanding that wider context is one reason why development is more than writing syntax.

AI is very good at answering the question “how can we build this?” It matters more that someone still asks:

Should we build this the way it is being built?

The most dangerous code is not the code that fails

When code fails, the problem is usually easy to see. The app shows an error, a test breaks, or the feature cannot be used. Code that works but is built badly causes far more trouble.

It can hide a security hole that nobody finds for months. It can pull in a library the project did not need. It can send too many queries to the database and run fine until the app gets many users. It can copy logic that already exists elsewhere, or solve today's problem in a way that makes every later change harder.

AI models are convincing because the result often looks fully correct. The code is neatly structured and the function names make sense. A comment explains what happens, and everything looks like the author is sure of the work. That is not proof that the solution is good.

Security matters even more

As coding tools become more capable, we no longer talk only about AI that suggests text inside an editor. An agent can read files, run commands, install dependencies, browse the internet and use other development tools.

That is why OWASP published dedicated guidance in 2026 for working safely with AI coding agents. The recommendations include isolated development environments and limited access to tools and data. They also ask for control over dependencies and required human review of changes that touch CI/CD, infrastructure or production.

In other words, the more an agent can do, the more important it is to decide what it is allowed to do. Nobody gives a developer admin access to every company system because of one feature. The same rule applies to AI.

Code review matters less? No, it matters more.

If AI produces more code in less time, a team finishes more work. But the amount of code that someone must understand and check grows too. So the role of code review does not shrink. It changes.

It is no longer enough to see whether a function returns the expected result. The reviewer checks how it fits the existing architecture and whether it adds new dependencies. The reviewer also checks how it handles errors, what happens in edge cases, whether it follows project standards and whether it opens a new security risk.

AI is already used on the other side of this process as well. GitHub, for example, is building AI-assisted code review and security review that find some kinds of vulnerabilities before the code reaches the main branch. AI writes part of the code and helps check it. The team that builds the product still owns what ships.

And where are the tests?

If there is one place where AI clearly helps a development team, it is testing. It quickly proposes unit tests, finds scenarios the developer missed and checks many variations of a feature. But a large number of tests does not mean the app is well tested.

AI can write a test that only confirms the same assumption it made while writing the code. If the assumption is wrong, the result looks great: a wrongly built feature that passes every test. So we test what was written and also what was meant to be written. That is where the specification, the business logic, the team's experience and an understanding of users come back in.

So does AI really speed up development?

Yes, and probably by more than we can measure today. For some tasks the difference is huge. Boilerplate code, migrations, documentation, tests, code cleanup and small isolated features all get much faster. So does exploring an existing codebase. A developer no longer starts every problem from a blank screen.

But typing speed never decided how fast a good product appears. Before development starts, someone must understand what we are building. Someone designs the architecture, picks the technology, learns the business logic and defines how the systems talk to each other. After the code exists, it must be tested, reviewed, deployed, monitored and maintained.

AI can speed up almost every one of those steps. It cannot make them disappear.

What changes in a developer's job?

Less time goes to typing code and more goes to making decisions. The developer increasingly gives the agent context and splits a problem into smaller parts. Then the developer judges the proposed solution and notices when something that looks right is wrong for the system being built.

So experience does not lose value now that AI programs too. It gains value. A junior developer with AI builds something that works much faster. A senior developer sees much more easily why that solution does not belong in production. That is probably where the biggest gap will open.

AI-generated is not the same as production-ready

Today you build a landing page, a prototype, a small app or even a fairly complex feature in a time that was unthinkable not long ago. That is great. The problem starts when we mistake speed of delivery for quality of the product.

Production-ready software is reliable, secure, maintainable and built to keep growing. It works when something goes wrong, when the number of users grows and when the developer who built it no longer maintains it.

So the question is no longer whether AI can write code. It can. The more interesting question for the coming years is: who knows enough to judge whether that code is really good?

AUTHOR

Nemanja Pavlović · CEO / Founder

Writes about strategy, sales and digital products.