Your Coding Agent Doesn't Know What It Doesn't Know. Qodo Wants a Second Agent Watching.

Your Coding Agent Doesn't Know What It Doesn't Know. Qodo Wants a Second Agent Watching.

BackerLeader ●48 ●330 ●509
calendar_today • schedule4 min read

A coding agent looks at one service, sees a change that makes sense, and ships it. What it doesn't see is the billing service three hops away that depends on that exact behavior staying put.

"There could be another part of the system — one service using another service," said Itamar Friedman, co-founder and CEO of Qodo. "A billing service that's using a service that computes a certain part of what's being done in the product. You change something that looks logical, and suddenly your billing mechanism is completely broken. It sounds obvious, but it happens constantly, because billing is looking at a specific contract between two microservices, and the product is changing so often because you're using AI."

That's the failure mode Qodo built its new Agentic Toolbox to catch, and it's a believable one. Autonomous coding agents are good at local reasoning — the file in front of them, the function they're editing. They're much worse at knowing that the function they're editing is load-bearing for something three services away that nobody documented.

Qodo's fix is to give the coding agent a second agent with a different job: read the plan before the code ships, check it against everything the company has learned about its own system, and push back if something doesn't line up. Friedman calls it adversarial review, and he's specific that it's not just a second AI rereading the first one's output with a slightly different prompt — which he says is what most "AI code review" amounts to today.

"Usually, because of lack of resources or incentive, all they're doing is taking the same harness, changing the prompt a little, and asking it to review," Friedman said. He reaches for an analogy: bookkeeping and auditing look similar from the outside, but the incentives are different, so the work is different. Cloud performance and cloud observability are the same story. "Qodo intentionally has a dedicated harness, a dedicated context, that all comes down to adversarial review."

What that reviewer checks against is what Friedman calls a wisdom base — a standing, continuously updated model of a company's codebase, its history, its past incidents, and its own quality standards, built before the coding agent ever starts a task rather than assembled on the fly. He frames the difference as code review finally happening early enough to matter. "Until now, that codifying part was only served after the developer finished coding." Too late, in other words. The toolbox tries to surface that context — the kind a principal engineer would carry around in their head — during the work, not after a PR is already open.

Friedman is candid that this changes what "10x productivity" actually means. "Real-world software, 2x is amazing. No technology has ever done that. But that productivity is borrowed — from the senior developers, from the principals who know the system." His pitch is that Qodo is codifying what those people know, partly so the load doesn't fall entirely on the few engineers left who actually understand how the system fits together — a group he says is shrinking as senior engineers leave and take that knowledge with them.

The harder problem underneath all of this is that most of what makes code "good" isn't something a linter can check. Friedman breaks software quality into roughly a dozen categories — maintainability, testing, security, compliance, and others — each with its own set of questions, some of which have answers that are specific to a given company's own standards. A linter can catch a missing null check. It can't catch a policy like "never let a customer's tooling reach our database directly," which he gave as a real example pulled from companies letting customers build on their platforms. It also can't catch something closer to a values judgment: he described a rule against pulling a customer's data from Facebook to personalize a product recommendation, even though nothing in the code itself would look wrong. Those are semantic, not deterministic, and Qodo's admins write them in plain language rather than code.

None of this ships as an on/off switch. Friedman says Qodo deliberately avoids auto-approving or auto-blocking anything in a team's first week — the toolbox spends that time learning the codebase and building what he called a software graph, then spends the following week or two suggesting rules and standards before a team would reasonably trust it with real gatekeeping. He says that onboarding period has already shrunk from three months to about one, with a goal of getting it down to a few days.

He's willing to put numbers on where this goes, though they're his own projections, not an audited benchmark: 20 to 30% auto-accepted PRs and 30 to 40% auto-blocked by the end of 2026, climbing toward 80% auto-approved by the end of 2027. If that holds, he expects the pull request itself to stop being the natural unit of review, in favor of larger "work packages" — a bet worth watching rather than taking at face value, since it depends entirely on the adversarial agent's judgment being trustworthy at scale, which is exactly the part that isn't proven yet.

Qodo open-sourced its PR-Agent tool, and Friedman was upfront that the toolbox isn't going the same route. His comparison was Databricks and Spark: an open source tool can review a pull request, but a platform that accumulates a company's own standards and history over time is a different kind of product to run, and Qodo is betting enterprises will pay to have it managed rather than build it themselves.

The open question is the one Friedman didn't fully answer: how much of this adversarial judgment is actually earning trust through outcomes, versus being trusted because it sounds like the kind of thing a principal engineer would say. For a team handing over merge authority, that's the distinction that matters most — and it's one only time in production will settle.

1 Comment

2 votes
🔥 Join developers growing publicly
Share your knowledge, build in public, and grow your developer presence with a global community.

More Posts

Cisco's Amy Chang: A Model's "Passport" Doesn't Tell You Where It Actually Came From

Tom Smithverified - Aug 27

Your AI Doesn't Just Write Tests. It Runs Them Too.

Kevin Martinez - May 12

What Developers Already Know About Data Center Delays

Tom Smithverified - Sep 28

Systems Thinking: Thriving in the Third Golden Age of Software

Tom Smithverified - Apr 15

EKS Auto Mode: What It Actually Changes (and What It Doesn’t)

Alexandre Vazquez - Jul 27
chevron_left
19.1k Points • 887 Badges
260Posts
151Comments
127Connections
LLM Training & Evaluation Specialist with hands-on experience building major AI models. As one of th... Show more

Related Jobs

Commenters (This Week)

3 comments
1 comment
1 comment

Contribute meaningful comments to climb the leaderboard and earn badges!