A steel gold pan holding a small pile of wet gravel and pebbles, resting in dark water inside a weathered wooden sluice frame, with no people in frame

Reviewing agent code is a trust problem, not a diff problem

From Mexico, the part of AI coding that eats my time isn’t the writing anymore. It’s the reading. An agent hands me a change that touches a dozen files, and I have to decide if I trust it. On October 6, JetBrains Research published a post that names that feeling better than I could: “Our Framework for Reviewing AI-Generated Code”, by Katie Fraser, Agnia Sergeyuk, and Ilya Zakharov. Their opening line is the whole problem: “The bottleneck is no longer code generation; it’s review, and nobody’s quite sure how to do it well yet.”

The post summarizes a paper from JetBrains’ Human-AI eXperience team and researchers at Lund University, published in the ESEM 2026 proceedings and set to be presented at Empirical Software Engineering International Week in October. It’s a design study, not a product launch, so there’s nothing to install. But it gave me a better way to think about my own review habits.

Why my usual review instincts don’t fit

When I review a teammate’s pull request, I bring a lot of context I don’t even notice. I know which parts of the codebase they know well and which parts they rush. If something looks odd, I can just ask. The authors point out that “none of that exists when the author is an LLM.”

Worse, a model “presents every line of generated code with the same apparent confidence, regardless of how uncertain it actually was when it produced that line.” Their example landed for me: “There’s no signal telling you that the authentication logic was straightforward and the database migration was a stretch.” Everything reads the same. So the rational move is to read every line, and that, in their words, “scales terribly” as agents produce bigger change sets.

Trust calibration, not diffing

This is the core idea. A classic diff viewer assumes the reviewer’s job is to understand what changed. With agent output, the authors argue, understanding what changed is the easy part. “Knowing whether to trust it – and where – is the hard part.” They call the skill trust calibration and define it as “the capacity to allocate review effort proportionate to segment-level risk when the author can’t be interrogated about their confidence or reasoning.”

In plain words: knowing where to look hard and where you can skim. With human code, social cues help you do that. With AI code, the post says, current tools give “almost no support for at all.” The line I’d put on a sticky note: “The diff viewer was the right tool for reviewing what your colleague wrote. It may not be the right tool for reviewing what your agent wrote.”

Cream-paper schematic read left to right: a grid of identical grey page tiles, an arrow to a single column of tiles that each carry a heat bar (mostly pale, one gold, one red), and a large magnifying loupe over the red tile where one enlarged line is underlined in red
First every file looks the same. Then each file gets a risk signal. Only then do you zoom in on the lines that matter. Original schematic for this post, based on the three-level workflow described by JetBrains Research.

Three levels: overview, files, then lines

The proposed workflow follows an old information-visualization rule from Shneiderman: overview first, then zoom and filter, then details on demand. The reviewer first forms high-level hypotheses about the change, then drills down selectively to check them, instead of starting with line one of file one.

For tool builders, the post turns that into three jobs. Overview-level tools should stand in for the conversation you’d normally have with the author. File-level tools should “do risk stratification before the reviewer reads a single line.” And snippet-level tools should bring back the careful line-by-line work, helped by “chunk decomposition and chain-of-thought linkage.” The question to organize around, they write, isn’t how to display the diff. It’s “At what granularity does the reviewer need to allocate attention, and what signal do we need to surface there?”

The paper lists seven design constructs behind this: chunk, risk-per-line, risk-per-file, judge, walk-through, zooming in/out, and security cage. The blog post also notes that some existing review tools already cover pieces of it, like prose walkthroughs of a pull request or findings tagged by severity, but says no single tool fully brings it together, especially not around trust calibration.

How strong is the evidence?

I want to be fair about scale. According to the paper’s abstract, 17 industry practitioners took part in the discovery workshops, seven of them came back for the design phase, and a follow-up survey of 43 practitioners evaluated a semi-interactive prototype. All three workflow levels scored above the neutral midpoint, with means of 3.50 to 3.91 on a five-point scale. 63% of respondents expected less overall review effort and 52% expected less effort spent judging trust, compared with their current tools.

Those are expectations about a prototype, not measured time savings on real teams. The authors frame it as “a positive direction for future tool development,” and I’d read it the same way. It’s a well-argued design direction, not proof that reviews get faster.

Teams are feeling the same squeeze

A day later, on October 7, JetBrains’ Air team published “Faster Developers Don’t Make a Faster Team” by Gleb Malkov, based on discovery calls with engineering organizations. It’s from a product team building a team tool, so I read it as field notes, not research. One section heading says it all: “Implementation got cheap. Review didn’t.” The post quotes a mapping company that measured it: “If we looked at the pure numbers, PR time actually increased.” The post’s own conclusion is sharp: “If verifying the output costs more than the savings it generated, the team has actually regressed.”

What I’m taking into my own reviews

I don’t need a new tool to borrow the shape of this. Before I open the diff on an agent’s change, I want a short summary of what it thinks it did and why. Then I sort the files by risk myself: anything touching auth, data migrations, money, or deletes goes to the top, and generated boilerplate or test fixtures go to the bottom. Only then do I read lines, and I read the risky ones slowly.

The warning in the post is the part I’d underline for anyone building review tooling, including AI review bots. Tools that only help you understand the code “may improve efficiency for low-stakes changes while leaving the more consequential failure mode untouched.” That failure mode is spending your attention on the safe parts and skimming the dangerous ones. If agents are going to keep writing more of the code, I’d rather get good at deciding where to look than pretend I can read everything.

Featured image: Gold Pan, October 2004, by Nate Cull, Wikimedia Commons, CC BY 2.0. Cropped, resized to 1400×900, and color-graded from the original upload.