Everyone Ships Agent-Generated PRs. This Week the Tools to Review Them Arrived.

Three AI review tools for agent-written code launched in 48 hours. What the wave gets right, where it is thin, and how to review agent PRs with local deterministic scanners plus your own key.

Everyone Ships Agent-Generated PRs. This Week the Tools to Review Them Arrived.

This morning a tool called Jev hit the Hacker News front page with 40 points and 44 comments (as of this writing). The repo is six days old and already at 45 stars. The day before, two more launches landed: Critic, a review tool that positions itself around understanding the code your agents write, and Canary (YC), pitched as independent verification for AI code. Three AI code review tools in 48 hours, all attacking the same moment in the workflow: the pull request nobody has fully read.

The input changed

The first wave of AI review tools, roughly 2024 through 2025, reviewed human pull requests. A person wrote the diff, a model commented on it, and the human author decided what to take. Accountability was never in question because the author had a name.

The new wave inverts the setup. Jev's README leads with reviewing behavior rather than diffs. Critic's homepage tagline (accessed September 25, 2026) is one sentence: understand the code your agents write. Canary skips explanation entirely and sells verification. The diff author is no longer a person with a name. It is an agent that produced 900 lines at 2 a.m. while nobody watched.

That inversion is the whole story of this week. The tools are not competing on review quality in the abstract. They are competing on a new question: what does it mean to approve work that was generated, not written?

The accountability gap

Here is the question none of the three launches answers. When a person writes a PR and a reviewer flags an issue, responsibility is clear: the author fixes it or defends it. When agent A writes the PR and model B approves it, the merge button is still yours. What exactly are you approving?

Our take, after shipping review tooling ourselves: the hardest agent-written bugs are not style or syntax. They are behavior bugs, and catching a behavior bug requires knowing what the code was supposed to do. That spec usually lives in the agent's prompt or the issue thread, and it rarely survives into the PR description. A verifier with no spec can only check invariants, which is why the deterministic layer matters more in this wave than the last one.

This is also where the wave's economics get interesting. All three tools rent intelligence from the same handful of frontier models. If the writing model and the reviewing model share a blind spot, the verification is weaker than it looks. Independence, the thing Canary sells, is mostly a property of using different systems, not of adding a second API call.

Where local processing fits

Jev's own README is refreshingly blunt about data flow: analyze sends changed code and nearby context to TypeSafe and OpenAI. For a cloud-native team, that is a reasonable trade. For everyone else, it re-creates the exact problem review tooling was supposed to reduce: your code sitting in someone else's pipeline, this time including agent-written code that no human on your team has read end to end.

We are not neutral here. We build one of the alternatives. cora-code is our CLI-first review tool, and its deterministic layer runs with no model and no network call: 12 built-in rules, 13 security patterns, and 15 secret detection patterns, all local. When you do want model review, it is bring-your-own-key against whichever provider you already pay, including a local Ollama endpoint. Findings come out as SARIF for GitHub Code Scanning, hooks catch issues pre-commit, and an MCP server exposes 18 tools so your coding agent can request a review before it opens the PR.

The honest numbers: v0.15.0 shipped August 31 on GitHub and crates.io the same day, the repo sits at 28 stars, and the license is Apache-2.0. That is an early-stage project, and we would rather state it plainly than round it up. The deterministic scanners and the BYOK setup are the parts that matter for this debate, and both work today.

Picking a tool this month

Match the tool to your constraint, not to the demo. If your code can leave your machine, pick on workflow fit: Jev for GitHub-heavy teams that want a browser extension in the loop, Critic for keeping agent context reviewable across devices, Canary if verification theater is what your compliance process needs.

If your code cannot leave, your option set shrinks to a local deterministic layer plus a model you point at yourself. Then, whatever you pick, keep two rules that no tool enforces for you: a human presses merge, and the agent ships its tests with the diff. A review tool that cannot see intent is a lint pass with better marketing. The tests are where intent lives.

What's next

The wave is three days old. Watch which bets survive contact: explanation (Critic), verification (Canary), or behavior-first review (Jev). Our own next step is smaller and unglamorous. Every PR our agents open gets the deterministic layer first, the model second, and a human last. If that order feels obvious to you too, the interesting question is why the front page keeps discovering it fresh.

cora-code lives at github.com/codecoradev/cora-code, installable via the quick installer or cargo install --git https://github.com/codecoradev/cora-code. The Jev discussion is on Hacker News, Critic is at critic.run, and Canary is at runcanary.ai.