An old iron key left in a round, worn metal escutcheon on a warm wooden door, with a vertical door frame on the left and no people in frame

Agents aren’t making developers careless. They’re making more code.

From Mexico, every time I let a coding agent push a branch, a small voice in my head asks the same thing: did it just commit a password? On October 7, GitHub published an essay that gave me real numbers for that worry: “Secret protection must scale with software”, by Erin Havens, the product lead for secret scanning at GitHub. It opens with a stat that stopped me: “Today, one in three pull requests on GitHub involves an AI agent. A year ago, that number was fewer than one in 10.”

The easy story would be that agents are sloppy and developers stopped paying attention. The essay argues the opposite, and it brings nine quarters of data to back it up. Its one-line version is the subtitle: “Developers aren’t becoming more careless; they’re being outpaced.”

More code, not worse habits

Here’s the part I found most useful. According to GitHub, “A new secret appears in publicly visible code about once every two seconds, doubling yearly for the past three years.” That sounds like a carelessness problem until you look at the denominator. Between Q2 2024 and Q2 2026, screened pushes grew 2.84 times, while pushes carrying credentials grew 2.59 times. GitHub says it “found no statistically detectable trend regarding per-push prevalence.” GitHub’s chart puts Q2 2026 at 574 million public pushes, with 0.47% of them carrying a detected secret.

So the rate per push is about the same. There’s just a lot more pushing. And developers seem to take the warnings more seriously, not less: over the same period, the share of push blocks that developers overrode fell from 6.63% to 3.93%. The essay’s conclusion is that these numbers “challenge the common claim that agents are causing developers to become more careless.”

I like this framing because it’s honest about where the risk actually comes from. If the same small share of pushes leaks a secret, and we’re pushing far more than we used to, the number of leaked secrets goes up even if nobody gets worse at their job.

Prevention scales with compute. Cleanup doesn’t.

The cost shows up after the leak. GitHub says “The mean time to manually revoke a secret hovers around 40 days; roughly one in five took more than 90 days.” That’s weeks or months where a credential can still work. If each exposure needs the same human response, doubling the activity doubles the workload too.

GitHub’s push protection is the early catch. It stops recognizable credentials before they land in repository history. But the essay is clear about its limits: including additional secret types, push protection stops about 30% of newly detected secrets before they enter history, and the other 70% are found after the credential is already out. That’s where the essay’s key line comes from, and it’s the one I’d put on the wall: “Prevention scales with compute, but remediation still scales with people.”

Cream-paper schematic read left to right: a widening funnel of grey dots with small gold keys mixed in evenly, a dashed vertical line with a green comb that catches keys into a tray below, then stacked pale bands holding three dark keys circled in red, tied by a red dotted line to an hourglass with most of its sand still on top
More pushes, same share of secrets. What gets caught at the push is cheap. What slips into history waits on a person. Original schematic for this post, based on the data in GitHub’s essay.

A small model right at the push

The hard cases are secrets with no pattern. A provider token often has a recognizable prefix. An internal database password can be completely random, and the only clue is the code around it. GitHub calls the tradeoff a “four-body problem”: precision, latency, throughput, and cost all pull on each other. A false positive interrupts a developer and makes the next block harder to trust. A slow or expensive check can’t run on every push.

Their answer isn’t a big chatbot. It’s a fine-tuned ModernBERT classifier, built with Microsoft Applied Sciences, that reads the surrounding code and scores candidate secrets “without generating code or prose.” GitHub says it evaluates candidate batches in under two milliseconds, and that it’s more precise than existing LLM-based pipelines. The example in the post: it blocks password-like values in a database URL, a Kubernetes Secret manifest, and a Dockerfile, while letting the placeholder changeme through. GitHub says adding the model to push protection could more than double the number of secrets it’s able to prevent.

I think that’s the real engineering lesson here, beyond secrets. Not every AI check in your pipeline needs a large model. A small classifier that’s fast and cheap enough to sit in the critical path can do more good than a smarter model you can only afford to run later.

What’s shipping, and what it costs

The GitHub changelog entry from the same day lays out the rollout. Customers with AI-detected password alerts were moved to the new model automatically, and those alerts stay included with GitHub Secret Protection and GitHub Advanced Security at no extra charge. AI-detected secrets in push protection is in private preview, and the essay says it’ll reach organizations with Secret Protection on Enterprise Cloud and GitHub Teams later this month. The model is also planned for GitHub Enterprise Server 3.23 in public preview for AI-detected alerts.

For individual developers, the classifier is coming to the /security-review command in the Copilot CLI and Copilot app, which the changelog lists as available soon in private preview. The checks can be used without a Secret Protection license. Two details I’d flag before anyone turns these on: the new push protection and security-review checks consume GitHub AI Credits, and the changelog notes that a check “can consume credits even if it doesn’t block a push.” The new review checks are off by default, and the changelog even says “Agents shouldn’t enable credit-consuming features or change policies or budgets without explicit authorization.” I appreciate that line. Set a budget first.

What I’m doing in my own repos

None of this changes the basics, it just raises the stakes. My agents don’t get real credentials in the files they edit. Config reads from the environment, and example files use obvious placeholders. If push protection blocks something, I treat it as a real alarm and rotate, not as noise to bypass. And when something does slip, I revoke first and clean history second, because a 40-day average is way too long for a key to stay live.

The essay ends with a line I agree with: “We want people to build more software. Our capacity to protect it should grow with our capacity to create it.” If agents are going to write a bigger share of the code, I’d rather have the boring, fast checks at the push than trust that everyone, human or agent, will just be more careful.

Featured image: Old key in door lock, August 2016, by Santeri Viinamäki, Wikimedia Commons, CC BY-SA 4.0. Cropped, resized to 1400×900, and lightly adjusted from the original upload; this adapted image is shared under the same CC BY-SA 4.0 license.