A canal towpath and grassy bank beside a lock gate fading into thick fog, with no people in frame

GitLab says coding agents have to earn their autonomy

From Mexico, the question I keep asking about coding agents isn’t “can it write the fix?” It’s “who let it merge?” GitLab put a public answer on the table on October 6. The page is called the GitLab Security Standard, it’s marked v1.0, and its core line is short: “autonomy isn’t given. It’s earned, one stage at a time.” I haven’t tried GitLab’s agent platform. That sentence is theirs.

It came out the same day as GitLab’s Transcend roundup, which says they made “over a dozen announcements across all four layers of our architecture for agentic software engineering.” Most of that list is about making agents go faster. The standard is about how far they get to go. I think the pairing is the story.

The speed half

The Transcend post opens on a complaint I recognize. “Today, each step in your agentic workflow still waits for a person to approve it before the next one starts.” Their answer is goal-driven flows, the /goal command, listed as GA this month. They say it takes a stated objective and runs it “end to end, through review, tests, security scans, and approvals from the Duo CLI, headless mode, or Agentic Chat.” When GitLab first described /goal back in September, the pitch was that you stop “effectively acting as its continue button until the task is done,” and that “a separate model checks that work against your stated goal at each step.” I haven’t run it.

Around that, the same post lists Custom Flows to chain agents (GA), an MCP Server so “external AI applications and coding agents” can reach a GitLab instance “under the same policy your own agents follow” (GA this month), GitLab-hosted open-weight models, and Orbit, a context graph of the software lifecycle coming GA next month. GitLab claims Orbit gives “up to 45x fewer retries” and says more than 3,500 organizations used it in beta. Those are their numbers. I haven’t checked them.

The brakes half

The standard starts from a premise I agree with. “Organizations are transitioning to a world where agents build software without a human in every step.” Then it says how to get there without pretending the risk went away. “Code is abundant. Trust is scarce.”

The shape is a pyramid. At the bottom sits a foundation of policy, audit, and ownership. “Every agent, service account, and control is accountable to a named human.” Above that are five stages: Authorize, Isolate, Verify, Release, and Respond. “The more an agent proves, the more autonomy it earns.”

Cream-paper schematic of a five-step staircase on a base slab, with three solid steps, two faint hatched steps, a filled disk resting on the third step, and a curved arrow from the top step back to the slab
Three stages proven, two still faint. Failures loop back to the base before the next step. Original schematic for this post, based on GitLab’s Oct 6 standard.

Each stage asks one plain question and lists example controls. Isolate wants work to run “in a contained environment whose boundaries hold even when the agent follows a malicious instruction.” Verify has two lines I’d tape to my monitor: “The actor cannot edit the pipeline or policy that judges it” and “The author, human or agent, cannot be the sole approver.” Release says “Deployment decisions bind to the exact artifact that passed verification, identified by digest and provenance, not by name or version label.” Respond starts with a tested kill switch: “Stop a task, revoke its access, pause releases.”

Three levels, per workflow

The part I like most is that autonomy isn’t one global switch. “Autonomy is granted per workflow, and each workflow must pass every stage’s tests before it earns more.” GitLab spells out three levels. Level 1, Suggest: “Recommend; humans make the change.” Level 2, Propose: “Open the MR; a human approves release.” Level 3, Act with guardrails: “Fix, test, and merge; humans review by exception; volume and blast-radius caps apply.” Level 3 needs all five stages, “including a tested stop.”

The examples are where it gets honest. The Level 3 example is “Dependency updates for reachable, known-exploited vulnerabilities.” The Level 2 example is application security fixes. The Level 1 example is “Pipeline and policy config (changes to the controls themselves).” So the agent can merge a narrow dependency bump on its own, but when it touches the rules that judge it, it only gets to suggest. That’s the line I’d draw too.

There’s also a loop for when it goes wrong. “When something fails, learn from it and update the policy, so the rules improve before autonomy is restored.” The page maps the stages to OWASP SAMM, the OWASP Top 10 for Agentic Applications, MITRE ATLAS, and NIST SSDF and SP 800-53, so a security team has something familiar to hold it against. GitLab says it “runs this standard against its own software factory,” and that it’s “public, ungated, and always evolving.” I haven’t audited that claim.

Where the agent adds packages on its own

One of the Isolate controls is “Trusted dependencies,” and GitLab shipped a product for it the same day. The Dependency Firewall post, in early access, names the problem directly: “AI coding agents now add open source dependencies on their own, often without anyone reviewing or even seeing what landed in the build.” The firewall checks malicious status, vulnerability severity, license, and package age before install. It starts in a warn mode that only records what a policy would catch, then moves to block. A logged bypass exists for real exceptions. I like the package-age rule most, since it stops a version published minutes ago from going straight into an agent’s build.

What I’d actually check

I haven’t used these GitLab features, so I’m reading this as a checklist more than a product. For any coding agent setup, I’d ask the Verify questions first. Can the agent edit the CI file that grades its own work? Can it approve its own merge request? If the answer is yes, it’s sitting at Level 3 without having earned Level 1.

Then I’d ask whether I can actually stop it. The standard treats a tested kill switch as the price of real autonomy, not a nice-to-have. Starting a run is one command in most tools. I’d want stopping it, revoking its access, and pausing a release to be just as fast, and tested before I need it.

I’m also not claiming GitLab solved this. Several pieces named on the page are labeled beta, closed beta, or early access, and goal-driven flows are “GA this month,” not today. The standard is v1.0 and dated October 6, 2026. What I like is that the same company selling the fast path also wrote down, in public, when the fast path should stay closed.

Featured image: Caen Hill locks on a misty morning, England, 26 August 2019, by Wingrider13, Wikimedia Commons, CC BY-SA 4.0. This image is a derivative, shared under the same license: a 1400×900 crop of the original upload, with a mild contrast lift, slight warm grade, and light grain.