Search

· Strategy Â· 12 min read

Agentic Coding Security: Governing AI-Written Code in the SDLC

Agentic coding security starts by treating agent-written code as a supply-chain input: track provenance, scope permissions and keep audit trails.

Featured image for: Agentic Coding Security: Governing AI-Written Code in the SDLC

In my work with engineering teams adopting coding agents, the first security question is almost always the same: “Can we detect which code was written by AI?” It is a natural question, and it is the wrong one. I’ve watched teams spend a quarter evaluating detectors, only to discover that the pull request they most needed to worry about was one a human had pasted from an agent session, lightly edited, and merged on a Friday.

Detection fails for a boring reason. Code is text, and once a person has touched it, reformatted it or run it through a linter, the statistical fingerprint is gone. Even if detection worked, it would answer a question nobody needs answered. Your risk was never that a machine typed the characters. Your risk is that code entered your build without anyone being able to say where it came from, what it was allowed to touch, or who approved it.

That is a supply-chain problem, and the industry already knows how to think about those. We do not ask whether a dependency “looks malicious”. We ask where it came from, who signed it, what it can reach at build time, and what we recorded when we accepted it. Agent-authored code deserves the same treatment, and the same tools mostly apply.

Key Takeaways
  • Stop trying to detect "AI code". Record provenance at the point of authorship instead, so every change carries the agent, model, prompt context and human approver.
  • The largest risk is what the agent can do while it writes, not what it wrote. Scope credentials, network access and tool permissions per task.
  • Keep the existing gates (review, tests, scanners, signed builds) and make them stricter on trust boundaries, because agents produce plausible code faster than reviewers can read it.
  • For leaders: this is a control-design problem with a measurable cost, and it is cheaper to put in place before an incident than after an audit finding.

Why “is this AI code?” is the wrong control

Consider what a detection-based policy asks of you. You need a classifier with an acceptable false-positive rate, applied to every diff, in an environment where developers have every incentive to route around it. Developers who get flagged will rewrite a few lines or switch tools. Developers who are honest will be penalized for it. You end up with a control that punishes transparency.

Now consider a provenance-based policy. You require that every change records how it was produced. An agent session that opens a pull request attaches metadata automatically: which agent, which model version, which repository it had access to, which tools it called, and which human reviewed the result. Nobody has to guess. A change with no provenance record is not “suspected AI code”. It is an unattributed change, and your pipeline treats it the way it treats an unsigned artifact.

This reframing matters for the business reader too. Detection gives you a statistic you can put in a slide. Provenance gives you something an auditor, an insurer or a customer’s security questionnaire can actually verify. When a regulator asks how a defect reached production, “we believe about 30 percent of our code is AI-assisted” is not an answer. “Here is the record for that change” is.

Treat the agent as a build-time actor with an identity

The mental shift I find most useful is this: a coding agent is not a faster keyboard. It is a non-human principal that reads your repository, runs commands, calls tools and writes to systems. Security teams already have a word for that. It is a service account, and service accounts get least privilege, short-lived credentials and logs.

In practice, that means a few concrete rules.

Give each agent session its own identity. Not a shared developer token, not a long-lived personal access token pulled from someone’s shell. A short-lived credential minted for the task, scoped to one repository or one branch, expiring when the task ends. If an agent session is compromised, the blast radius is one task.

Separate read from write from execute. Many tasks need an agent to read a lot and write a little. Reading the whole monorepo is fine. Pushing to protected branches, editing CI configuration, or touching deployment manifests is not. Treat changes to pipeline definitions, dependency manifests, lockfiles and infrastructure code as a higher trust tier that always needs a human approval, regardless of who or what authored them.

Constrain the sandbox. Agents run commands. Whatever they run should run somewhere with no standing access to production secrets, restricted egress, and a read-only mount of anything it does not need to modify. The question to ask is simple: if this agent were fed a malicious instruction inside a file it reads, what could it reach?

That last question is not hypothetical. The MITRE ATLAS knowledge base, which catalogs adversary techniques against AI systems, now includes techniques aimed specifically at agents, such as AI Agent Tool Invocation (AML.T0053), where an adversary uses their access to an agent to invoke the tools it has been given, and AI Agent Context Poisoning (AML.T0080). An agent that reads untrusted text (an issue, a README in a third-party package, a web page) and holds powerful tools is exactly the shape those techniques target. You do not defend against that with a better code classifier. You defend against it with narrower permissions.

A lifecycle view: where the controls go

It helps to lay the software delivery lifecycle out as a series of trust boundaries and ask what evidence should exist at each one.

Supplemental Explainer

Read it left to right. At the start, the task and its context are inputs, and anything the agent ingests (issues, docs, third-party files) is untrusted data. In the middle, the agent runs under a scoped identity and every tool call is logged. At the pull request, provenance is attached. Review and scanners run as they always have. The build is signed and produces an attestation, and the deploy step checks policy before anything ships. The audit log at the bottom is not a separate system bolted on at the end. It collects evidence from every earlier step.

Nothing here is exotic. Frameworks such as SLSA (Supply-chain Levels for Software Artifacts) already describe build provenance, and tooling like Sigstore and the in-toto attestation format already exist for signing and recording it. The new part is extending the same idea one step upstream, to the authoring step, which used to be invisible because it happened inside a person’s head.

What to record, concretely

The record does not need to be elaborate. A useful minimum for each agent-authored change:

  • The agent and model identifier, including version, so you can tell later which behavior produced a change.
  • The task or ticket that triggered it, and the instruction the agent was given.
  • The repository, branch and credentials scope the session ran with.
  • The list of tools and external resources the agent invoked.
  • The human who reviewed and approved the merge.

Commit trailers, pull request metadata and your CI system’s own logs can carry most of this. You do not need a new platform to start. You need a convention and a check that fails the pipeline when the convention is missing.

Having said that provenance records are only as trustworthy as the system that writes them. If a developer can type any trailer they like, the record is a claim, not evidence. The stronger version has the agent runtime or CI system write the record and sign it, so a developer cannot quietly edit it. Start with the convention, then move to signed records for the repositories that matter most.

Use agents on the defensive side too

There is a counterweight worth taking seriously. The same capability that creates risk can find it. Cloudflare has published an open-source security-audit skill that turns a coding agent into a security auditor, running a multi-phase pipeline of reconnaissance, vulnerability hunting, adversarial validation, reporting, structured output and independent verification, aimed at findings with real, demonstrable impact. It is a single-source signal and early, so I would treat it as an example of a pattern rather than a recommendation to adopt a specific tool.

The pattern is the interesting part: the verification and validation phases. Agents are good at generating hypotheses about vulnerabilities and bad at being trusted about them. A pipeline that makes a second, independent pass try to confirm each finding before reporting it is applying the same discipline we want on authored code. Do not let the agent that wrote the change be the only thing that checks it.

The tradeoff is real. Agent-driven audits cost compute and tokens, they produce findings that still need triage by a person, and they can miss things a specialist would catch. Used as an additional layer on top of existing static analysis and review, they are valuable. Used as a replacement, they create a new single point of failure.

A worked example: the dependency bump

Here is a failure mode I think every team should walk through before they are in it.

A developer asks a coding agent to fix a failing build. The agent notices a version conflict, edits the dependency manifest, bumps a package, and updates the lockfile. The tests pass. The diff looks small. A reviewer, busy and trusting a green check, approves.

What could go wrong? The agent may have chosen a version based on a blog post or a package description it read, which an attacker controls. It may have added a transitive dependency from a typosquatted name. It may have run an install script during its own session, on a machine with access to credentials. None of these show up as “this code looks AI-written”. All of them show up if your policy says that manifest and lockfile changes are a high-trust tier, that agent sessions have no standing credentials, that new dependencies require an allowlist or a review from the owning team, and that the session’s tool calls are logged.

The cost of that control is a slower merge on a small number of files. The cost of skipping it is a compromised build pipeline. That is an easy trade once it is written down, and a hard one to see in the moment.

The review problem

Controls on the pipeline help, but reviewers are still the last human line, and agents change the economics of review. An agent can produce a large, well-formatted, plausible diff in minutes. Reading it carefully takes longer than writing it did. If review capacity does not change, the likely result is rubber-stamping.

Three adjustments help:

  1. Smaller, task-bounded changes. Make the agent’s unit of work something a person can hold in their head. A policy limiting diff size per agent-authored pull request is crude and effective.
  2. Review by risk tier, not by author. Spend human attention on the trust boundaries (authentication, authorization, input handling, secrets, deployment config, dependencies). Let automated checks carry the rest.
  3. Tests the agent did not write. If the same agent writes the code and the tests that bless it, passing tests prove little. Keep a body of independently authored tests, and treat agent-written tests as a suggestion rather than evidence.

Where this approach does not fit

Provenance and scoped permissions cost effort, and they are not equally valuable everywhere. A weekend prototype or an internal script with no sensitive access does not need signed attestations. Teams without basic hygiene, such as branch protection, CI that actually blocks merges, and a working secrets manager, should fix that first, because agent governance sits on top of it. And if your agents run on a hosted service versus your own infrastructure, the controls differ: a hosted agent gives you less visibility into its runtime but may bring built-in guardrails and vendor audit logs, while a self-run agent gives you full control of the sandbox and full responsibility for building it. Either can be made safe. Neither is safe by default.

The executive version

For a leader deciding how much to invest, here is the short framing. Agentic coding raises throughput, and throughput without traceability is how incidents become unexplainable. The investment is modest: identity for agent sessions, a provenance convention enforced in CI, a stricter tier for pipeline and dependency changes, and retained logs. The return is the ability to answer three questions quickly after something goes wrong. Where did this change come from? What was the author allowed to touch? Who approved it?

Analyst commentary points the same way. Gartner’s reporting on the Hype Cycle for Application Security, as relayed by secondary sources, now treats securing MCP connections (tool-definition poisoning, hijacked tool calls and missing access control) as a topic in its own right, and separately predicts that over 40 percent of agentic AI projects will be cancelled by the end of 2027, with weak risk controls among the reasons. I would treat those as directional signals, not precise forecasts, but they match what I see: the projects that survive are the ones that can show their work.

A starting plan

If you want to begin this month, keep it small.

  1. Inventory where agents already run: developer laptops, CI, hosted tools. You will find more than you expect.
  2. Replace long-lived tokens with short-lived, task-scoped credentials for agent sessions.
  3. Add a required provenance field to pull requests, enforced by a CI check.
  4. Mark manifests, lockfiles, CI config and infrastructure code as protected paths requiring named human approval.
  5. Log agent tool calls somewhere you will retain and can query.
  6. Pilot an independent agent-driven audit pass on one repository, and compare its findings against your existing scanners.

None of this requires deciding whether AI-written code is good or bad. It requires deciding that unattributed, over-privileged and unlogged change is not acceptable, whoever or whatever produces it. That has been true for as long as we have shipped software. Agents just make it urgent.

Enjoying this insight?

Join the distribution list to get deep dives on AI transitions and agency economics directly in your inbox. No spam, ever.

Back to Blog

Related Posts

View All Posts »
Governance: The "Human in the Loop" Fallacy

Governance: The "Human in the Loop" Fallacy

Humans cannot keep pace with AI outputs at scale. Here is why enterprise growth relies heavily on Constitutional AI, rather than just throwing more human reviewers at the problem.