AI Across the SDLC, Part 3: What to Build Before You Give Agents Write Access

Part 2 ended its list of controls with a sentence that deserves its own article: "Write in branch" only means something when the credential itself lacks broader access.

This is that article.

A chatbot talks. An agent acts. Once an agent can push to Git, publish a package, trigger CI, or touch cloud resources, prompt injection stops being a content problem. It becomes an authorization problem. A poisoned README no longer produces a bad answer. It produces a commit.

Most teams still choose agents by model quality: benchmark scores, context window, which one feels smarter in the IDE this month. That question matters less than it looks. A hijacked agent with a narrow token, a policy check it cannot reach, a disposable sandbox, and a trace gives you an incident report. The same agent with a developer's credentials gives you an incident.

Here is what I would build before an agent gets write access, in the order I would build it.

1. Give the agent its own identity, not yours

The easiest way to connect an agent is to hand it whatever the developer already has: a personal access token in an environment variable, the cloud CLI session, the SSH key in ~/.ssh. It works on the first try. That is the problem.

Picture the failure. The agent picks up an issue and reads a linked document. The document contains instructions written for the model, not for a human. The agent now acts with every permission the developer holds: other repositories, package publishing, maybe the production read access granted for on-call. Nothing in that chain was a model failure. The model did what the text in its context told it to do. The credential decided the blast radius.

Agent identity is not developer identity. A human authorizes a task. The task gets its own token, scoped to the actions that task needs and nothing else.

An agent task gets its own short-lived identity with only the scopes the task needs. The token doesn't include production, IAM, secrets, or package publishing.
An agent task gets its own short-lived identity with only the scopes the task needs. The token doesn't include production, IAM, secrets, or package publishing.

What to change. Issue one agent identity per task, not one per developer. Make the tokens short-lived and action-scoped: repository:read, branch:create, pull_request:create, ci:read. Leave production:write, iam:write, secrets:list, and package:publish out of the token entirely, so a hijacked agent cannot use them no matter what it is told. On GitHub, that means a GitHub App with fine-grained permissions instead of a user PAT (its installation tokens expire after one hour). And the agent identity must not be able to approve its own pull request.

2. Put the policy check outside the model

The guardrail most teams start with is a sentence in the system prompt: "Never push to main. Never read .env files."

That is a request, not a control. The rule and the attacker's text sit in the same context window, and the model weighs both.

The MCP specification is blunt about this for its own protocol: MCP "cannot enforce these security principles at the protocol level." Enforcement belongs to whatever executes the action.

So the agent should not execute anything directly. It produces an action request: a tool name plus parameters. A policy decision point outside the model decides whether to run that request.

Agent action flow through a policy decision point, sandbox, verification, human approval, state change, and trace record
Every action request passes a policy decision point outside the model. Denials go to the audit log, allowed actions are verified, and the model never controls its own policy.

Denied requests go to the audit log. Allowed requests run in a sandbox or against an approved API, then pass through verification. Some actions need a named human approver before any state changes. Every path ends in a trace record.

The one arrow that must never exist is the model controlling its own policy. Do not give the model final authority over the policy protecting the action. That includes the softer versions: an agent that can edit the policy file in the repository it works on, or a "self-check" step where the same model decides whether its own request is safe.

What to change. Write the rules as code (Open Policy Agent or an equivalent) and evaluate them before each sensitive action: which tool, which repository, which paths, which parameters. Then count the denied tool calls. A rising number means the policy is too tight for the work, or something is steering the agent.

3. Run the agent somewhere you can throw away

Policy decides what the agent may request. The sandbox decides what a mistake can reach.

Run agent work in a disposable container with a scoped clone of one repository and no production credentials. Limit network egress to the AI gateway, your package registry mirror, and the APIs the policy allows. When the task ends, the container goes away, along with everything the agent installed, cached, or wrote to disk.

This also covers the hallucinated-package problem from Part 1 from a different side. A dependency the model invented can still get installed inside the sandbox. It cannot reach a developer laptop or a long-lived CI runner, and the lockfile check still has to pass before merge.

4. Treat MCP servers as dependencies with attacker-written text

If you need the protocol itself explained, see our guide to the Model Context Protocol. The interesting part here is the trust model.

MCP is no longer one vendor's project. Anthropic donated it to the Linux Foundation's Agentic AI Foundation on December 9, 2025, next to Block's goose and OpenAI's AGENTS.md. Anthropic's announcement counted 10,000 active MCP servers at that point. The current specification revision is dated July 28, 2026.

Every one of those servers ships text that lands in your model's context window: tool names, descriptions, input schemas. The specification itself says tool descriptions "should be considered untrusted, unless obtained from a trusted server."

Invariant Labs showed what that means on April 1, 2025. In their "tool poisoning" proof of concept, a tool that claimed to add two numbers carried hidden instructions in its description, telling the model to read ~/.cursor/mcp.json and ~/.ssh/id_rsa and pass the contents along as a tool parameter. A follow-up a week later used a poisoned server to pull WhatsApp message history out through a different, legitimately connected server. Both are proofs of concept, not incident counts. In both, the user saw a normal-looking tool call.

I'd treat every MCP server as a dependency with attacker-controlled text. Not a plugin. Not a line of config.

What to change.

  • Pin server versions. An update can change a tool description without changing a line of your code.
  • Review tool descriptions the way you review a new dependency: who publishes it, what it asks for, what it can reach.
  • Route MCP tool calls through the same policy decision point as every other action. The client's approval prompt is a user convenience. It is not your control.

5. Let deterministic checks own the build

Part 1 listed non-reproducible CI as a failure mode: a model inside your pipeline produces different output for the same commit on different runs.

Temperature zero does not fix this. Thinking Machines Lab ran one prompt ("Tell me about Richard Feynman") 1,000 times at temperature 0 on a 235B-parameter Qwen3 model and got 80 different completions. Sampling was not the cause. Inference kernels return slightly different numbers depending on how the server batches requests, so your output depends on other people's traffic. With batch-invariant kernels, all 1,000 completions came out identical, and the unoptimized run took about twice as long (55 seconds against 26). That is one model and one prompt, but it is enough to stop treating temperature 0 as a determinism switch.

A gate that passes on Tuesday and fails on Wednesday for the same commit is not a gate.

 CI pipeline where deterministic checks pass or fail the build and the AI reviewer only posts an advisory artifact
Deterministic checks decide whether the build passes. The AI reviewer produces an advisory artifact that a human reads, and it never blocks the build.

Deterministic checks decide pass or fail: compiler, tests, SAST, SCA, secret scanning, policy. Model output becomes an artifact a human reads. I'm much more conservative with execution authority than with advice, and a blocking gate is execution authority. It is the same rule as in remediation: the scanner owns detection, and the model helps with what comes after.

If one specific control really needs a blocking AI gate, make it a deliberate trade. Log the model version, the prompt template version, the retrieved context, and the raw output as evidence. Accept that someone will rerun a build and get a different answer.

A CI reviewer that stays advisory

The security boundary matters more than the prompt. Four properties do the work: the reviewer has no tool access, it executes nothing, its output is advisory, and repository content never becomes a trusted instruction.

The secret scan runs on the diff, not on the working tree. A pull request that removes a hard-coded key still sends that key to the model on a - line, and a scan of the checked-out files would pass. gitleaks stdin scans exactly the text the model will see. (Older examples use gitleaks detect --no-git. The detect command has been deprecated since v8.19.0.)

The report is untrusted too. Injected text in the diff can come back out through the model's summary, straight into a PR comment a human reads. plain() escapes HTML, breaks code fences and markdown links, and neutralizes @-mentions. Findings that point at files outside the diff are dropped.

Schema drift after a model upgrade shows up as its own CI annotation instead of an empty report. And every path returns 0. The reviewer fails open, because deterministic checks own the build.

6. Trace every material action

When a generated change causes an incident, someone has to reconstruct how the change was produced: which model, which prompt version, which context, which tools, who approved. Part 2 listed the audit trail as a control. This is the shape of one record.

{
  "event_id": "a8e6b131-e120-45e6-a468-cbf249245767",
  "workflow": "issue_to_pull_request",
  "workflow_version": "3.4.1",
  "agent_identity": "agent-task-7f3c",
  "actor": "developer@example.org",
  "model": "enterprise-coding-agent",
  "model_version": "2026-08-14",
  "prompt_template_version": "7",
  "repository": "payments-api",
  "commit": "74fa42b662d13e98c4736f114e5003d5684af009",
  "context_artifacts": ["issue-4812", "AGENTS.md", "secure-coding-standard-v5"],
  "tools_requested": ["branch:create", "pull_request:create", "package:publish"],
  "tools_denied": ["package:publish"],
  "human_approval_required": true,
  "approver": "reviewer@example.org",
  "input_tokens": 18273,
  "output_tokens": 2134,
  "latency_ms": 8412
}

For an agent with write access, two fields matter most. tools_denied is how you notice an agent being steered. approver is how you answer "who owned this change?" six months later.

Do not build a parallel logging stack for this. OpenTelemetry's GenAI semantic conventions define spans for model calls, agents, and MCP tool calls. They still carry "Development" status, so attribute names can change, and they recently moved into their own repository. Emit them next to your existing traces anyway. A renamed attribute is cheaper than a second observability system.

Earn the next level

Write access is one rung on a longer ladder.

Eight-level autonomy ladder for AI agents from read and explain to narrow reversible production actions
The autonomy ladder. Write access is level 4. Stop at level 5 until your measurements justify the next level.

Everything above is the price of moving from level 3 to level 4. I would stop at level 5 until the numbers say the next level is earned: denied tool calls, reverted AI changes, review time, false positives.

At every level, answer four questions:

  • What context does the model receive?
  • What action is the model allowed to request?
  • What deterministic system verifies the output?
  • Who owns the result?

The best AI-driven SDLC is not the one with the most AI. It is the one that puts AI on expensive cognition, deterministic systems on verification and enforcement, and accountable people on consequential decisions.


Tags:


Comments:

Please log in to be able add comments.