Enterprise11 min read

Air-Gapped AI Code Generation: A Reference Architecture for Defense Contractors

Defense contractors cannot send source code to a SaaS provider. The reference architecture that runs the full multi-agent pipeline inside a customer's perimeter, with no outbound traffic.

EnsureFix Solutions Team · Solutions Engineers, EnsureFix
Air-Gapped AI Code Generation: A Reference Architecture for Defense Contractors, EnsureFix

The Constraint That Reshapes Everything

For most software teams, "send the code to a SaaS API" is invisible, it just works. For defense contractors, federal agencies, and certain critical-infrastructure operators, that single design choice eliminates 95% of AI coding tooling from consideration. Source code under ITAR controls, in environments under FedRAMP High, or running on classified networks cannot leave the perimeter. Period.

This post is the reference architecture we deploy for those customers. It assumes a fully air-gapped environment: no outbound internet traffic, no external API calls, no SaaS dependencies. Everything (model serving, pipeline orchestration, audit storage), runs inside the customer's boundary.

The Components

Five components, all customer-hosted:

  • Model serving. A locally-hosted inference stack. For most defense customers, this is an on-prem GPU cluster running an enterprise-licensed model. The model weights are delivered on physical media or via cleared download channels and live entirely within the perimeter.
  • Pipeline orchestrator. The multi-agent coordinator, planner, coder, reviewer, security, test, root-cause. Containerized, deployed via the customer's standard orchestration stack (typically OpenShift or air-gapped Kubernetes).
  • Source integration. Connectors to on-prem GitLab, Bitbucket Server, or Azure DevOps Server. No cloud SaaS connectors.
  • Audit and observability. Logs and traces ship to the customer's existing SIEM. No external telemetry endpoints.
  • Update channel. A cleared, manual update process for the orchestrator and model. Typically quarterly, with each update reviewed by the customer's security team before deployment.

The full deployment fits on a 3-node cluster for small teams, scaling to 20+ nodes for the largest deployments.

What You Lose, And How To Compensate

Three things you cannot have in this architecture:

A frontier proprietary model accessed over the public internet. You are running an on-prem-eligible model. The capability gap with public frontier models is roughly 6-9 months. For ticket-to-PR workflows, that gap is not material. For research-grade reasoning, it can be.

SaaS-side observability and dashboards. You operate your own. The orchestrator emits OpenTelemetry traces and Prometheus metrics that plug into existing customer infrastructure.

Vendor-pushed updates. Quarterly cleared updates instead of weekly SaaS pushes. The compensation is operational discipline: tested, audited releases that move slower but never surprise.

The Network Architecture

The deployment sits in three zones:

  • Code zone. GitLab/Bitbucket Server, the pipeline orchestrator, and the model serving. All in the same trust boundary, with mTLS between components.
  • Audit zone. Long-term audit storage, SIEM integration, and the compliance reporting tier.
  • Operations zone. Update staging and operator workstations.

No traffic crosses the boundary outbound. Inbound traffic is restricted to authenticated operators and code-zone clients (developer workstations or CI runners).

The Compliance Story

For FedRAMP High and DoD IL5/IL6 customers, the architecture supports:

  • Boundary discipline. All data stays in-boundary. No SaaS data flows.
  • Audit trails. Every AI action (plan, change, validation, decision), is logged with timestamps, agent identity, and human approval state. Audit logs are immutable and retained per the customer's retention policy.
  • Identity integration. Operators authenticate through the customer's existing IdP (typically Okta or Active Directory Federation Services). No vendor-managed identity.
  • Cryptographic posture. FIPS 140-2 validated cryptographic modules throughout. TLS 1.3 with approved cipher suites only.
  • Vulnerability management. Quarterly CVE assessment on the deployed software stack, tied to the customer's STIG compliance process.

The Audit Trail Specifically

This is the question every accreditor asks first. The architecture produces, per AI-generated change:

  • The triggering event (ticket creation, label assignment, scheduled scan).
  • The planner's decomposition, with reasoning.
  • The coder agent's diff and the inputs it received.
  • The reviewer agent's findings and recommendations.
  • The security agent's findings.
  • The test agent's generated tests and results.
  • The decision engine's confidence score and routing decision.
  • The human approver, if any, with timestamp and IP.
  • The merge event and the eventual production deployment, linked to the PR.

This is more audit detail than most human-generated changes carry. For regulated environments, this is the selling point. For details on the audit framework, see SOC 2 compliance for AI code generation.

The Performance Profile

For a 3-node cluster (modest deployment, ~50 engineers):

  • Ticket parse to first plan: 8-15 seconds.
  • Plan to first diff: 45-90 seconds.
  • Full pipeline including validation: 3-6 minutes per ticket.
  • Concurrent tickets supported: 8-12.

This is slower than a SaaS deployment with elastic capacity, but comfortably faster than the human review and merge cycle. The bottleneck is human review, not the pipeline.

What Does Not Work In This Environment

Three patterns that do not translate to air-gapped:

Vector retrieval over a SaaS-hosted embedding service. You need on-prem embeddings. The trade-off is acceptable embedding quality with smaller, cheaper-to-host embedding models.

Public-internet documentation lookups. The agent cannot Google a Stack Overflow answer. It works from the customer's internal documentation, code comments, and learned context. Onboarding documentation matters more in this environment.

Telemetry to the vendor. No outbound telemetry. The vendor cannot see what is failing. Diagnostic data is collected by customer ops and shared through cleared channels if a support case requires it.

What Surprised Us

Two things that emerged from deploying this:

The on-prem model is "good enough" sooner than expected. For ticket-to-PR work, the local model is competitive after the first month of per-repo learning calibration. The frontier model gap matters less when the task is well-scoped.

The audit trail is a productivity feature, not just compliance. Engineers who can see exactly what the AI did, why, and on what input file the issue at much faster pace. The audit detail produced for the accreditor is also the debugging information operators need.

Who Should Consider This

Three signals you are in the air-gapped tier:

  • Your code is subject to export controls (ITAR, EAR).
  • Your environment is FedRAMP High, DoD IL5/IL6, or classified.
  • Your security team will not approve sending source to a SaaS endpoint, regardless of contract language.

If any apply, the SaaS tier is not an option for you. The air-gapped tier is. For the broader compliance scaffolding, see AI coding in regulated industries. For the self-hosted buyer's guide more broadly, see self-hosted AI coding agent buyer's guide.

Frequently asked questions

Can defense contractors use AI code generation with ITAR-controlled source code?

Yes, but only with an air-gapped deployment where the entire pipeline runs inside the customer's perimeter. Source code under ITAR or EAR export controls cannot be sent to a SaaS API, so model serving, orchestration, and audit storage all live in-boundary with no outbound traffic. This is the only architecture that satisfies security teams who will not approve sending source to an external endpoint regardless of contract language.

How does air-gapped AI code generation work without internet access?

Five customer-hosted components handle everything locally: an on-prem GPU cluster serves an enterprise-licensed model, a containerized orchestrator runs the multi-agent pipeline, connectors integrate with on-prem GitLab or Azure DevOps Server, logs ship to the customer's existing SIEM, and updates arrive through a cleared quarterly process. The full deployment fits on a 3-node cluster for small teams and scales to 20+ nodes. Model weights are delivered on physical media or through cleared download channels.

How much slower is on-prem AI code generation than SaaS?

On a 3-node cluster serving roughly 50 engineers, ticket parse to first plan runs 8-15 seconds, plan to first diff 45-90 seconds, and a full validated pipeline 3-6 minutes per ticket. That is slower than an elastic SaaS deployment but comfortably faster than the human review and merge cycle. In practice the bottleneck is human review, not the pipeline.

Is air-gapped AI coding compliant with FedRAMP High and DoD IL5/IL6?

The architecture is built for it: all data stays in-boundary, every AI action is logged immutably with timestamps and human-approval state, operators authenticate through the customer's own IdP, and cryptography uses FIPS 140-2 validated modules with TLS 1.3. Vulnerability management ties to the customer's STIG process with quarterly CVE assessments. For the broader scaffolding, see AI coding agents in regulated industries and the SOC 2 compliance checklist.

What gets lost in an air-gapped AI coding deployment and how do you compensate?

You give up a public frontier model, SaaS-side dashboards, and vendor-pushed updates. You compensate with an on-prem-eligible model that closes the gap for scoped work after per-repo calibration, your own OpenTelemetry and Prometheus observability, and tested quarterly releases that move slower but never surprise. If you are weighing the broader self-hosted decision, see the self-hosted AI coding agent buyer's guide.

EnsureFix Solutions Team

Solutions Engineers, EnsureFix

The EnsureFix solutions team helps engineering leaders evaluate, pilot, and roll out autonomous coding agents, drawing on real deployment data across customer teams.

air-gappeddefenseself-hostedFedRAMPITARenterprise security

From reading to running

Ready to automate your tickets?

Watch EnsureFix take a real item from your backlog all the way to a pull request.