BLUE / External Release · Technical Product White Paper

Oak & Sparrow Gatekeeper

Deterministic Execution Control for Enterprise AI Agents

Handling: BLUE / External Release. Approved for disclosure outside Oak & Sparrow Systems Enterprise LLC. Quantitative claims in this paper are tied to named internal test artifacts and are available for inspection under an executed evaluation agreement. The reference-candidate record is not a production certification.

Enterprise AI is moving from conversational recommendation toward autonomous workflow. Agents now invoke tools, query sensitive databases, modify infrastructure, execute financial transactions, and trigger automated security responses. Gatekeeper places a deterministic authority boundary between those probabilistic proposals and the enterprise systems they can affect.

AI models and agent frameworks may propose an action. A deterministic, non-AI authority must decide whether that exact action is permitted before it executes.

1Executive overview

Probabilistic machine-learning systems are strong at reasoning, pattern synthesis, and task decomposition. Those capabilities do not create invariant execution authority. Gatekeeper evaluates versioned policy, identity, requested authority, human approvals, and operational context before an action crosses the commit boundary.

OutcomeAuthority verdictOperational response
GREENAuthorizedExecution may proceed within bounded parameters. A Safety Receipt is produced.
YELLOWHuman review requiredExecution pauses and a review interval is inserted before commit.
REDDeniedExecution is blocked and the fail-closed decision is recorded.

2The enterprise AI execution problem

Traditional security controls were designed around software with predictable code paths. Instruction-tuned language models operate in probabilistic proposal space. When those models can call consequential tools, four recurrent risks become architectural rather than merely conversational.

3The Gatekeeper architectural principle

Inference is proposal. Authority exists only at commit.

Gatekeeper separates the probabilistic reasoning engine, which proposes actions, from the deterministic authority that decides whether the exact proposed transition can execute. The architecture is built around three primitives.

  1. Pre-execution authorization. The exact canonical request is evaluated before the downstream API call, database write, infrastructure change, or other consequential effect.
  2. Rules as code with deontic structure. Policy is versioned and deterministic. Permission is an explicit grant backed by a named rule. Missing authority does not silently become permission.
  3. Least-data minimization. Decisioning can operate on structured metadata, identity, target resources, and policy provenance rather than requiring raw prompts, complete chat histories, or sensitive enterprise payloads.

4Scope boundaries

Gatekeeper isGatekeeper is not
A non-AI Policy Decision Point; a pre-execution authorization gate; a versioned rules-as-code evaluation engine; a decision-to-execution evidence recorder; a cross-system verification authority. An LLM or prompt scanner; an AI gateway or provider router; a Policy Enforcement Point; a replacement for SIEM, SOAR, identity systems, secure browsers, or existing enterprise platforms.

5Architecture and execution lifecycle

Stage 1: pre-execution request

The calling agent or application sends a structured canonical action envelope containing requesting identity, action type, target resource, relevant thresholds, and security context.

Stage 2: deterministic authorization

Gatekeeper evaluates the request against active, versioned policy. A forbidden action, missing credential, or absent required approval produces RED or YELLOW. An authorized action produces GREEN and can be bound to short-lived, single-use capability material.

Stage 3: post-execution validation and evidence

The connected system returns the execution result or resulting state delta. The completed action can be compared with the authorized scope, and the finalized Safety Receipt is sealed into ProofVault.

Latency and availability. Machine-only decision latency is engineered to a target of 250 milliseconds at the ninety-fifth percentile. This is a product target, not a universal measured guarantee. High-consequence operations can be configured to fail closed. Lower-risk exceptions require explicit policy and evidence.

6Deterministic policy and decisioning

Policies are immutable, versioned rule packs in the decision path. GREEN, YELLOW, and RED are the external product verdict vocabulary. Internal harness or integration tokens may differ, but any published mapping must remain explicit and reconcilable against the underlying test record.

YELLOW is not a weaker GREEN. It re-inserts human review before execution. RED is not a warning after the fact. It is a denial before system change.

7Safety Receipts and ProofVault

Each governed authorization request produces a structured Safety Receipt. The record is designed to preserve the relationship among the exact proposed action, active policy, requesting identity, resulting verdict, and execution outcome.

Receipt elementRecorded content
Canonical action digestSHA-256 digest of the exact proposed action parameters.
Policy provenancePolicy version, rule-pack hash, and the rule references applied.
Identity and contextAuthenticated agent, user, tenant, and approver identifiers where applicable.
Deterministic verdictIssued outcome with explicit reason codes.
Execution matchEvidence of whether the resulting state matched the authorized intent.
Chain linkageCryptographic linkage to the prior governed record.

8Deployment and integration models

Gatekeeper can be integrated as a centralized API decision service, a sidecar near agent workloads, a pre-execution step in an orchestration playbook, or an authorization surface connected to an AI gateway or agent framework. The integration pattern can change without moving final authority back inside the model.

9Representative enterprise use cases

ScenarioProposed actionTypical governed outcome
Sensitive data accessBulk export exceeds an approved data threshold without required approval.YELLOW, hold and escalate.
Account changesAgent proposes a privileged role change requiring multi-party consent.YELLOW, hold until approval.
Infrastructure modificationAgent proposes an explicitly forbidden open-internet firewall rule.RED, block before change.
Production deploymentRelease is attempted without a required security gate.RED, block deployment.
Financial actionTransfer exceeds a dual-control threshold.YELLOW, hold for dual authorization.
Automated remediationSecurity automation targets a tier-zero asset requiring human signoff.YELLOW, hold before isolation.

These are representative policy configurations, not claims about specific deployed customer environments.

10How Gatekeeper complements existing systems

AI gateways and provider routers manage model access, routing, quotas, and latency. Identity systems establish who or what is authenticated. SIEM and orchestration platforms collect events and coordinate procedures. Gatekeeper consumes relevant signals from those systems and answers the narrower execution question: is this exact action authorized in this context now?

11Product differentiation

DimensionCommon control patternGatekeeper pattern
Decision basisProbabilistic risk scoring, prompt patterns, or semantic similarity.Deterministic, versioned rules as code.
Execution timingPost-hoc logging or prompt/response filtering.Authorization at the proposal-to-commit boundary.
Human escalationAlert queued after or beside workflow activity.Synchronous YELLOW hold before commit.
EvidenceApplication logs and event traces.Cryptographically linked Safety Receipts with policy and action provenance.
Post-commit verificationOften outside the authorization decision.Resulting-state comparison against authorized scope.

12Validation record and scope boundaries

Oak & Sparrow separates benchmark evidence from field deployment. The figures below are the documented state of the Gatekeeper V2 reference-candidate record represented by External Release 1.0.

BenchmarkRecorded result
Reference-candidate build 357 unit tests across 511 tracked codebase files, with 17,565 deterministic attack and fuzz cases executed.
1,000-request evaluation 99.6 percent local artifact-sealing coverage, 996 sealed receipts, with continuous verified artifact-chain sequencing across the run.
Warrant reference E4 enforcement certification 47 authorizing verdicts produced 47 authorized effects; 0 effects occurred across the remaining 953 non-authorizing and timeout cases; all 18 adversarial and bypass cases failed closed.
Recorded status: Controlled reference candidate GO. Production topology NO-GO. The reference candidate is validated in benchmark harnesses. Production authorization is topology-specific and requires site-specific deployment certification. A benchmark result is not represented as a production result.

Explicit non-claims

Gatekeeper does not claim to guarantee legal compliance, provide nonrepudiation, operate as a blockchain, use zero-knowledge proofs, or replace an existing security platform.

13Evaluation questions for enterprise leaders

  1. Where does final execution authority reside when an autonomous agent prepares to invoke a consequential tool or API?
  2. How does the current architecture prevent a prompt injection from translating directly into an unauthorized infrastructure change or database write?
  3. Does the current stack bind the exact proposed action, policy version, human approval, and resulting system state into one reconstructable record?
  4. How are fail-closed defaults enforced for high-consequence actions while routine authorized actions retain low-latency execution?

14Conclusion

As enterprise AI moves toward autonomous operation, security architecture must extend beyond probabilistic content filtering. Gatekeeper supplies a deterministic execution-control layer so models can propose actions while authority over commit remains governed, verifiable, and capable of failing closed.

The recommended evaluation starts with one consequential workflow: identify the proposal-to-commit boundary, define the policy authority that governs it, bind the exact action to that authority, and test both authorized and denied paths against reconstructable evidence.