Egress is an action boundary: what our controlled study found.
Exact-action grants contained the observed scope violations that destination filtering missed. The same study exposes where revocation, single use and completion guarantees end.
Exact-action grants contained the observed scope violations that destination filtering missed. The same study exposes where revocation, single use and completion guarantees end.

An allowed network destination does not establish permission to send every payload, write every object or act for another task. In our incident-inspired controlled study, exact-action grants contained observed scope violations that destination filtering missed. A conventional grant matched OSuite's result. Neither result eliminated late-revocation effects, cross-store reuse or work lost after admission.
Public OpenAI reports about DNS-mediated access, unauthorized repository writes and publication side effects motivated the scenarios. We did not reproduce OpenAI's private environment, inspect its incident logs or establish that OSuite would have prevented those incidents.
The tasks were synthetic, but execution used real sockets, processes, signatures and PostgreSQL on an isolated Azure VM. An unprivileged actor could reach a mandatory private broker, which could reach the private receivers. All comparison modes shared this outer safety perimeter; none sent payloads to a third-party exfiltration destination.
OSuite's reference integration called actual Studio guard and repository code against a disposable lab database, not the public Studio authentication stack. The broker was trusted. CAVA bound the request's action identity to approval evidence, and ZeroGate checked admission. These results depend on enforced mediation, not an agent voluntarily respecting a prompt.
Here, v3 identifies the egress study revision, not a product version. Earlier v1/v2 work provides development history rather than extra trials for the primary result. V1 also showed that network cutoff plus process termination stopped observed post-stop effects. V2 exposed a stale-pass failure retained below.
The primary protocol covered 14 mechanism families and 39 episode definitions, tested in six modes with ten repetitions: 2,340 episode records and 3,000 request attempts. Some episodes contained multiple attempts. These are repeated authored scenarios, not 3,000 independent attacks or samples of real-world attack prevalence.
Each mode had 500 attempts, 210 planned authorized effects and 10 intentionally unknown receiver-capture episodes. Launch-only approval was an ablation with no inner action check. The conventional comparator implemented exact task, principal and request binding with durable one-use consumption; it was not an installation of a commercial competitor.
| Configuration | Observed unauthorized deliveries | Authorized effects completed | Unknown episodes |
|---|---|---|---|
| Transport only | 280 | 210/210 | 10 |
| Destination ACL | 200 | 140/210 | 10 |
| Launch-only approval ablation | 280 | 210/210 | 10 |
| Conventional exact-action grant | 0 | 210/210 | 10 |
| OSuite slow / online readback | 0 | 210/210 | 10 |
| OSuite fast / prepared pass plus live authority | 0 | 210/210 | 10 |
An unauthorized delivery is an observed request outside the frozen grant. It includes disclosure to a receiver that rejects the requested mutation, not necessarily successful theft. Zero means no unauthorized delivery among observed, fully captured cases. Missing capture remains unknown, not safe: the whole-suite evidence-valid flag stays false and strong-mode security scores remain null.
The conventional grant matched OSuite's containment and utility. OSuite's contribution exercised here is the connection between approval, canonical action, durable admission and replayable evidence, not exclusive ability to enforce exact authorization.
Routing was not perfect either. Each strong mode matched 450/500 expectations; 40 differences were review-versus-deny outcomes and ten involved missing capture. Designed caller timeouts were not counted as blocks when receiver evidence showed delivery.
The fast path retains a prepared signed Action Pass but performs a fresh authority read before admission. It does not treat a valid signature as proof that an approval remains current, and it has no positive-cache or offline fallback in the tested integration.
These timings cover 220 admitted attempts per strong mode, including ten unknown-capture cases.
| Matched v3 path | Gate p50 | Gate p95 | Broker p50 | Broker p95 |
|---|---|---|---|---|
| Conventional exact-action grant | 4.740 ms | 11.375 ms | 47.124 ms | 195.402 ms |
| OSuite slow | 17.917 ms | 39.640 ms | 60.020 ms | 208.723 ms |
| OSuite fast with live authority | 12.044 ms | 21.802 ms | 53.726 ms | 198.666 ms |
Against the matched OSuite slow path, fast-path gate p50 was 32.8% lower, gate p95 was 45.0% lower and broker p50 was 10.5% lower. The conventional grant remained faster. The older local-only pass timing is not the timing of this repaired configuration.
Gate timing ends before dispatch and excludes preparation, model inference, human approval, reviewer UI and public WAN. Broker timing includes dispatch and response, including deliberately delayed responses. These are neither customer end-to-end SLAs nor throughput measurements.
In the separate regression, the controller recorded remote revocation and a separate readback before the next authority check, while local invalidation was initially withheld. The historical fast path produced receiver effects in 5/5 trials; the repaired fast path produced 0/5. Both slow-path versions produced 0/5. Five trials demonstrate a bounded regression result, not a zero population failure rate or independent timestamp attestation.
That repair covers revocation before the fresh authority read. It does not make admission and a later remote effect atomic with revocation. A separate local process-fault study repeated seven scenarios three times, yielding 21 episodes and 39 worker attempts. It retained these counterexamples:
These are execution limits, not protection successes. Recovery needs executor-side idempotency and reconciliation; automatically issuing another pass after a timeout does not establish exactly-once execution. The injected process crashes also do not establish power-loss durability.
The separate v3 model cohort ran 24 episodes: four synthetic tasks, three modes and two repeats. Each mode completed 8/8 episodes and 10/10 expected effects. There were no unauthorized model requests, including in the misleading retrieved-source task. This demonstrates task utility, not attack interception or defeat of an adversarial model.
The egress v3 evidence package is the entry point for offline captured-evidence recomputation. Follow its instructions and retain the exact revision used. It is not an independent rerun of the cloud experiment or independent certification.
Recomputing records, counts and signatures checks consistency under retained laboratory keys. It does not attest that the operator captured every event honestly, independently validate the scenario labels or establish production safety. Earlier egress studies and the local process-fault follow-up retain separate protocols and denominators; their results must not be silently pooled into this matrix.
The practical conclusion is specific: destination controls and exact-action authorization answer different questions. A useful evaluation should measure both authorized work and observed violations, then keep unknown evidence, stale authority and incomplete execution visible alongside the successful cases.
Request enterprise access and send your first governed decision today.