PyxGrant / Evidence

Evidence you can rerun.

These are our own tests, run on our own machine, and published with the command that produces them. They are not a customer result and not a third-party audit. Build the binary and run them yourself.

Numbers below are from a full Windows run on 26 September 2026. Reproduce them with the same commands on the binary you have. A later re-run on this machine was blocked by Windows application control.

90 of 90checks pass in the hostile-server walkthrough
26 of 26sections pass, with no planted secret reaching the client or the log
0 of 45attacks missed (allowed). Flagged-but-not-blocked still counts as caught; under_enforced is printed and does not fail the run
133Go packages passed go test on that run, none failed
Hostile-server walkthrough

Every section of pyxgrant demo.

The demo starts a hostile MCP server behind the gateway and runs each control against it in one process. It exits non-zero if any check fails.

SectionWhat it checksPassed
1. Tool poisoningA tool that shadows another and asks for ~/.ssh is hidden from the model, and calling it is refused3 of 3
2. Poisoned tool resultSSN, email, API key, card, and MRN redacted; injected instruction stripped; benign text kept4 of 4
3. Confused deputyHR data can't leave through a public tool; unrelated content can2 of 2
4. Exfiltration in argumentsA markdown-image beacon is refused and an opaque query-string payload redacted2 of 2
5. Corporate IPA confidential roadmap is withheld from the context window1 of 1
6. Memory writeAn injected instruction is neutralised before it is stored1 of 1
7. Human approvalA high-impact call is held, then approved, with an undo recorded2 of 2
8. Rug pullA tool whose description changed after approval is quarantined and removed from the list2 of 2
9. FreezeA request is held during a freeze and released on resume1 of 1
10. Self-protectionWrites to the agent's MCP config and to PyxGrant's policy are refused2 of 2
11. Blast radiusA recursive wipe and a whole-table DELETE are refused before they run2 of 2
12. Semantic data classificationA layoff memo with none of the usual keywords is withheld1 of 1
13. Memory poisoningA poisoned write is refused and fingerprinted; a reworded read-back is refused2 of 2
14. Grant revocationRevoking a grant kills its child, refuses the bound session, and signs a receipt that verifies offline5 of 5
15. Cross-org delegationA partner can narrow and re-delegate a grant, never widen it, with no shared secret6 of 6
16. Agentic browserThe person's own profile, banks, a new site, a cross-origin jump, the host desktop, and a changed page6 of 6
17. SaaS inventoryOver-broad OAuth scopes, unverified publishers, shadow agents, and standing tokens flagged6 of 6
18. Shadow AISanctioned, personal-tenant, shadow, and imitation services told apart; a local model found6 of 6
19. Identity providerStanding credentials, admin service accounts, ownerless identities, and unregistered actors flagged6 of 6
20. Sandbox pre-flightWorkspace escape, unlisted egress, escape primitives, and download-then-run refused6 of 6
21. eDiscoverySearch by person and decision, agent attribution, subject-access export, and legal hold6 of 6
22. PaymentsThe approved cart pays; a swapped cart, a tainted cart, and an amount over the cap are refused4 of 4
23. Industrial controlsReads allowed; acts need an attested link, a rollback plan, and two approvers6 of 6
24. VoiceA spoken injection and a poisoned CRM note are blocked; irreversible actions held4 of 4
25. BenchmarkThe open corpus scores zero misses and zero false alarms1 of 1
26. Audit chainThe hash chain and its Ed25519 seal verify, with no plaintext secret in the log3 of 3
Open benchmark

No misses, on a small corpus we wrote.

pyxgrant benchmark runs a labeled corpus of 45 attacks and 22 benign calls through the real engine. The last completed Windows run, on 26 September 2026, missed none and flagged none of the benign calls. A miss means the attack was allowed. An attack expected to be blocked that was only flagged is counted as caught, recorded as under_enforced, and does not fail Pass. Read that field on the report; we are not guessing it here. With 45 attacks, the 95% range on the miss rate is 0 to 8%. The corpus fingerprint is c8fbe4ba8f6887a6169d074fe1857673, so you can check you ran the same one.

Speed

Measure it on your hardware.

We don't publish a latency figure. pyxgrant perf measures what the gateway adds per call on the machine you run it on, at the median, 95th, and 99th percentile, by payload size, along with sustained throughput and cold-start time. -max-p95-ms turns it into a CI gate.

What these numbers don't show

  • They are our tests on our machine, not an independent audit or a customer deployment.
  • The demo checks that known attack shapes are stopped. It does not measure how often a novel attack gets through.
  • The benchmark corpus is ours and small. Prompt injection is not solved; detection is pattern-based, and new phrasing can pass.
  • The demo's audit key sits beside its log, and PyxGrant warns about it. In production, keep the key elsewhere.

Check it yourself

  • pyxgrant demo reruns the walkthrough and prints a summary.
  • pyxgrant benchmark reruns the corpus and prints its fingerprint.
  • go test ./... runs the test suite.
  • pyxgrant redteam runs the adaptive attack harness, and pyxgrant bypass proves reconciliation names a call that went around the gateway.

Run the engine in your browser →