Assurance
Evidence over assertion.
Tajalli is built around deterministic replay, frozen decision rules, controlled validation and explicit claim boundaries.
01Methodology
How a result earns its place on this page.
- Calibration is not validation
- Rules and thresholds are set on calibration data. Results are reported only on data that played no part in setting them.
- Fresh and holdout evaluation
- Validation uses unseen periods, assets or cases. Where a holdout exists, it is evaluated once, after rules are frozen.
- Frozen thresholds
- Decision rules are versioned and frozen before evaluation. A result is attributed to a specific policy version.
- No outcome-driven rescue
- Rules are not adjusted after seeing evaluation outcomes in order to improve them. A changed rule is a new version requiring new evaluation.
- Historical failures preserved
- Failed versions and their results are retained. Progress is measured against them, not by replacing them.
- Deterministic replay
- Every evaluated decision is re-executed from recorded state and policy. Replay must reproduce the original decision exactly.
- Shadow before control
- Live deployments begin in shadow mode. Decisions are computed and compared but not enforced until evidence supports it.
- Evidence lineage and versioned policy
- Each decision references the evidence it consumed and the policy version that governed it.
02Status
Validated, designed and under evaluation — kept separate.
Validated means demonstrated in controlled evaluation. Designed means specified and implemented in architecture, not yet independently validated. Under evaluation means in active study.
| Capability | Scope | Status |
|---|---|---|
| Deterministic certification and replay | QISTAS | Validated |
| Boundary-triggered withdrawal and recertification | QISTAS | Validated |
| Conservative portfolio certification | QISTAS | Validated |
| Live commitment gating in customer operations | QISTAS | Under evaluation |
| Integration adapters (REST, streams, OCPP / OCPI) | QISTAS | Designed |
| Deterministic action authorization | RADM | Validated |
| Provenance-bound tool mediation | RADM | Designed |
| Execution broker, staged commit / abort, receipts | RADM | Designed |
| Multi-tenant enterprise deployment | RADM | Designed |
| Industrial / OT control applications | Architecture | Under evaluation |
03Selected evidence
Results, with their scope attached.
QISTAS
Fresh unseen validation
Distributed energy flexibility. Frozen rules evaluated on data excluded from calibration.
- external publication compression
- 98.7248%
- decision-relevant status transitions captured
- 23 / 23
- ADMISSIBLE → INFEASIBLE transitions captured
- 17 / 17
- downward overstatement events
- 0
- stale publications
- 0
- deterministic replay
- 26 / 26
Results shown are from controlled public-data validation and shadow evaluation. They are not customer-production performance claims.
QISTAS
Portfolio validation
Conservative portfolio certification composed from asset-level certificates.
- of internally admissible capacity-time preserved by the conservative certification layer
- 95.61%
- portfolio overstatement events
- 0
- deterministic portfolio replay in the evaluated dataset
- 100%
Results shown are from controlled public-data validation and shadow evaluation. They are not customer-production performance claims.
RADM
Controlled authorization benchmarks
A ground-truth-contract configuration of the public AgentDojo benchmark, and a separate holdout set of authorization cases.
- safe cases in the evaluated ground-truth-contract AgentDojo configuration
- 949 / 949
- successful evaluated attacks in that configuration
- 0 / 949
- holdout authorization cases in a separate controlled evaluation
- 484
- recall in that evaluated holdout
- 100%
- false-positive rate in that evaluated holdout
- ~0.495%
Benchmark results apply to the stated evaluated contracts and configurations. They do not constitute universal security completeness.
04Claim discipline
What we will not say.
- We do not describe controlled benchmark performance as universal proof.
- We do not present shadow economic results as customer savings.
- We do not treat missing evidence as permission.
Enterprise enquiries
Evaluate it on your own evidence.
The most credible validation is the one run on your data. We structure evaluations as shadow studies with pre-agreed metrics.