Server-side request time
The published P50-P99 figures measure concurrent API work on the production server path while ingestion is active. They exclude a buyer's browser, internet route, DNS and public TLS termination.
Evidence and performance methodology
Jylus publishes model outcome, evidence integrity, token use, compiler timing and sustained ingest as separate measurements with explicit gates.
Model evidence benchmarks
Select a cohort to inspect its complete same-model result. The frozen benchmark and supporting real-data cohort remain separate because their datasets, sample sizes and evaluation contracts are different.
| Evidence path | Strict accuracy | Evidence recall | Hallucination trials | Avg model input | P95 end-to-end | Cost/query |
|---|---|---|---|---|---|---|
| Full context | 78.79% | 86.74% | 7.58% | 205,129 | 9,352.59 ms | US$0.053305 |
| BM25-only RAG | 76.14% | 90.03% | 8.71% | 8,128 | 6,359.19 ms | US$0.003149 |
| Jylus Context Pack | 100.00% | 100.00% | 0.00% | 2,532 | 3,404.00 ms | US$0.001160 |
528 unique adversarial questions across four data domains. Full context, BM25-only RAG and Jylus used the same Gemini 3.1 Flash Lite settings and deterministic scorer.
What changed
Full context mixed current facts with stale values, conflicts, unrelated records, nested-field collisions and untrusted text. BM25 reduced volume but could omit or mis-rank required evidence. Jylus supplied a compact, proof-bound Context Pack around the requested scope. On this defined benchmark that changed accuracy, latency and cost.
Jylus strict-accuracy Wilson 95% interval: 99.28-100.00%. Hallucination-trial interval: 0.00-0.72%. Zero observed trials is not a universal guarantee. Results apply to this defined benchmark, and the evaluation has not been independently reproduced.
Measured economics
Scaled linearly from provider-reported usage of US$0.053305 versus US$0.001160 per query at the frozen standard price. Actual savings vary by model, workload, output length, caching, discounts and usage.
Frozen comparison contract
Public methodology exposes the evaluation contract and limitations required for credibility. Compiler algorithms, ranking and fusion logic, evidence selection, temporal and relationship processing, optimisation and tuning remain proprietary.
Test conditions
Latency scope
The published P50-P99 figures measure concurrent API work on the production server path while ingestion is active. They exclude a buyer's browser, internet route, DNS and public TLS termination.
Public measurements include client-to-server network time, TLS and ingress. They vary by endpoint, payload size, route, cache state and client location, so Jylus reports them independently from server-side request time.
Sustained validation
The passing 30-minute profile published all 1.8 billion events while 18,659,349 API requests ran concurrently. Focused short runs remain useful for regression checks, but they are not presented as sustained-load proof.
Validated sustained ingest, concurrent query traffic, final drain and reliability deltas.
Validate longer-term resource stability, compaction behaviour and recovery margin.
Measured result
Queryability
The gate ran concurrent API requests while events were being accepted, then required the accepted workload to settle without failures, new redeliveries, new dead letters or remaining backlog. The test verifies continued query service and final zero backlog. It does not claim that every individual event was queried immediately after its own publish.
Context compiler
This benchmark measures the deterministic decision-context compiler only. Retrieval and model inference are excluded. Source records are deterministic arbitrary-schema fixtures covering structured JSON records, events, telemetry and document-like payloads across AV, payments, healthcare and industrial scenarios, including conflicting and unverified evidence cases.
Representative payments case: 100 documents, 300 runs, 2,000-token budget. Token counts estimate UTF-8 JSON characters divided by four; model tokenizers vary. The full seven-case suite completed 2,110 runs with zero correctness failures.
Grounding comparison
A read-only AI Sentry architecture query expected four facts: Jylus, MongoDB, source of truth and reindex. The published gate required complete fact coverage with verifiable provenance; internal retrieval and compilation logic is not disclosed.
Prior evaluations
Before the current frozen benchmark, this evaluation used 168 unique questions across seven failure-oriented scenarios and four arbitrary data schemas. Each successful call compared one of three evidence paths without exposing the path label or expected answer.
What this means
Instead of asking a model to reason across the entire evidence universe, Jylus supplied a bounded Context Pack containing the relevant current state, historical evidence, relationships and provenance. On this defined workload, average model input fell from 175,289 to 2,093 tokens while strict accuracy increased from 70.81% to 97.99% compared with full context.
| Evidence path | Strict exact | Passed | Avg input | Hallucination trials |
|---|---|---|---|---|
| Full context | 70.81% | 211/298 | 175,289 | 7.38% |
| BM25 RAG | 75.92% | 227/299 | 7,515 | 6.35% |
| Jylus Context Pack | 97.99% | 293/299 | 2,093 | 1.34% |
| Evidence path | Exact | Passed | Avg input | Avg cost |
|---|---|---|---|---|
| Full context | 40.00% | 52/130 | 154,529 | US$0.116257 |
| BM25 RAG | 48.85% | 64/131 | 7,150 | US$0.005730 |
| Jylus Context Pack | 95.42% | 125/131 | 1,948 | US$0.001839 |
| Evidence path | Exact | Passed | Avg input | Avg cost |
|---|---|---|---|---|
| Full context | 94.64% | 159/168 | 191,353 | US$0.145387 |
| BM25 RAG | 97.02% | 163/168 | 7,799 | US$0.007239 |
| Jylus Context Pack | 100% | 168/168 | 2,206 | US$0.002399 |
GPT-5.4 mini: 98.4% lower average model input cost while strict accuracy increased from 40.00% to 95.42%.
Gemini 3.7 Flash: 98.35% lower average model input cost while strict accuracy increased from 94.64% to 100%.
Defined evaluation only. Costs use the published provider-reported input usage and list rates used for this run.The strict gate required all three expected values, no unsupported structured claim, no invented source proof ID and a current proof citation when required. Across 299 successful Jylus trials, evidence recall was 99.78% and structured hallucination-trial incidence was 1.34%. Gemini completed the full 504-call matrix. GPT completed 392 successful calls from the deterministic seeded order before API credit exhaustion; 111 GPT calls were not run and are not counted as failures. These 896 successful evaluations measure the defined workload, not universal accuracy.
Public methodology describes the evaluation contract, measured inputs and outputs, failure rules and limitations. Context Compiler algorithms, ranking and fusion logic, evidence selection, temporal and relationship processing, optimisation and tuning remain proprietary.