Evidence and performance methodology

Measure the answer, not just retrieval.

Jylus publishes model outcome, evidence integrity, token use, compiler timing and sustained ingest as separate measurements with explicit gates.

100% strict accuracy / 528 trials98.77% less model input97.82% lower measured cost/query1M+ events/s for 30 min

Model evidence benchmarks

Compare Jylus with standard and tuned RAG.

Select a cohort to inspect its complete same-model result. The frozen benchmark and supporting real-data cohort remain separate because their datasets, sample sizes and evaluation contracts are different.

528unique adversarial questions
1,584completed model evaluations
100%strict accuracy across frozen Jylus trials
98.77%less model input than full context
FROZEN BENCHMARK528 questions
Evidence pathStrict accuracyEvidence recallHallucination trialsAvg model inputP95 end-to-endCost/query
Full context78.79%86.74%7.58%205,1299,352.59 msUS$0.053305
BM25-only RAG76.14%90.03%8.71%8,1286,359.19 msUS$0.003149
Jylus Context Pack100.00%100.00%0.00%2,5323,404.00 msUS$0.001160

528 unique adversarial questions across four data domains. Full context, BM25-only RAG and Jylus used the same Gemini 3.1 Flash Lite settings and deterministic scorer.

What changed

The model did not improve. The evidence supplied to it improved.

Full context mixed current facts with stale values, conflicts, unrelated records, nested-field collisions and untrusted text. BM25 reduced volume but could omit or mis-rank required evidence. Jylus supplied a compact, proof-bound Context Pack around the requested scope. On this defined benchmark that changed accuracy, latency and cost.

+21.21 pointsstrict accuracy vs full context
97.82%lower measured model cost/query
98.77%less model input than full context
0%scored hallucination trials in 528 Jylus evaluations

Jylus strict-accuracy Wilson 95% interval: 99.28-100.00%. Hallucination-trial interval: 0.00-0.72%. Zero observed trials is not a universal guarantee. Results apply to this defined benchmark, and the evaluation has not been independently reproduced.

Measured economics

One million requests on this benchmark workload

Full context
US$53,305
Jylus Context Pack
US$1,160
Measured difference
US$52,145

Scaled linearly from provider-reported usage of US$0.053305 versus US$0.001160 per query at the frozen standard price. Actual savings vary by model, workload, output length, caching, discounts and usage.

Frozen comparison contract

No tuning between accepted arms

  • 528 unique designs across four data domains and 11 adversarial classes
  • Zero missing pairs, duplicate cells, execution errors or malformed response hashes
  • Raw answers, provider request IDs, timings, hashes and checkpoints retained
  • No core, portal, prompt, scorer, dataset, budget or model-setting changes between arms

Public methodology exposes the evaluation contract and limitations required for credibility. Compiler algorithms, ranking and fusion logic, evidence selection, temporal and relationship processing, optimisation and tuning remain proprietary.

Test conditions

Reproducible workload definition

Host
AMD EPYC 4464P, 12 cores / 24 threads
Memory
128 GB DDR5 ECC
Storage
2 x 960 GB NVMe, software RAID 1
Network
1 Gbps host interface
Event shape
Structured device telemetry with nested identity, state and measurements
Representative raw size
1,342 bytes per JSON envelope; values vary by sequence and payload
Load generation
Isolated server-side generator using a fixed, declared test profile
Working set
7,000 synthetic device identities in a deterministic workload
Duration
30 minutes at the 1M-events/s profile; 1.8 billion published events
Reliability gate
Accepted-event accounting, zero-failure checks and verified final zero backlog

Latency scope

Server-side and public network time are separate measurements.

REFERENCE PROFILE

Server-side request time

The published P50-P99 figures measure concurrent API work on the production server path while ingestion is active. They exclude a buyer's browser, internet route, DNS and public TLS termination.

PUBLIC HTTPS

Customer-perceived request time

Public measurements include client-to-server network time, TLS and ingress. They vary by endpoint, payload size, route, cache state and client location, so Jylus reports them independently from server-side request time.

Sustained validation

Sustained and focused runs stay separate.

The passing 30-minute profile published all 1.8 billion events while 18,659,349 API requests ran concurrently. Focused short runs remain useful for regression checks, but they are not presented as sustained-load proof.

Passed and published

30 minutes

Validated sustained ingest, concurrent query traffic, final drain and reliability deltas.

Memory
Start, peak and end
Queues
Peak and final depth
Latency
P50, P90, P95 and P99
Future validation profile

2 hours

Validate longer-term resource stability, compaction behaviour and recovery margin.

Memory
Start, peak and end
Queues
Peak and final depth
Latency
P50, P90, P95 and P99

Measured result

Latency and reliability gates

Accepted rate1,000,135.7 events/s
Published events1,800,000,000
Concurrent API requests18,659,349
P505.44 ms
P9010.78 ms
P9513.30 ms
P9920.99 ms
Failures0
Redeliveries0 added by the run
Dead letters0 added by the run
Pending queue0 after verified drain

Queryability

Reads remained active during ingest.

The gate ran concurrent API requests while events were being accepted, then required the accepted workload to settle without failures, new redeliveries, new dead letters or remaining backlog. The test verifies continued query service and final zero backlog. It does not claim that every individual event was queried immediately after its own publish.

Context compiler

Reduction is published with correctness gates.

This benchmark measures the deterministic decision-context compiler only. Retrieval and model inference are excluded. Source records are deterministic arbitrary-schema fixtures covering structured JSON records, events, telemetry and document-like payloads across AV, payments, healthcare and industrial scenarios, including conflicting and unverified evidence cases.

Source tokens74,414
Final packet1,892
Reduction97.46%
Compiler P9521.04 ms

Representative payments case: 100 documents, 300 runs, 2,000-token budget. Token counts estimate UTF-8 JSON characters divided by four; model tokenizers vary. The full seven-case suite completed 2,110 runs with zero correctness failures.

✓Expected decision readiness returned✓Expected latest values retained✓Zero irrelevant entities admitted✓Zero secret-marker leakage✓Every retained fact carries proof IDs✓Final packet remains within the 2,000-token budget✓One context identity across shuffled input order

Grounding comparison

Evidence availability, not model accuracy.

A read-only AI Sentry architecture query expected four facts: Jylus, MongoDB, source of truth and reindex. The published gate required complete fact coverage with verifiable provenance; internal retrieval and compilation logic is not disclosed.

Without retrieved evidence
0%
With raw Jylus records
100%
With compiled Jylus evidence
100%
This is a deterministic grounding-readiness score for one defined four-fact gate. It is not an LLM answer-quality percentage. No model A/B result is published from this run.

Prior evaluations

Earlier multi-model evidence comparison.

Before the current frozen benchmark, this evaluation used 168 unique questions across seven failure-oriented scenarios and four arbitrary data schemas. Each successful call compared one of three evidence paths without exposing the path label or expected answer.

293/299strict Jylus answers across both models
95.69-99.08%combined Wilson 95% interval
98.81%fewer input tokens than full context

What this means

Jylus did not change the models. It changed the evidence they received.

Instead of asking a model to reason across the entire evidence universe, Jylus supplied a bounded Context Pack containing the relevant current state, historical evidence, relationships and provenance. On this defined workload, average model input fell from 175,289 to 2,093 tokens while strict accuracy increased from 70.81% to 97.99% compared with full context.

Combined successful evaluations

Evidence pathStrict exactPassedAvg inputHallucination trials
Full context70.81%211/298175,2897.38%
BM25 RAG75.92%227/2997,5156.35%
Jylus Context Pack97.99%293/2992,0931.34%

OpenAI GPT-5.4 mini

Evidence pathExactPassedAvg inputAvg cost
Full context40.00%52/130154,529US$0.116257
BM25 RAG48.85%64/1317,150US$0.005730
Jylus Context Pack95.42%125/1311,948US$0.001839

Gemini 3.7 Flash

Evidence pathExactPassedAvg inputAvg cost
Full context94.64%159/168191,353US$0.145387
BM25 RAG97.02%163/1687,799US$0.007239
Jylus Context Pack100%168/1682,206US$0.002399

GPT-5.4 mini: 98.4% lower average model input cost while strict accuracy increased from 40.00% to 95.42%.

Gemini 3.7 Flash: 98.35% lower average model input cost while strict accuracy increased from 94.64% to 100%.

Defined evaluation only. Costs use the published provider-reported input usage and list rates used for this run.

The strict gate required all three expected values, no unsupported structured claim, no invented source proof ID and a current proof citation when required. Across 299 successful Jylus trials, evidence recall was 99.78% and structured hallucination-trial incidence was 1.34%. Gemini completed the full 504-call matrix. GPT completed 392 successful calls from the deterministic seeded order before API credit exhaustion; 111 GPT calls were not run and are not counted as failures. These 896 successful evaluations measure the defined workload, not universal accuracy.

Public methodology describes the evaluation contract, measured inputs and outputs, failure rules and limitations. Context Compiler algorithms, ranking and fusion logic, evidence selection, temporal and relationship processing, optimisation and tuning remain proprietary.