Q2 2026 Results · 15 May 2026

Same fraud. Different answers.
Caliber tells us which one to ship.

Q2 2026 test: 20-chain battery run on the Nikola Corporation S-1 fraud case. The model that won is the one CevantAI ships in production. The framework is replicable; the verdict isn't ours to argue with.

Model Composite D1 Sharpness D2 Halluc D3 Fidelity D4 Depth D5 Consist D6 Cost Verdict
Claude Opus 4.7 8.53 8.2 8.2 9.6 8.2 7.7 10.0 Production-ready
GPT-5.5 Pro 6.89 5.8 8.3 6.8 5.2 7.4 10.0 Supplementary only
Claude Opus 4.7: “IMMEDIATE_RED, do not proceed.”
GPT-5.5 Pro: “Warrants further scrutiny.” Same fraud. Different answers. Caliber tells us which one to ship.

Source: Caliber Q2 2026 Evaluation · 15 May 2026 · Nikola Corporation S-1 fraud case · 20-chain battery.
Gemini 3.1 Pro excluded from Q2 due to SDK transport reliability; re-inclusion planned for Q3.

The six dimensions

What we measure, and why.

Every Caliber evaluation scores each model across six dimensions chosen for adversarial financial diligence. Scores are 0–10 per dimension; composite is weighted by relevance to institutional DD workflows.

D1 · Sharpness

Decisiveness on flagged risk

How directly does the model commit to a verdict when the evidence supports one? Hedging without justification is penalised.

D2 · Hallucination Resistance

Fidelity to the source corpus

Cross-checks every assertion against source text. Fabricated facts or invented references are heavily penalised.

D3 · Instruction Fidelity

Adherence to the chain spec

Does the model do exactly what the chain asked, in the format requested? Drift from spec compounds across multi-step workflows.

D4 · Analytical Depth

Second-order reasoning

Surfaces implications beyond restating the document. Hunts cross-document contradictions standard review misses.

D5 · Consistency

Repeatability across runs

Same chain, same data, same conclusion. Stochastic variation that flips verdicts is penalised.

D6 · Cost

Production economics

Tokens-per-dollar at production scale. Cheap is necessary but not sufficient. Cost only matters if D1–D5 clear the institutional bar.

The widening gap

Q1 → Q2: the lead grew.

Q1 2026

Claude 4.6 composite 8.73 vs GPT-5.4 7.81. Lead: 0.92.

Q2 2026

Claude 4.7 composite 8.53 vs GPT-5.5 Pro 6.89. Lead: 1.64.

GPT-5.5 Pro dropped a tier (“Supplementary only” from a Q1 production-ready verdict). The frontier-Claude lead on adversarial reasoning is growing, not closing. The institutional implication is one CevantAI builds on: the underlying model choice is not a static answer.

Why this exists

Vendor-neutral. Replicable. Quarterly.

No incumbent ships an institutional benchmark of frontier reasoning models for adversarial finance workflows. CevantAI built one because we had to. Every chain we ship is run through Caliber. The model layer is vendor-swappable; the reasoning library and calibration logic are the IP.

If you are an investor making model-vendor commitments today (or a CTO writing the firm's AI policy), Caliber is the framework that survives the next model release.