We don't choose your model.
We built the framework that does.
Vendor-neutral evaluation. Quarterly. Six adversarial DD dimensions. Every frontier model. Publicly replicable methodology, built to be the institutional standard for AI model selection. The Moody's of frontier model evaluation in financial services.
Same fraud. Different answers.
Caliber tells us which one to ship.
Q2 2026 test: 20-chain battery run on the Nikola Corporation S-1 fraud case. The model that won is the one CevantAI ships in production. The framework is replicable; the verdict isn't ours to argue with.
| Model | Composite | D1 Sharpness | D2 Halluc | D3 Fidelity | D4 Depth | D5 Consist | D6 Cost | Verdict |
|---|---|---|---|---|---|---|---|---|
| Claude Opus 4.7 | 8.53 | 8.2 | 8.2 | 9.6 | 8.2 | 7.7 | 10.0 | Production-ready |
| GPT-5.5 Pro | 6.89 | 5.8 | 8.3 | 6.8 | 5.2 | 7.4 | 10.0 | Supplementary only |
GPT-5.5 Pro: “Warrants further scrutiny.” Same fraud. Different answers. Caliber tells us which one to ship.
Source: Caliber Q2 2026 Evaluation · 15 May 2026 · Nikola Corporation S-1 fraud case · 20-chain battery.
Gemini 3.1 Pro excluded from Q2 due to SDK transport reliability; re-inclusion planned for Q3.
What we measure, and why.
Every Caliber evaluation scores each model across six dimensions chosen for adversarial financial diligence. Scores are 0–10 per dimension; composite is weighted by relevance to institutional DD workflows.
Decisiveness on flagged risk
How directly does the model commit to a verdict when the evidence supports one? Hedging without justification is penalised.
Fidelity to the source corpus
Cross-checks every assertion against source text. Fabricated facts or invented references are heavily penalised.
Adherence to the chain spec
Does the model do exactly what the chain asked, in the format requested? Drift from spec compounds across multi-step workflows.
Second-order reasoning
Surfaces implications beyond restating the document. Hunts cross-document contradictions standard review misses.
Repeatability across runs
Same chain, same data, same conclusion. Stochastic variation that flips verdicts is penalised.
Production economics
Tokens-per-dollar at production scale. Cheap is necessary but not sufficient. Cost only matters if D1–D5 clear the institutional bar.
Q1 → Q2: the lead grew.
Claude 4.6 composite 8.73 vs GPT-5.4 7.81. Lead: 0.92.
Claude 4.7 composite 8.53 vs GPT-5.5 Pro 6.89. Lead: 1.64.
GPT-5.5 Pro dropped a tier (“Supplementary only” from a Q1 production-ready verdict). The frontier-Claude lead on adversarial reasoning is growing, not closing. The institutional implication is one CevantAI builds on: the underlying model choice is not a static answer.
Vendor-neutral. Replicable. Quarterly.
No incumbent ships an institutional benchmark of frontier reasoning models for adversarial finance workflows. CevantAI built one because we had to. Every chain we ship is run through Caliber. The model layer is vendor-swappable; the reasoning library and calibration logic are the IP.
If you are an investor making model-vendor commitments today (or a CTO writing the firm's AI policy), Caliber is the framework that survives the next model release.