● LIVE · UPDATED IN REAL TIME

Evaluation Readiness

Loading the evaluated outcome cohort. No performance conclusion is shown until at least one prediction has been evaluated.

Prediction Stream

ORION's own ledger, as it's actually written — every row here is a real, persisted prediction. Nothing on this page is staged for display.

LEDGER · LIVE
Loading…
Loading live accuracy data…
Validation status loading…
Type: · Methodology: · Data generated:
Ledger rows retained: · eligible cohort: · excluded: · reconciliation: · duplicates: · orphans: · invalid state: .
Excluded from this public model cohort: client-supplied, unverified-source, and incomplete point-in-time provenance rows.

Overall

Persisted and evaluated cohort counts.

Platform-Wide Predictions
Evaluated
Pending
Ungradable Legacy
Overall Accuracy

By Recommendation Type

Buy, Sell/Avoid, and Hold calls are graded against different thresholds — never blended into one number.

Buy Accuracy
Sell / Avoid Accuracy
Hold Accuracy

Outcome-Based Accuracy

A second, independent lens — graded against a real subsequent FDA event (approval/CRL), never blended with the price-based numbers above.

Evaluated
Accuracy

Trial-Status-Based Accuracy

A third, independent lens — graded against ClinicalTrials.gov: did a trial that was still running when the prediction was made since stop for a bad reason (terminated/withdrawn). Never blended with the price- or FDA-event-based numbers above.

Evaluated
Accuracy

By Event Type

Only predictions matched to a verified FDA Calendar catalyst carry an event type — everything else groups under "Unknown," it's never guessed.

No gradeable predictions yet

By Therapeutic Area

Same taxonomy source as Event Type above — indication is only known when a real FDA Calendar catalyst was matched at prediction time.

No gradeable predictions yet

Confidence Calibration

Completed, gradeable predictions only. This calibrates recommendation confidence against recommendation correctness; it is not an FDA approval-probability calibration. A lower Brier score and calibration error indicate closer agreement.

Calibration Sample
Brier Score
Expected Calibration Error
No gradeable calibration sample yet.

Best & Worst Calls

We show the miss, not just the win — hiding it would defeat the point of a public track record.

✓ Best Call
Not enough evaluated predictions yet.
✕ Worst Call
Not enough evaluated predictions yet.

Methodology

Every included prediction is stored before evaluation and a rating is not edited after the fact. Full historical source snapshots are not yet complete, so this page is labeled a prospective observational ledger rather than a point-in-time reproducible backtest. Price-based grading: a Buy-type rating is graded correct if the stock later moved ≥+10%, partially_correct between 0–10%, else incorrect (mirrored for Sell/Avoid; Hold is correct within a ±10% band). Predictions whose reference price was never captured are labeled ungradable_legacy and excluded from return-based accuracy; no current or estimated price is substituted. Outcome-based grading: a rating is separately checked against any real FDA event that occurred after the prediction was made — approval-type events validate Buy-type ratings, rejection-type events validate Sell/Avoid ratings; ambiguous or Hold cases are marked inconclusive rather than forced. Trial-status-based grading: a third, independent check against ClinicalTrials.gov — if a trial that was still running when the prediction was made has since gone TERMINATED or WITHDRAWN, that validates a Sell/Avoid-type rating and invalidates a Buy-type rating; a status of SUSPENDED is left inconclusive rather than treated as a failure, since a suspension can be administrative or temporary. Every graded verdict across all three lenses links to the actual evidence it was graded against — a live quote page, an FDA source, or a ClinicalTrials.gov study page. Below 30 evaluated predictions, results are labeled statistically insignificant rather than presented as a real number.

TOTAL counts persisted ORION-generated prediction rows (analyze_pipeline and house_prediction) independent of UI pagination. Client-supplied and unknown-source rows remain in the append-only ledger but are reported as exclusions, not used in ORION accuracy. Failed, pending, and ungradable evaluations remain visible in their applicable cohort states.

Curious what happens when we backtest our own model against 83,000+ real historical trials? Read the full 2022 backtest case study →

Want to know if a "VERY HIGH IMPACT" tag actually means the stock moved more? See the Catalyst Impact Backtest →

The 2022 study above is one entry in a running, dated log — every time ORION checks a formula against real outcomes again, successful or not, it's added here rather than replacing the last one. See the full Backtest & Validation Log →