How the Obside benchmark is measured
The homepage comparison is a production benchmark, not an illustrative chart. This page publishes the aggregate results, evaluation protocol, and limitations needed to interpret it.
Benchmark snapshot
Production assistants, live reference data, blind scoring.
Aggregate results
Same published scores, one source of truth
The homepage chart and the downloadable result use this same dataset. Scores combine accuracy, relevance, and freshness on a 10-point scale.
| Assistant | Score / 10 |
|---|---|
| Obside | 9.7 |
| ChatGPT | 8.4 |
| Gemini | 8.1 |
| DeepSeek | 8.0 |
| Claude | 7.9 |
Evaluation protocol
Equal prompts, anonymized answers
- 01Every assistant receives the same finance prompt set.
- 02Answers are collected with live-data or web-search access enabled where the product supports it.
- 03Model names are removed before evaluation.
- 04A separate neutral judge scores each answer against live reference data.
- 05Dimension scores are aggregated and normalized to a 10-point scale.
Interpretation
A snapshot, not a universal ranking
Results describe this prompt set and the products available during the July 2026 run. Model versions, search systems, market conditions, and data availability change over time, so future runs may produce different rankings.
The benchmark measures answer quality for finance questions. It does not measure investment returns and is not financial advice.