Published comparison

How the Obside benchmark is measured

The homepage comparison is a production benchmark, not an illustrative chart. This page publishes the aggregate results, evaluation protocol, and limitations needed to interpret it.

Benchmark snapshot

Production assistants, live reference data, blind scoring.

Aggregate results

Same published scores, one source of truth

The homepage chart and the downloadable result use this same dataset. Scores combine accuracy, relevance, and freshness on a 10-point scale.

AssistantScore / 10
Obside9.7
ChatGPT8.4
Gemini8.1
DeepSeek8.0
Claude7.9

Evaluation protocol

Equal prompts, anonymized answers

  1. 01Every assistant receives the same finance prompt set.
  2. 02Answers are collected with live-data or web-search access enabled where the product supports it.
  3. 03Model names are removed before evaluation.
  4. 04A separate neutral judge scores each answer against live reference data.
  5. 05Dimension scores are aggregated and normalized to a 10-point scale.

Interpretation

A snapshot, not a universal ranking

Results describe this prompt set and the products available during the July 2026 run. Model versions, search systems, market conditions, and data availability change over time, so future runs may produce different rankings.

The benchmark measures answer quality for finance questions. It does not measure investment returns and is not financial advice.