DICE™ — AI Evaluator

Per-Model Analysis

Four models score it.You see all four.

  • A single AI opinion is one opinion — DICE runs Claude, GPT, Gemini and Perplexity independently.
  • You get the consensus, the spread, and every place the models disagreed, with their reasoning.

Why the ensemble holds up

Four independent reads

Each model scores the app on its own, from the same evidence and the same rubric. No model sees another's answer before scoring.

Consensus and spread

You get the blended score plus how tightly the models agreed. A tight spread is a confident score; a wide one is a flag worth reading.

Where they disagree

The disagreements are listed line by line with each model's reasoning, so you can judge the argument instead of trusting one black box.

Bias-resistant by design

One model's blind spot gets outvoted. Running four families of models is the closest thing to a second opinion an AI report can give you.

What's inside this section

Delivered as part of your DICE report, in the same structure every time.

  • Score table — one column per model
  • Blended consensus score
  • Agreement spread and confidence read
  • Line-by-line disagreements with reasoning
  • Each model's standout observation
  • Where the ensemble flagged low confidence
  • How to read the spread before acting

Who this is for

Investors who want the confidence interval, not just the number — and founders who want to know which criticisms held up across every model before they rebuild anything.

For investors

Questions about this report?

We'll walk you through a real ensemble breakdown, model by model.