AI Methodology
BestBest Reviews asks several AI models to review the same subject, then shows where their independent scores and explanations agree or differ. The aim is comparison, not a claim that the combined result is objectively correct.
The shared prompt
ChatGPT, Claude, Gemini, Grok, and DeepSeek receive the same concise prompt with only the topic name changed. Each model is asked for valid JSON containing a score from 0 to 5, a one-sentence verdict, and brief reasoning. The prompt requires a neutral, analytical tone and says not to force agreement or disagreement.
The complete prompt used for a review is displayed on that review page. Each response is labeled with its provider, model, prompt version, and generation date.
Summary and jury reviews
The introductory summary, pros, cons, related topics, and source links are separate from the short jury ballots. When a legacy review exists, those fields are reused from the site's existing OpenAI-generated archive. For a genuinely new topic, ChatGPT prepares one stored summary before the jury votes. The jury cards contain the models' independently generated scores, verdicts, and reasoning.
Publishing and updates
Model responses are generated in a controlled background process and saved in MongoDB. Viewing a completed page does not call every model again. A new-format review can be published after at least three jurors have responded; later runs can fill missing jurors and refresh stale responses.
Visits to review pages can add those subjects to a promotion queue, whether or not they already exist in the legacy archive. The queue has daily page and model-call limits, preserves existing content when available, and never makes a reader wait for model generation.
How to interpret the score
The displayed jury score is the arithmetic mean of the available model scores. It is a summary of model opinions, not a scientific measurement, factual verdict, or substitute for expert advice.
Limitations
- AI can produce incorrect or outdated statements.
- Scores can change when a prompt or model changes.
- Models can share biases or repeat common misconceptions.
- Not every statement has received individual human fact-checking.
- A small jury should not be treated as representative of people.
Important claims should be checked against current, authoritative sources, especially for medical, legal, financial, historical, and other consequential subjects.