Methodology v1

Benchmark notice

Benchmark pages define transparent evaluation frameworks and reusable scenarios. They are not product rankings and do not invent measurement scores.

Editorial status

Published 2026-07-23 · Last reviewed 2026-07-23 · Next review due 2027-07-23

  • Review cadence: Annually
  • Verification badge: Verified
  • Review status: Current
  • Evidence level: methodology
  • Content owner: ONULSURI Editorial

Read the AI editorial policy

This methodology describes how to run transparent, evidence-based AI evaluations using fixed scenarios and recorded evidence. It is a reusable framework for later measured reports — not a leaderboard and not a claim of scientific certification.

A single scenario cannot establish general product superiority. Model choice, plan, region, enabled tools, settings, and test date can all change outcomes. Prefer “insufficient evidence” over guessing.

Purpose

The purpose of these benchmarks is to make evaluation criteria explicit, comparable across runs, and reviewable by humans — not to declare a single top-ranked product.

Suites define scenarios and evidence expectations. They do not ship pre-filled product winners, numeric scores, or aggregate rankings.

Benchmark design principles

Scenarios are written as static specifications that a reviewer can execute manually against a product they already have permission to use.

Outcomes are qualitative labels tied to observable evidence. They are not weighted into totals and must not be converted into marketing rankings.

  • Prefer concrete, inspectable tasks over vague “quality” claims
  • Record environment details that can change results
  • Preserve inputs, outputs, and reviewer notes for later audit
  • Treat variance across runs as expected, not as failure of the method alone

Scenario construction

Each scenario states an objective, setup constraints, a fixed input, expected evidence, evaluation criteria, likely failure modes, and safety notes.

Scenarios should avoid requiring hidden model internals, undisclosed vendor telemetry, or credentials the reviewer should not share.

Test environment recording

Before judging an outcome, record the environment so another reviewer can interpret the run later.

  • Product name
  • Model or mode when visible
  • Account or plan type
  • Platform (web, desktop, IDE, mobile)
  • Enabled tools or connectors
  • System or workspace settings that affect answers
  • Test date
  • Exact test input
  • Complete output
  • Reviewer notes
  • Known interruptions or failures

Execution procedure

Execute one scenario at a time with the recorded environment held constant for that run.

Do not silently edit the scenario input mid-run. If the product interface forces a change, note it and mark evidence incomplete when needed.

This methodology does not invoke models automatically and does not assume access to private evaluation APIs.

Evidence collection

Evidence is the material a second reviewer could inspect: the prompt, the full response, screenshots or exports when relevant, and notes about tool use or refusals.

If required evidence is missing, choose Insufficient evidence rather than inferring an unobserved result.

Qualitative outcome rubric

Use the five qualitative labels below. Do not assign points, weights, or totals. Do not publish aggregate product rankings from these labels.

Reviewer judgment may differ. Document why a label was chosen using the scenario’s observable-evidence guidance.

Repetition and variance

Repeated runs can produce different outputs even with the same input. Note sampling, temperature-like settings when visible, and tool non-determinism.

Variance is a reason to collect more evidence, not a license to invent a single definitive score.

Human review

A human reviewer remains responsible for labeling outcomes, checking safety constraints, and refusing to run unsafe prompts.

For coding and research scenarios, human review includes inspecting diffs, running tests when available, and verifying citations against primary sources.

Safety and privacy

Do not paste secrets, production credentials, private customer data, or regulated records into evaluation prompts unless an approved policy explicitly allows it.

Scenarios must not request malware, phishing, credential theft, impersonation, or evasion of security controls.

Limitations

This framework cannot prove universal product superiority, predict future model behavior, or certify compliance.

It does not claim to be audited, standardized, or endorsed by an external benchmarking body.

Plan changes, regional availability, and tool access can invalidate older runs without updating the scenario text.

Versioning and review dates

Each suite and this methodology carry a last-reviewed date. When scenario text or rubric guidance changes, update the review date and note the methodology version used in any later report.

Older runs should cite the methodology version and suite revision that were active at test time.

Qualitative outcome rubric

These labels are evidence judgments for a single scenario run. They are not product scores and must not be totaled into rankings.

Meets

Observable evidence shows the response satisfied the scenario criteria without material gaps.

Partially meets

Some criteria are satisfied, but important gaps, omissions, or inconsistencies remain.

Does not meet

Evidence shows the response failed one or more required criteria in a material way.

Not applicable

The criterion does not apply to this product surface, plan, or allowed tool configuration.

Insufficient evidence

The run cannot be judged fairly because required evidence is missing, incomplete, or interrupted.

Environment checklist

  • Product name
  • Model or mode when visible
  • Account or plan type
  • Platform
  • Enabled tools
  • System or workspace settings
  • Test date
  • Exact test input
  • Complete output
  • Reviewer notes
  • Known interruptions or failures

FAQ

Does this methodology publish product rankings?

No. It defines scenarios and qualitative labels. It does not assign scores, weights, or leaderboard positions.

Can one scenario prove general product superiority?

No. A single scenario cannot establish general superiority. Results also depend on model, plan, region, tools, settings, and date.

What should I do when evidence is incomplete?

Use Insufficient evidence. Guessing from partial logs undermines later review.

Is this a certified scientific standard?

No. It is an internal, transparent framework for reproducible manual evaluation. It is not an external certification.

Do you call vendor APIs automatically?

No. These pages are static specifications. Reviewers run products manually with accounts they already control.

Why record plan and tool settings?

Because those factors often change capability and refusal behavior. Without them, later readers cannot interpret the run.