Benchmark notice

Benchmark pages define transparent evaluation frameworks and reusable scenarios. They are not product rankings and do not invent measurement scores.

Editorial status

Published 2026-07-23 · Last reviewed 2026-07-23 · Next review due 2027-01-19

  • Review cadence: Every 6 months
  • Verification badge: Verified
  • Review status: Current
  • Evidence level: framework
  • Content owner: ONULSURI Editorial

Read the AI editorial policy

Read the shared methodology and outcome rubric

This suite helps reviewers evaluate general-purpose chat assistants on ordinary work and learning tasks using fixed inputs and recorded evidence.

Outcomes are qualitative labels tied to observable criteria. A single run does not establish product superiority, and results can change with model, plan, settings, and date.

Intended use

  • Compare how a specific assistant handles the same fixed prompts across repeated runs
  • Document instruction-following, clarity, and revision behavior with screenshots or exports
  • Train reviewers on a shared qualitative rubric before publishing later measured reports
  • Capture environment notes that explain variance between runs

Not intended for

  • Declaring a top-ranked assistant or publishing ranking tables
  • Converting qualitative labels into numeric leaderboards
  • Evaluating coding agents, research browsing, or image generators (use the dedicated suites)
  • Prompts that request malware, phishing, credential theft, or real-person impersonation

Evaluation dimensions

Instruction following

Whether the response respects stated constraints, format, length, and required sections.

Clarity and structure

Whether the output is readable, organized, and usable without unnecessary filler.

Factual caution

Whether the assistant avoids inventing details when information is missing and flags uncertainty.

Revision quality

Whether requested edits preserve intent while applying the stated change.

Refusal and boundaries

Whether the assistant stays within safe, appropriate scope for the scenario.

Qualitative outcome rubric

These labels are evidence judgments for a single scenario run. They are not product scores and must not be totaled into rankings.

Meets

Observable evidence shows the response satisfied the scenario criteria without material gaps.

Partially meets

Some criteria are satisfied, but important gaps, omissions, or inconsistencies remain.

Does not meet

Evidence shows the response failed one or more required criteria in a material way.

Not applicable

The criterion does not apply to this product surface, plan, or allowed tool configuration.

Insufficient evidence

The run cannot be judged fairly because required evidence is missing, incomplete, or interrupted.

Scenarios

instruction

Constrained bullet summary

Objective: Check whether the assistant follows hard length and format constraints when summarizing a short passage.

Setup

  • Use a fresh chat with default settings unless the product requires a mode selection
  • Do not enable browsing or file uploads for this scenario
  • Paste the scenario input exactly once and capture the full reply

Input

Summarize the following passage in exactly 5 bullet points. Each bullet must be one sentence and at most 20 words. Do not add a title, intro, or closing. Passage: A neighborhood library launched evening study hours three nights a week. Staff noticed more high-school students using quiet rooms after dinner. The library also offered free printing for school assignments up to ten pages. Volunteers helped new visitors find databases and citation guides. Attendance rose through the school year without extending weekend hours.

Expected evidence

  • Exact prompt used
  • Full assistant response
  • Count of bullets and approximate word counts per bullet
  • Notes on any extra sections or titles added

Evaluation criteria

Exact bullet count

The response contains exactly five bullets and no additional list items.

Observable evidence: Number of list items in the output equals five.

Outcome guidance: Meets if there are exactly five bullets; partially meets if four or six with no other violations; does not meet if the format is a paragraph or the count is clearly wrong.

Length limits

Each bullet is one sentence and stays near the stated word limit.

Observable evidence: Per-bullet sentence count and word count notes from the reviewer.

Outcome guidance: Meets if all bullets are one sentence and within or very near 20 words; partially meets if one bullet mildly exceeds; does not meet if multiple bullets ignore the limit.

No extra framing

The response omits titles, introductions, and closings as instructed.

Observable evidence: Presence or absence of preamble or postscript text.

Outcome guidance: Meets if only the five bullets appear; partially meets if a short label appears; does not meet if a full intro or closing essay is added.

Failure modes

  • Adds an introduction or concluding paragraph
  • Uses more or fewer than five bullets
  • Writes multi-sentence bullets that ignore the word limit
  • Invent details not present in the passage

Reviewer notes

  • Count words generously for hyphenated terms; note borderline cases
  • Do not grade stylistic preference if constraints are met

Safety notes

  • Use only the provided passage; do not substitute private documents
  • Do not request or paste personal student records

instruction

Ambiguous request handling

Objective: Observe whether the assistant asks clarifying questions or states assumptions before over-committing on an underspecified task.

Setup

  • Start a new conversation
  • Do not provide extra context beyond the input
  • Record whether clarification questions appear before a full plan

Input

Help me plan the rollout. Use a short checklist. Keep it practical.

Expected evidence

  • Full response including any clarifying questions
  • Whether assumptions were listed explicitly
  • Whether a checklist was produced and at what length

Evaluation criteria

Handles ambiguity

The assistant acknowledges missing context instead of inventing a specific product or company setting.

Observable evidence: Clarifying questions or clearly labeled assumptions in the reply.

Outcome guidance: Meets if it asks useful clarifiers or states assumptions; partially meets if it proceeds with generic advice while noting gaps; does not meet if it fabricates a specific rollout context as fact.

Checklist format

When it answers, it provides a short practical checklist rather than a long essay only.

Observable evidence: Presence of a concise checklist structure.

Outcome guidance: Meets if a short checklist is present; partially meets if mixed prose with a short list; does not meet if only a long narrative with no checklist.

Avoids false specificity

The reply does not invent deadlines, budgets, or stakeholder names that were not supplied.

Observable evidence: Invented proper nouns, dates, or figures not in the prompt.

Outcome guidance: Meets if no fabricated specifics appear; partially meets if one soft example is clearly marked as hypothetical; does not meet if invented facts are presented as given.

Failure modes

  • Assumes a specific industry or tool stack without labeling it as an example
  • Produces a very long plan with no checklist
  • Ignores that the request is underspecified

Reviewer notes

  • Clarifying questions alone can still meet if they are substantive
  • Generic checklists are acceptable when assumptions are labeled

Safety notes

  • Do not inject real internal project names or confidential timelines into the prompt

analysis

Structured analysis of a small table

Objective: Evaluate careful reading of structured text and whether conclusions stay within the provided numbers.

Setup

  • Disable browsing if it can be toggled
  • Paste the table exactly as written
  • Ask the assistant only what the input asks; do not add hints

Input

Here is a small attendance table for a community workshop (numbers are total attendees):
Week 1: 12
Week 2: 15
Week 3: 9
Week 4: 18
Tasks: (1) Identify the week with the lowest attendance. (2) State the average attendance rounded to one decimal place. (3) List one limitation of drawing strong conclusions from only four weeks. Do not invent causes.

Expected evidence

  • Full response with answers to all three tasks
  • Reviewer calculation of the average for comparison
  • Notes on any invented causal explanations

Evaluation criteria

Lowest week identified

Correctly identifies Week 3 as the lowest attendance week.

Observable evidence: Explicit mention of Week 3 or attendance 9 as lowest.

Outcome guidance: Meets if Week 3 is correctly identified; does not meet if another week is named; insufficient evidence if the answer is missing.

Average calculation

Reports the average attendance as 13.5 when rounded to one decimal place.

Observable evidence: Numeric average stated in the response.

Outcome guidance: Meets if 13.5 is reported; partially meets if the unrounded 13.5 equivalent is clear but formatting differs; does not meet if the average is wrong.

Limitation without invention

States a methodological limitation without inventing causes for the dip or rise.

Observable evidence: Limitation statement and absence of fabricated causal claims.

Outcome guidance: Meets if a clear limitation is given without invented causes; partially meets if a mild speculative cause is hedged; does not meet if confident false causes are asserted.

Failure modes

  • Misidentifies the lowest week
  • Miscalculates the average
  • Invent weather, marketing, or staffing causes not in the data

Reviewer notes

  • Average check: (12+15+9+18)/4 = 13.5
  • Focus on fidelity to the table, not prose elegance

Safety notes

  • Do not replace the table with real customer analytics

revision

Tone-preserving email revision

Objective: Check whether a revision request changes tone as asked while preserving factual content.

Setup

  • Provide the draft and revision instructions in one message
  • Do not send the email; evaluation is on the draft only
  • Capture the revised email text in full

Input

Revise the email below to be warmer and more concise. Keep the meeting time and the agenda items unchanged. Do not add new agenda items. Do not change the meeting to a different day.

Draft:
Hello team,
I am writing to inform you that we will convene tomorrow at 10:00. Agenda: project status, open questions, next steps. Please arrive prepared.
Regards,

Expected evidence

  • Original draft and revision instructions
  • Full revised email
  • Side-by-side notes on tone, length, and preserved facts

Evaluation criteria

Preserves facts

Meeting time remains tomorrow at 10:00 and agenda items stay the same.

Observable evidence: Time and agenda strings in the revised email.

Outcome guidance: Meets if time and agenda are unchanged; does not meet if time, day, or agenda items are altered or expanded.

Warmer tone

The revision reads warmer than the stiff original without becoming unprofessional.

Observable evidence: Greeting, phrasing, and closing compared with the draft.

Outcome guidance: Meets if warmth improves while staying professional; partially meets if only minor wording changes; does not meet if tone is unchanged or becomes inappropriate.

More concise

The revised email is shorter or tighter without dropping required facts.

Observable evidence: Relative length and redundancy removed.

Outcome guidance: Meets if clearly more concise; partially meets if similar length but cleaner; does not meet if longer and more verbose.

Failure modes

  • Changes the meeting time or day
  • Adds new agenda items
  • Makes the email longer while claiming concision

Reviewer notes

  • Judge warmth and concision qualitatively; document examples
  • Do not require a specific greeting formula

Safety notes

  • Use only the fictional draft; do not paste real employee emails with private details

instruction

Multi-constraint how-to

Objective: Evaluate adherence to several simultaneous constraints in a practical how-to answer.

Setup

  • Use default chat settings
  • Do not enable tools that fetch live pages unless required by the product UI
  • Record whether all sections appear in the requested order

Input

Explain how to prepare a shared meeting agenda document for a five-person team. Constraints: (1) Use exactly three numbered steps. (2) After the steps, include a section titled Risks with exactly two bullets. (3) End with one sentence labeled Tip:. (4) Do not mention any specific software brand names.

Expected evidence

  • Full response
  • Checklist of whether each constraint was met
  • Notes on any brand names mentioned

Evaluation criteria

Three numbered steps

Contains exactly three numbered steps before the Risks section.

Observable evidence: Numbered step count and ordering.

Outcome guidance: Meets if exactly three numbered steps appear; partially meets if three steps are present but unnumbered; does not meet if step count is wrong.

Risks and Tip sections

Includes a Risks section with exactly two bullets and ends with a one-sentence Tip: line.

Observable evidence: Section headings, bullet count, and final Tip sentence.

Outcome guidance: Meets if both structures match; partially meets if one is slightly off; does not meet if either is missing.

No brand names

Avoids specific software brand names as instructed.

Observable evidence: Any named commercial tools in the reply.

Outcome guidance: Meets if no brand names appear; partially meets if a brand appears once then corrected in the same reply; does not meet if brands are recommended as required tools.

Failure modes

  • Uses bullet steps instead of numbered steps
  • Wrong number of risks
  • Mentions specific document or chat product brands

Reviewer notes

  • Generic terms like shared document or calendar are fine
  • Order matters: steps, then Risks, then Tip

Safety notes

  • Do not expand the prompt to include production access credentials

analysis

Uncertainty disclosure on incomplete facts

Objective: See whether the assistant separates knowns from unknowns when asked to advise with incomplete information.

Setup

  • New chat, no prior context
  • Do not supply extra facts after the first reply unless documenting a follow-up as a separate note
  • Capture the first complete answer

Input

Our volunteer group may host an outdoor reading event next month. We do not yet know the park rules, the weather plan, or how many chairs we can borrow. Give a brief preparation outline. Clearly separate What we know, What we still need to confirm, and Suggested next questions. Do not invent park rules or weather forecasts.

Expected evidence

  • Full structured response
  • Notes on any invented rules, forecasts, or headcounts
  • Whether the three requested sections appear

Evaluation criteria

Required sections

Includes the three requested section labels with relevant content.

Observable evidence: Section headings and content under each.

Outcome guidance: Meets if all three sections are present; partially meets if content is present but labels differ slightly; does not meet if sections are missing.

No invented facts

Does not invent park rules, weather forecasts, or borrowed-chair counts.

Observable evidence: Any specific claims not supported by the prompt.

Outcome guidance: Meets if unknowns stay unknown; partially meets if soft examples are clearly hypothetical; does not meet if invented rules or forecasts are stated as fact.

Actionable next questions

Suggested questions are concrete and tied to the stated unknowns.

Observable evidence: Quality and relevance of the questions list.

Outcome guidance: Meets if questions map to rules, weather plan, and chairs; partially meets if questions are generic but useful; does not meet if no questions or irrelevant questions only.

Failure modes

  • Fabricates park permit requirements
  • States a weather forecast as known
  • Omits the required section structure

Reviewer notes

  • Hypothetical examples are acceptable only when labeled as examples
  • Brevity is preferred if structure is complete

Safety notes

  • Do not use real home addresses or private volunteer contact lists in the prompt

Evidence requirements

Environment log

Record product name, visible model or mode, plan type, platform, enabled tools, settings, and test date.

Prompt and full output

Keep the exact scenario input and the complete assistant response for each run.

Constraint checklist

Note which stated constraints were met, partially met, or missed, with brief quotes.

Interruption notes

Document refusals, tool errors, truncated replies, or UI changes that blocked a fair run.

Execution guidance

  • Run one scenario per chat unless the product forces continued context; note if context carried over
  • Do not edit the scenario input mid-run; if the UI forces a change, mark evidence incomplete
  • Apply qualitative outcome labels only; do not assign points or ranks
  • Prefer insufficient evidence when the reply is truncated or tools fail mid-answer

Reproducibility notes

  • Same prompt can yield different wording across runs; preserve each full output
  • Model or mode switches invalidate direct comparison unless both environments are logged
  • Regional UI differences and feature flags can change available tools
  • Cite the suite last-reviewed date and methodology version in any later report

Limitations

Not a leaderboard

This suite does not produce rankings, scores, or claims of a best general assistant.

Everyday task scope

Scenarios focus on common chat tasks and do not cover specialized professional certification domains.

Environment sensitivity

Plan, model, and tool settings can change outcomes without changing the scenario text.

Reviewer variance

Qualitative judgments may differ between reviewers; evidence notes should make disagreements inspectable.

Related comparisons

Related tool overviews

  • AI HubOverview of ONULSURI AI guides and where each section fits.
  • AI CompareSide-by-side comparisons of assistants and tools.
  • AI Tool DirectoryCategory directory and tool overviews.
  • AI PricingPlan structure and upgrade guidance without fabricated prices.
  • Prompt LibraryReusable prompts for coding, writing, and everyday work.
  • AI GuidesEvergreen topic guides for choosing tools and workflows.

FAQ

Does this suite declare a winning assistant?

No. It defines scenarios and evidence expectations only. Reviewers apply qualitative labels without publishing product rankings.

Can I change the prompt to fit my product?

For comparable runs, keep the scenario input fixed. If you must adapt it, treat the result as a separate experiment and document the change.

What if the assistant asks clarifying questions?

Record the questions as part of the evidence. For ambiguity scenarios, clarifiers can support a meets outcome when they address missing context.

Should I enable browsing?

Only if the scenario setup allows it. For this suite’s core scenarios, browsing is usually off so answers stay tied to the provided text.

How should ties or mixed results be reported?

Report per-criterion qualitative labels and preserve evidence. Do not collapse mixed results into a single product score.