Benchmark notice
Benchmark pages define transparent evaluation frameworks and reusable scenarios. They are not product rankings and do not invent measurement scores.
Editorial status
Published 2026-07-23 · Last reviewed 2026-07-23 · Next review due 2027-01-19
- Review cadence: Every 6 months
- Verification badge: Verified
- Review status: Current
- Evidence level: framework
- Content owner: ONULSURI Editorial
Read the AI editorial policy
Read the shared methodology and outcome rubric
This suite helps reviewers evaluate general-purpose chat assistants on ordinary work and learning tasks using fixed inputs and recorded evidence.
Outcomes are qualitative labels tied to observable criteria. A single run does not establish product superiority, and results can change with model, plan, settings, and date.
Intended use
- Compare how a specific assistant handles the same fixed prompts across repeated runs
- Document instruction-following, clarity, and revision behavior with screenshots or exports
- Train reviewers on a shared qualitative rubric before publishing later measured reports
- Capture environment notes that explain variance between runs
Not intended for
- Declaring a top-ranked assistant or publishing ranking tables
- Converting qualitative labels into numeric leaderboards
- Evaluating coding agents, research browsing, or image generators (use the dedicated suites)
- Prompts that request malware, phishing, credential theft, or real-person impersonation
Evaluation dimensions
Instruction following
Whether the response respects stated constraints, format, length, and required sections.
Clarity and structure
Whether the output is readable, organized, and usable without unnecessary filler.
Factual caution
Whether the assistant avoids inventing details when information is missing and flags uncertainty.
Revision quality
Whether requested edits preserve intent while applying the stated change.
Refusal and boundaries
Whether the assistant stays within safe, appropriate scope for the scenario.
Qualitative outcome rubric
These labels are evidence judgments for a single scenario run. They are not product scores and must not be totaled into rankings.
Meets
Observable evidence shows the response satisfied the scenario criteria without material gaps.
Partially meets
Some criteria are satisfied, but important gaps, omissions, or inconsistencies remain.
Does not meet
Evidence shows the response failed one or more required criteria in a material way.
Not applicable
The criterion does not apply to this product surface, plan, or allowed tool configuration.
Insufficient evidence
The run cannot be judged fairly because required evidence is missing, incomplete, or interrupted.
Scenarios
instruction
Constrained bullet summary
Objective: Check whether the assistant follows hard length and format constraints when summarizing a short passage.
Setup
- Use a fresh chat with default settings unless the product requires a mode selection
- Do not enable browsing or file uploads for this scenario
- Paste the scenario input exactly once and capture the full reply
Input
Summarize the following passage in exactly 5 bullet points. Each bullet must be one sentence and at most 20 words. Do not add a title, intro, or closing. Passage: A neighborhood library launched evening study hours three nights a week. Staff noticed more high-school students using quiet rooms after dinner. The library also offered free printing for school assignments up to ten pages. Volunteers helped new visitors find databases and citation guides. Attendance rose through the school year without extending weekend hours.
Expected evidence
- Exact prompt used
- Full assistant response
- Count of bullets and approximate word counts per bullet
- Notes on any extra sections or titles added
Evaluation criteria
Exact bullet count
The response contains exactly five bullets and no additional list items.
Observable evidence: Number of list items in the output equals five.
Outcome guidance: Meets if there are exactly five bullets; partially meets if four or six with no other violations; does not meet if the format is a paragraph or the count is clearly wrong.
Length limits
Each bullet is one sentence and stays near the stated word limit.
Observable evidence: Per-bullet sentence count and word count notes from the reviewer.
Outcome guidance: Meets if all bullets are one sentence and within or very near 20 words; partially meets if one bullet mildly exceeds; does not meet if multiple bullets ignore the limit.
No extra framing
The response omits titles, introductions, and closings as instructed.
Observable evidence: Presence or absence of preamble or postscript text.
Outcome guidance: Meets if only the five bullets appear; partially meets if a short label appears; does not meet if a full intro or closing essay is added.
Failure modes
- Adds an introduction or concluding paragraph
- Uses more or fewer than five bullets
- Writes multi-sentence bullets that ignore the word limit
- Invent details not present in the passage
Reviewer notes
- Count words generously for hyphenated terms; note borderline cases
- Do not grade stylistic preference if constraints are met
Safety notes
- Use only the provided passage; do not substitute private documents
- Do not request or paste personal student records
instruction
Ambiguous request handling
Objective: Observe whether the assistant asks clarifying questions or states assumptions before over-committing on an underspecified task.
Setup
- Start a new conversation
- Do not provide extra context beyond the input
- Record whether clarification questions appear before a full plan
Input
Help me plan the rollout. Use a short checklist. Keep it practical.
Expected evidence
- Full response including any clarifying questions
- Whether assumptions were listed explicitly
- Whether a checklist was produced and at what length
Evaluation criteria
Handles ambiguity
The assistant acknowledges missing context instead of inventing a specific product or company setting.
Observable evidence: Clarifying questions or clearly labeled assumptions in the reply.
Outcome guidance: Meets if it asks useful clarifiers or states assumptions; partially meets if it proceeds with generic advice while noting gaps; does not meet if it fabricates a specific rollout context as fact.
Checklist format
When it answers, it provides a short practical checklist rather than a long essay only.
Observable evidence: Presence of a concise checklist structure.
Outcome guidance: Meets if a short checklist is present; partially meets if mixed prose with a short list; does not meet if only a long narrative with no checklist.
Avoids false specificity
The reply does not invent deadlines, budgets, or stakeholder names that were not supplied.
Observable evidence: Invented proper nouns, dates, or figures not in the prompt.
Outcome guidance: Meets if no fabricated specifics appear; partially meets if one soft example is clearly marked as hypothetical; does not meet if invented facts are presented as given.
Failure modes
- Assumes a specific industry or tool stack without labeling it as an example
- Produces a very long plan with no checklist
- Ignores that the request is underspecified
Reviewer notes
- Clarifying questions alone can still meet if they are substantive
- Generic checklists are acceptable when assumptions are labeled
Safety notes
- Do not inject real internal project names or confidential timelines into the prompt
revision
Tone-preserving email revision
Objective: Check whether a revision request changes tone as asked while preserving factual content.
Setup
- Provide the draft and revision instructions in one message
- Do not send the email; evaluation is on the draft only
- Capture the revised email text in full
Input
Revise the email below to be warmer and more concise. Keep the meeting time and the agenda items unchanged. Do not add new agenda items. Do not change the meeting to a different day.
Draft:
Hello team,
I am writing to inform you that we will convene tomorrow at 10:00. Agenda: project status, open questions, next steps. Please arrive prepared.
Regards,
Expected evidence
- Original draft and revision instructions
- Full revised email
- Side-by-side notes on tone, length, and preserved facts
Evaluation criteria
Preserves facts
Meeting time remains tomorrow at 10:00 and agenda items stay the same.
Observable evidence: Time and agenda strings in the revised email.
Outcome guidance: Meets if time and agenda are unchanged; does not meet if time, day, or agenda items are altered or expanded.
Warmer tone
The revision reads warmer than the stiff original without becoming unprofessional.
Observable evidence: Greeting, phrasing, and closing compared with the draft.
Outcome guidance: Meets if warmth improves while staying professional; partially meets if only minor wording changes; does not meet if tone is unchanged or becomes inappropriate.
More concise
The revised email is shorter or tighter without dropping required facts.
Observable evidence: Relative length and redundancy removed.
Outcome guidance: Meets if clearly more concise; partially meets if similar length but cleaner; does not meet if longer and more verbose.
Failure modes
- Changes the meeting time or day
- Adds new agenda items
- Makes the email longer while claiming concision
Reviewer notes
- Judge warmth and concision qualitatively; document examples
- Do not require a specific greeting formula
Safety notes
- Use only the fictional draft; do not paste real employee emails with private details
instruction
Multi-constraint how-to
Objective: Evaluate adherence to several simultaneous constraints in a practical how-to answer.
Setup
- Use default chat settings
- Do not enable tools that fetch live pages unless required by the product UI
- Record whether all sections appear in the requested order
Input
Explain how to prepare a shared meeting agenda document for a five-person team. Constraints: (1) Use exactly three numbered steps. (2) After the steps, include a section titled Risks with exactly two bullets. (3) End with one sentence labeled Tip:. (4) Do not mention any specific software brand names.
Expected evidence
- Full response
- Checklist of whether each constraint was met
- Notes on any brand names mentioned
Evaluation criteria
Three numbered steps
Contains exactly three numbered steps before the Risks section.
Observable evidence: Numbered step count and ordering.
Outcome guidance: Meets if exactly three numbered steps appear; partially meets if three steps are present but unnumbered; does not meet if step count is wrong.
Risks and Tip sections
Includes a Risks section with exactly two bullets and ends with a one-sentence Tip: line.
Observable evidence: Section headings, bullet count, and final Tip sentence.
Outcome guidance: Meets if both structures match; partially meets if one is slightly off; does not meet if either is missing.
No brand names
Avoids specific software brand names as instructed.
Observable evidence: Any named commercial tools in the reply.
Outcome guidance: Meets if no brand names appear; partially meets if a brand appears once then corrected in the same reply; does not meet if brands are recommended as required tools.
Failure modes
- Uses bullet steps instead of numbered steps
- Wrong number of risks
- Mentions specific document or chat product brands
Reviewer notes
- Generic terms like shared document or calendar are fine
- Order matters: steps, then Risks, then Tip
Safety notes
- Do not expand the prompt to include production access credentials
analysis
Uncertainty disclosure on incomplete facts
Objective: See whether the assistant separates knowns from unknowns when asked to advise with incomplete information.
Setup
- New chat, no prior context
- Do not supply extra facts after the first reply unless documenting a follow-up as a separate note
- Capture the first complete answer
Input
Our volunteer group may host an outdoor reading event next month. We do not yet know the park rules, the weather plan, or how many chairs we can borrow. Give a brief preparation outline. Clearly separate What we know, What we still need to confirm, and Suggested next questions. Do not invent park rules or weather forecasts.
Expected evidence
- Full structured response
- Notes on any invented rules, forecasts, or headcounts
- Whether the three requested sections appear
Evaluation criteria
Required sections
Includes the three requested section labels with relevant content.
Observable evidence: Section headings and content under each.
Outcome guidance: Meets if all three sections are present; partially meets if content is present but labels differ slightly; does not meet if sections are missing.
No invented facts
Does not invent park rules, weather forecasts, or borrowed-chair counts.
Observable evidence: Any specific claims not supported by the prompt.
Outcome guidance: Meets if unknowns stay unknown; partially meets if soft examples are clearly hypothetical; does not meet if invented rules or forecasts are stated as fact.
Actionable next questions
Suggested questions are concrete and tied to the stated unknowns.
Observable evidence: Quality and relevance of the questions list.
Outcome guidance: Meets if questions map to rules, weather plan, and chairs; partially meets if questions are generic but useful; does not meet if no questions or irrelevant questions only.
Failure modes
- Fabricates park permit requirements
- States a weather forecast as known
- Omits the required section structure
Reviewer notes
- Hypothetical examples are acceptable only when labeled as examples
- Brevity is preferred if structure is complete
Safety notes
- Do not use real home addresses or private volunteer contact lists in the prompt
Evidence requirements
Environment log
Record product name, visible model or mode, plan type, platform, enabled tools, settings, and test date.
Prompt and full output
Keep the exact scenario input and the complete assistant response for each run.
Constraint checklist
Note which stated constraints were met, partially met, or missed, with brief quotes.
Interruption notes
Document refusals, tool errors, truncated replies, or UI changes that blocked a fair run.
Execution guidance
- Run one scenario per chat unless the product forces continued context; note if context carried over
- Do not edit the scenario input mid-run; if the UI forces a change, mark evidence incomplete
- Apply qualitative outcome labels only; do not assign points or ranks
- Prefer insufficient evidence when the reply is truncated or tools fail mid-answer
Reproducibility notes
- Same prompt can yield different wording across runs; preserve each full output
- Model or mode switches invalidate direct comparison unless both environments are logged
- Regional UI differences and feature flags can change available tools
- Cite the suite last-reviewed date and methodology version in any later report
Limitations
Not a leaderboard
This suite does not produce rankings, scores, or claims of a best general assistant.
Everyday task scope
Scenarios focus on common chat tasks and do not cover specialized professional certification domains.
Environment sensitivity
Plan, model, and tool settings can change outcomes without changing the scenario text.
Reviewer variance
Qualitative judgments may differ between reviewers; evidence notes should make disagreements inspectable.
Related comparisons
Related tool overviews
- AI Hub — Overview of ONULSURI AI guides and where each section fits.
- AI Compare — Side-by-side comparisons of assistants and tools.
- AI Tool Directory — Category directory and tool overviews.
- AI Pricing — Plan structure and upgrade guidance without fabricated prices.
- Prompt Library — Reusable prompts for coding, writing, and everyday work.
- AI Guides — Evergreen topic guides for choosing tools and workflows.
FAQ
Does this suite declare a winning assistant?
No. It defines scenarios and evidence expectations only. Reviewers apply qualitative labels without publishing product rankings.
Can I change the prompt to fit my product?
For comparable runs, keep the scenario input fixed. If you must adapt it, treat the result as a separate experiment and document the change.
What if the assistant asks clarifying questions?
Record the questions as part of the evidence. For ambiguity scenarios, clarifiers can support a meets outcome when they address missing context.
Should I enable browsing?
Only if the scenario setup allows it. For this suite’s core scenarios, browsing is usually off so answers stay tied to the provided text.
How should ties or mixed results be reported?
Report per-criterion qualitative labels and preserve evidence. Do not collapse mixed results into a single product score.