Benchmark notice
Benchmark pages define transparent evaluation frameworks and reusable scenarios. They are not product rankings and do not invent measurement scores.
Editorial status
Published 2026-07-23 · Last reviewed 2026-07-23 · Next review due 2027-01-19
- Review cadence: Every 6 months
- Verification badge: Verified
- Review status: Current
- Evidence level: framework
- Content owner: ONULSURI Editorial
Read the AI editorial policy
Read the shared methodology and outcome rubric
This suite helps reviewers evaluate research-oriented assistants using fixed questions, clear rules about supplied sources versus browsing, and recorded evidence.
Citations must be verified by a human against primary sources. Qualitative labels are not rankings, and a single run cannot prove general research superiority.
Intended use
- Test how an assistant handles supplied excerpts without inventing outside facts
- Document browsing-on versus browsing-off behavior when the product allows a toggle
- Practice citation verification workflows before trusting linked claims
- Capture uncertainty disclosure when evidence is incomplete
Not intended for
- Publishing research leaderboards, accuracy percentages, or ranking tables
- Treating model citations as verified without opening the sources
- Prompts that seek malware, phishing, credential theft, or illegal evasion advice
- Substituting private medical, legal, or confidential business files without approval
Evaluation dimensions
Source fidelity
Whether claims stay within supplied sources when browsing is not permitted.
Citation quality
Whether citations are specific, relevant, and presented in a verifiable way.
Uncertainty handling
Whether the assistant separates confirmed statements from open questions.
Synthesis clarity
Whether comparisons and summaries are structured and usable for a human reviewer.
Browse policy compliance
Whether the run respects the scenario’s supplied-sources versus browsing rules.
Qualitative outcome rubric
These labels are evidence judgments for a single scenario run. They are not product scores and must not be totaled into rankings.
Meets
Observable evidence shows the response satisfied the scenario criteria without material gaps.
Partially meets
Some criteria are satisfied, but important gaps, omissions, or inconsistencies remain.
Does not meet
Evidence shows the response failed one or more required criteria in a material way.
Not applicable
The criterion does not apply to this product surface, plan, or allowed tool configuration.
Insufficient evidence
The run cannot be judged fairly because required evidence is missing, incomplete, or interrupted.
Scenarios
research
Supplied sources only — brief
Objective: Check whether the assistant answers strictly from provided excerpts and refuses to invent outside facts.
Setup
- Browsing is not permitted for this scenario — disable web tools if possible
- Provide only the excerpts in the input
- Verify any cited lines against the supplied text
Input
Using ONLY the sources below, answer: What two benefits of evening library hours are explicitly stated? List them as bullets. If a benefit is not stated, say Not stated. Do not use outside knowledge.
Source A: "Evening hours helped students find quiet study rooms after dinner."
Source B: "Volunteers guided new visitors to citation guides during evening sessions."
Source C: "Weekend hours stayed unchanged during the school year."
Expected evidence
- Full response
- Notes confirming browsing was off or unavailable
- Line-by-line check that bullets map to Sources A/B
Evaluation criteria
Stays within sources
Benefits mentioned are grounded in Sources A and B; Source C is not misread as a benefit.
Observable evidence: Bullet content compared with excerpts.
Outcome guidance: Meets if both stated benefits are captured without extras; partially meets if one is missing or lightly paraphrased past the text; does not meet if outside benefits are invented.
Respects Not stated
Does not present unsourced claims as facts from the packet.
Observable evidence: Presence of invented claims vs Not stated usage when needed.
Outcome guidance: Meets if no unsourced benefits appear; does not meet if outside knowledge is presented as sourced.
Format compliance
Uses a short bullet list as requested.
Observable evidence: Response structure.
Outcome guidance: Meets if bullets are used; partially meets if short prose equivalent; does not meet if long unsourced essay.
Failure modes
- Adds benefits from general knowledge
- Claims weekend expansion contrary to Source C
- Ignores the supplied-sources-only rule
Reviewer notes
- Correct benefits concern quiet study rooms and volunteer citation help
- Weekend hours unchanged is context, not a listed benefit
Safety notes
- Do not replace excerpts with confidential internal research memos
research
Citation verification drill
Objective: Evaluate whether the assistant provides checkable citations and whether a human can verify them.
Setup
- Browsing may be permitted if the product supports it — record the setting
- Human must open each cited source before labeling citation quality
- If a citation cannot be opened, mark insufficient evidence for that claim
Input
What is the purpose of the IETF RFC series at a high level? Provide 2-4 sentences and include at least two citations a reviewer can open (URLs or precise document identifiers). If you are unsure, say what you could not verify.
Expected evidence
- Full answer text
- List of citations returned
- Human verification notes for each citation (reachable, relevant, supports claim)
- Browse on/off setting used
Evaluation criteria
Provides checkable citations
Includes at least two citations with URLs or precise identifiers.
Observable evidence: Citation count and specificity.
Outcome guidance: Meets if two or more checkable citations appear; partially meets if only one is checkable; does not meet if citations are missing or too vague to verify.
Citations support claims
After human verification, cited sources reasonably support the stated purpose.
Observable evidence: Reviewer verification notes.
Outcome guidance: Meets if verified sources support the summary; partially meets if partly relevant; does not meet if citations are unrelated or fabricated; insufficient evidence if links cannot be checked.
Uncertainty disclosed
States limits when verification is incomplete rather than overclaiming.
Observable evidence: Uncertainty language in the answer.
Outcome guidance: Meets if limits are clear when needed; partially meets if mildly overconfident; does not meet if false certainty is asserted despite weak sources.
Failure modes
- Invented URLs or document IDs
- Citations that do not support the claim
- No citations despite the request
Reviewer notes
- Verification is mandatory — do not trust citation formatting alone
- Record browse setting because it changes available evidence
Safety notes
- Do not instruct the assistant to bypass paywalls or access controls
- Avoid scraping personal data from sources
research
Conflicting excerpts synthesis
Objective: See whether the assistant surfaces disagreement between supplied sources instead of forcing a false consensus.
Setup
- Browsing is not permitted — use supplied excerpts only
- Ask for explicit conflict callouts
- Verify that each side is attributed to the correct source label
Input
Sources only. Summarize where these excerpts agree and where they conflict. Use sections: Agreement, Conflict, Open question. Do not resolve the conflict with outside knowledge.
Source 1: "The pilot reduced average ticket wait time to about 8 minutes during weekday mornings."
Source 2: "Independent observers recorded weekday-morning waits closer to 15 minutes during the same pilot month."
Source 3: "Both teams agreed signage improvements were deployed in week two."
Expected evidence
- Full structured response
- Mapping of agreement and conflict to source labels
- Confirmation browsing stayed off
Evaluation criteria
Surfaces conflict
Clearly states the wait-time disagreement between Source 1 and Source 2.
Observable evidence: Conflict section content.
Outcome guidance: Meets if the numeric conflict is explicit; partially meets if conflict is vague; does not meet if it averages away the conflict without disclosure.
Attributes agreement
Notes shared agreement on signage improvements from Source 3 / shared ground.
Observable evidence: Agreement section accuracy.
Outcome guidance: Meets if signage agreement is captured; partially meets if agreement is incomplete; does not meet if agreement invents shared claims not present.
No outside resolution
Does not declare which wait-time figure is correct using unsourced authority.
Observable evidence: Presence of unsourced tie-breakers.
Outcome guidance: Meets if conflict remains open with an open question; does not meet if it falsely crowns a winner without evidence.
Failure modes
- Averages 8 and 15 into a single undisputed figure
- Ignores the conflict section requirement
- Imports outside studies to declare a winner
Reviewer notes
- A good open question asks how wait time was measured
- Keep evaluation inside the three excerpts
Safety notes
- Do not replace excerpts with real confidential operations metrics
research
Browse-on freshness check
Objective: When browsing is permitted, check whether the assistant dates its findings and separates confirmed pages from unsure claims.
Setup
- Browsing is permitted for this scenario — enable web tools if available
- Record the date/time of the run
- Human verifies at least two cited pages before final labels
Input
With browsing allowed, identify the current documentation landing page title for MDN Web Docs (or the closest official MDN entry point you can reach). State: (1) page title you found, (2) URL, (3) retrieval date as today if known, (4) one sentence on what you could not confirm. Do not invent a URL.
Expected evidence
- Assistant answer including title and URL
- Browse-enabled environment note
- Human verification of the URL and title
- Run date
Evaluation criteria
Reachable URL
Provides a URL that the reviewer can open and that relates to MDN Web Docs.
Observable evidence: Human open-and-check notes.
Outcome guidance: Meets if URL is reachable and relevant; partially meets if relevant but redirected; does not meet if fabricated or unrelated; insufficient evidence if browsing failed.
Title alignment
Reported title reasonably matches the opened page.
Observable evidence: Comparison of stated title vs browser title.
Outcome guidance: Meets if titles align; partially meets if approximate; does not meet if mismatched or invented.
Limits disclosed
Includes a clear sentence about what could not be confirmed.
Observable evidence: Uncertainty sentence presence.
Outcome guidance: Meets if limits are stated; partially meets if vague; does not meet if absolute certainty is claimed without basis.
Failure modes
- Invented MDN URLs
- Stale mirrors presented as the official entry without caveat
- No verification-friendly URL
Reviewer notes
- If the product cannot browse, mark insufficient evidence rather than forcing the scenario
- Redirects to locale-specific MDN pages can still meet if disclosed
Safety notes
- Do not ask the assistant to bypass login walls
- Stay on publicly reachable documentation pages
research
Claim triage: supported vs unsupported
Objective: Evaluate whether the assistant labels claims as supported, unsupported, or unclear against a fixed source packet.
Setup
- Browsing is not permitted
- Use only the packet in the input
- Require explicit labels for each claim
Input
Browsing off. For each claim, label Supported, Unsupported, or Unclear based ONLY on the packet. Then quote the best supporting sentence or write None.
Packet: "The workshop ran for four weeks. Attendance totals were 12, 15, 9, and 18. Printing was free up to ten pages."
Claims:
1) The workshop ran for four weeks.
2) Attendance never fell below 10.
3) Printing was free for up to ten pages.
4) The workshop will expand next year.
Expected evidence
- Labels for all four claims
- Quotes or None markers
- Reviewer agreement notes
Evaluation criteria
Correct labels
Expected pattern: 1 Supported, 2 Unsupported (week 3 is 9), 3 Supported, 4 Unsupported or Unclear without future plans in packet.
Observable evidence: Label set compared with packet facts.
Outcome guidance: Meets if all four labels are correct; partially meets if one borderline miss on claim 4 wording; does not meet if claim 2 is marked Supported.
Quotes or None
Provides a packet quote for supported claims and None when not supported.
Observable evidence: Quote lines in the response.
Outcome guidance: Meets if quotes/None are present and accurate; partially meets if quotes are paraphrased lightly; does not meet if fake quotes appear.
No browse leakage
Does not import outside facts to justify labels.
Observable evidence: Absence of external references presented as packet evidence.
Outcome guidance: Meets if reasoning stays in-packet; does not meet if outside sources are used while browsing is off.
Failure modes
- Marks claim 2 Supported despite attendance 9
- Invented quote for claim 4
- Skips labels
Reviewer notes
- Claim 4 should not be Supported; Unsupported is preferred, Unclear acceptable if carefully justified
- Keep browsing disabled for a fair run
Safety notes
- Do not insert real student attendance rosters
revision
Revision of an overconfident research draft
Objective: Check whether the assistant revises an overconfident draft to add hedges and source limits without inventing new citations.
Setup
- Browsing is not permitted unless needed to refuse fake citations — prefer no browsing
- Provide the draft exactly
- Verify the revision does not add new unsourced statistics
Input
Revise the draft so it is more careful. Keep it under 120 words. Add a short Limitations sentence. Do not add new statistics. Do not add new citations.
Draft: "Our city definitely cut commute times by 40% after one month of the pilot. Every rider benefited equally. The result is permanent and proven."
Expected evidence
- Revised text
- Word count note
- Checklist confirming no new stats or citations were added
Evaluation criteria
Reduces overclaiming
Softens absolute claims like definitely, every rider, permanent, and proven.
Observable evidence: Comparison of draft vs revision language.
Outcome guidance: Meets if absolute claims are clearly softened; partially meets if only partly softened; does not meet if overclaiming remains.
Adds Limitations sentence
Includes an explicit limitations sentence.
Observable evidence: Presence of a limitations statement.
Outcome guidance: Meets if present and relevant; does not meet if missing.
No new stats or citations
Does not introduce new percentages, studies, or citation markers.
Observable evidence: Scan for new numbers and citations.
Outcome guidance: Meets if none added; does not meet if new statistics or citations appear.
Failure modes
- Keeps “proven” / “definitely” language
- Invent a new study to support the claim
- Exceeds the word budget by a wide margin without note
Reviewer notes
- Removing the 40% figure is acceptable; replacing it with a new figure is not
- Count words approximately and note borderline cases
Safety notes
- Do not revise real regulated disclosures without policy approval
Evidence requirements
Browse mode log
State whether sources were supplied only or browsing was permitted, and whether tools actually ran.
Source packet and answer
Keep the exact excerpts or URLs used, plus the full assistant response.
Citation verification sheet
For each citation, record reachability, relevance, and whether it supports the claim after human checking.
Uncertainty notes
Document what remained unverified, blocked, or contradictory.
Execution guidance
- Read the scenario setup first: supplied sources vs browsing permitted is part of the test
- Never treat citations as verified until a human opens them
- If browsing fails, use insufficient evidence instead of guessing page contents
- Do not convert qualitative labels into research accuracy percentages or rankings
Reproducibility notes
- Live web content changes; record run date and final URLs after redirects
- Product citation formats differ; judge verifiability, not brand-specific formatting
- Supplied-source scenarios should disable browsing when the UI allows
- Retain verification sheets alongside answers for auditability
Limitations
Not an accuracy leaderboard
This suite does not publish research ranking tables or scientifically certified accuracy scores.
Verification still required
Human citation checking takes time and remains mandatory.
Web volatility
Browse-on results can drift as pages change, even with the same prompt.
Domain limits
Scenarios are general research hygiene checks, not legal, medical, or financial advice certification.
Related comparisons
Related tool overviews
- AI Hub — Overview of ONULSURI AI guides and where each section fits.
- AI Compare — Side-by-side comparisons of assistants and tools.
- AI Tool Directory — Category directory and tool overviews.
- AI Pricing — Plan structure and upgrade guidance without fabricated prices.
- Prompt Library — Reusable prompts for coding, writing, and everyday work.
- AI Guides — Evergreen topic guides for choosing tools and workflows.
FAQ
When is browsing allowed?
Only when the scenario setup says browsing is permitted. Supplied-sources scenarios should keep web tools off when possible.
Are model citations trustworthy by default?
No. A human must verify each citation for reachability, relevance, and support before treating a claim as evidenced.
What if a link is dead or blocked?
Record the failure and use insufficient evidence for claims that depend on that citation.
Can I score research quality as a percentage?
This suite uses qualitative labels only. Do not convert results into fabricated accuracy percentages or product rankings.
How should conflicts between sources be handled?
Surface the disagreement and keep an open question. Do not force a false consensus with outside knowledge when browsing is off.