Skip to content
AI benchmark

AI that cites its evidence, or says no. Here's the score.

A fixed set of test cases checks whether PSI's AI accepts only claims its evidence supports, rejects claims it doesn't, and refuses when evidence is too thin.

Latest run

100% correct

Correct
8 of 8
False supports
0 (claims wrongly accepted)
Missed refusals
0 (should have said no)
Run on
3 Oct 2026 · openai/gpt-6-astra
Method

The test cases

B1 · Should accept
Evidence: Our admins abandon setup when permissions are unclear.
Claim: Unclear permissions cause admins to abandon setup.
B2 · Should accept
Evidence: Three team leads asked for an audit preview before inviting colleagues.
Claim: Team leads want to preview access before inviting colleagues.
B3 · Should reject
Evidence: Exports to CSV take too long for large reports.
Claim: Customers want a PDF export option.
B4 · Should reject
Evidence: Two enterprise customers mentioned SSO in renewal calls.
Claim: All customers require SSO to renew.
B5 · Should accept
Evidence: Mobile usage grew 12% last quarter.
Claim: Mobile usage grew last quarter.
B6 · Should reject
Evidence: Users praised the new dashboard layout.
Claim: The new dashboard reduced churn by 30%.
B7 · Should refuse
Evidence: none
Claim: We should rebuild billing to increase revenue.
B8 · Should refuse
Evidence: none
Claim: Customers want dark mode.

Small fixed set; we'll publish a larger one as PSI grows. Scores come from PSI's own test runs, not from an outside auditor.

Build the product organisation that learns.

PSI is onboarding a small group of design partners.