Skip to main content

Safer Agentic AIAuto-Assessor

Assess your AI systems against the Safer Agentic AI Framework. Upload your evidence documents and get AI-generated scores and gap analysis across all 16 safety suites.

Scores and remediation suggestions are AI-generated and indicative only. They do not constitute legal, regulatory, or compliance advice; verify results independently before relying on them.

Not sure where to start? Talk to the Safety Advisor

Runs in your browser via WebGPU — no API key, but a ~5 GB one-time model download (Chrome/Edge 113+, Safari 18+). Cloud models available if you have a key.

Testing Ground

See how the auto-assessor works before committing. Browse pre-computed evaluation results using synthetic data from a fictional “Acme Autonomous Vehicles” company.

  • •63 pre-computed results across all 16 suites
  • •Scores from 1 to 5 with justifications and gap analysis
  • •No API key or account needed

AI Safety Assessment

Evaluate your own organization's AI systems with live Claude analysis. Paste your evidence documents and get real-time scoring against any of the 994 evidence requirements.

  • •Uses your own Anthropic API key — stored in your browser's session storage and cleared when you close the tab
  • •Live evaluation with detailed justifications and gap analysis
  • •Evidence is sent directly to Anthropic for evaluation, never to our servers. Data handling
16
Safety suites
9 drivers · 7 inhibitors
238
Sub-goals
with SFRs & evidence items
994
Evidence requirements
each individually assessable

How It Works

  1. 1Browse the 16 safety suites covering drivers (goals, values, security, transparency, governance) and inhibitors (deception, opacity, uncertainty).
  2. 2Select a sub-goal and review its Safety Functional Requirements (SFRs) and evidence requirements.
  3. 3Paste or upload evidence documentation and run automated evaluation against each requirement.
  4. 4Receive scores (1–5), justifications, identified gaps, and relevant excerpts for each criterion.

Scoring Rubric

5
Excellent
Evidence comprehensively addresses all aspects of the requirement.
4
Good
Most aspects addressed. Minor areas for improvement remain.
3
Average
Some aspects addressed, but significant gaps remain.
2
Poor
Some relevant information, but major aspects unaddressed.
1
Unacceptable
Evidence does not address the requirement at all.