EduSafeBench

Public, reproducible reliability benchmarking for AP CSA/AP CSP learning assistants. EduSafeBench measures factual correctness, pedagogy quality, hallucination risk, and unsafe guidance risk.

Current Snapshot

300
Cited benchmark items (v1.3)
2
Models in latest public leaderboard
4.0
Top model factual score (v1.3)
5.0
Top model safety score (v1.3)

Reports and Artifacts

Repository

View source and contribute on GitHub.