# EduSafeBench Admissions Evidence Packet

## Problem
AI coding assistants are widely used by beginner CS learners, but trustworthiness is rarely audited with transparent, reproducible methods.

## What I built
- Open benchmark infrastructure for AP CSA/AP CSP reliability.
- Public datasets with citations, eval tooling, CI, reports, and a live project website.
- External review workflow with impact tracking and adjudication pipeline.

## Measurable Outputs
- Core cited benchmark set: 300 items (`apcsa_csp_v1_3_300.jsonl`)
- Hard misconception benchmark: 6 items in v1.4, expanded pipeline prepared for 60 items in v1.5
- Models evaluated: 2 (v1.3 leaderboard), 2 (v1.4 hard leaderboard)
- Reviewer submissions logged: 1
- Decisions influenced: 1

## Public Artifacts
- Site: https://kaushikatla-cell.github.io/EduSafeBench/
- Repo: https://github.com/kaushikatla-cell/EduSafeBench
- v1.3 report: `reports/v1_3_report.md`
- v1.4 changelog: `reports/v1_4_changelog.md`

## Why this is high-impact
- Niche focus: reliability and pedagogy safety for beginner CS AI assistants.
- Reproducible engineering: versioned data, scripts, CI workflows, result artifacts.
- Community pathway: reviewer submissions, impact dashboard, and local publication outreach.

## Next 6-week targets
- Reach 10 real reviewer submissions with role diversity.
- Publish real adjudication agreement metrics.
- Expand hard misconception benchmark to 60+ items and publish v1.5 report.
