Measuring the Professional Educational Competence of Foundation Models
Foundation models tutor, assess, and instruct at population scale, requiring externally defined measures of educational competence. Existing benchmarks emphasize difficult academic problems or isolated synthetic educational tasks rather than authentic teacher-entry standards. We introduce EDU 1.0 (Educational Due Diligence for Foundation Models), which uses teacher-entry assessments as proxies for educational competence rather than substitutes for human qualification. EDU 1.0 comprises 10,012 questions from teacher certification and recruitment examinations in the United States, China, and India, including the U.S. Praxis series, China's National Teacher Qualification Examination, and India's Kendriya Vidyalaya Sangathan examinations. It covers foundational literacy and knowledge, pedagogical principles, and subject-specific pedagogical expertise across language arts, mathematics, science, social science, and education practice. Across 36 foundation-model variants, the strongest proprietary model reaches a response-balanced score of 96.5%; the leading open-weight model trails by 2.1 percentage points, while the leading system deployable on a single accelerator reaches 92.2%. These aggregates conceal a shared limitation: all three systems score higher on general pedagogical principles than on assessments requiring pedagogy to be applied within a discipline. Their subject-assessment scores span 4.1-8.8 points, and their shortfall relative to general pedagogy widens from 2.6 to 4.1 points as capability declines. The outstanding requirement is therefore pedagogical content knowledge, the capacity to make particular subject matter teachable to particular learners, rather than general pedagogy or model scale.