The public infrastructure underlying Artificial Intelligence (AI) in education is growing. Researchers, developers, and educational agencies around the world are depositing datasets, benchmarks, curriculum artifacts, assessment resources, and related research artifacts in open repositories at a scale that now amounts to a public-goods ecosystem.
We began with the K-12 AI Infrastructure Program’s frame of datasets, models, and benchmarks, asking what public resources are available to support responsible AI development for schools. As we expanded across Hugging Face, GitHub, and Digital Object Identifier (DOI) registered repositories, we narrowed the reported inventory to data-bearing records: datasets, benchmark datasets, curriculum and standards artifacts, assessment resources, tutoring research artifacts, and paper-linked data resources. Standalone models, apps, and code pipelines are not counted unless they are associated with a reusable data artifact.
We use “public goods” in two related ways. First, we treat datasets, models, and benchmarks as shared building blocks that can support responsible AI development for education when they are made available for appropriate reuse. Second, following digital public-goods standards and open data traditions, we focus on resources that are openly accessible, reusable, and governed with attention to privacy, safety, and public benefit.
In this blog, we report on the data-bearing part of that infrastructure, not the full universe of models, apps, or software tools. In our scan, teacher professional development records, such as teacher coaching feedback datasets or annotated classroom observation records, can inform instructional coaching systems. Curriculum and standards artifacts such as state K-12 standards datasets, Texas Essential Knowledge and Skills (TEKS) resources, and Next Generation Science Standards (NGSS)-aligned science frameworks, support alignment research. Classroom observation datasets can inform pedagogical AI. Education benchmarks, assessment instruments, and tutoring research artifacts each support a different layer of responsible AI development for education.
Three findings emerged as we built the combined inventory.
Targeted coordination steps would make existing data-bearing public-good records substantially more reusable. These recommended steps would ask repositories, funders, and research groups to standardize information already being recorded unevenly across the ecosystem.
Our scan suggests that a growing ecosystem of data-bearing public resources is already in place. The next step is making these resources easier to discover, understand, and reuse to support responsible AI development in education.
Earlier this week, DrivenData launched a new platform in collaboration with the K-12 AI Infrastructure Program. The platform is designed to support the work of building AI that actually serves students in three key ways: