Training for Scrutiny Without Cost to Adoption: A Two-Site Study of Teacher Professional Development with Generative AI
Description
Generative AI tools are entering classrooms faster than evidence about how teach- ers learn to use them well, and a central design worry is overreliance: that fluent, sometimes wrong outputs will be accepted uncritically. We report a two-site, pre– post professional development study with school teachers in Karnataka, India (122 pre- and 119 post-intervention responses). Both sites received hands-on practice on authentic teaching tasks; one additionally received structured training in prompt construction and output verification. Adoption intention was already at ceiling be- fore training and did not move. What moved was interaction readiness: perceived ease of use (pooled d = 0.42) and AI self-efficacy (d = 0.34). Where verification was trained, teachers’ reported propensity to scrutinise AI outputs rose most of all (d = 0.52) with no accompanying decline in adoption. We measured disposi- tion rather than detection performance, so this is an initial result: it indicates that guarding against overreliance need not be traded against adoption.
Files
Steps to reproduce
DATA COLLECTION 1. Teachers at two institutions in Karnataka, India, attended a single professional development session on generative AI. Site A was a teacher-training centre cohort drawn from private and government institutions; Site B was drawn entirely from one private school. 2. Both sites received hands-on practice with generative AI tools on authentic teaching tasks: drafting lesson plans, generating assessment questions, and refining student feedback. Site B additionally received structured training in prompt construction and output verification. 3. Participants completed an anonymous survey via Google Forms immediately before the session and a second immediately after. Both were administered in English. The opening item required an explicit click-through agreement authorising the University of Jyvaskyla to use responses for academic research; this was reiterated verbally by the facilitator, and participants who did not agree did not proceed. No identifying information was collected. 4. Items were rated on a seven-point Likert scale measuring AI self-efficacy, perceived usefulness, perceived ease of use, perceived enjoyment, and behavioral intention, plus critical thinking toward AI outputs at Site B only. Site B surveys also carried open-ended items on intended classroom use and anticipated challenges. Full item wording is in README.md. 5. Because responses are anonymous, pre- and post-surveys cannot be linked at the individual level. All comparisons are between independent cohorts. REPRODUCING THE ANALYSIS 6. Install dependencies: pip install pandas numpy scipy statsmodels 7. Place the four CSV files and the two scripts in the same directory. 8. Run: python3 05_analysis.py This prints every quantitative result in the article, in the order the results appear there: data integrity checks, participant characteristics, construct means with effect sizes and reliability coefficients, pooled fixed-effect estimates, the correlation matrix, the regression models with HC3 robust standard errors, ceiling diagnostics with a robustness check, item-level changes, and the continuation-platform items. 9. Run: python3 06_thematic_coding.py This prints the thematic frequencies for the two open-ended questions. Add --audit to list every response alongside its assigned codes. 10. Construct scores are the mean of available items, so a respondent contributes whenever at least one item was answered. Three cells in the original export recorded two mutually exclusive options simultaneously and are treated as missing, as are 20 blank item responses out of 4,327.
Institutions
- University of JyväskyläCentral Finland, Jyväskylä