This panel brings together speakers from two assessment organizations to discuss how they are using generative AI to support item development while maintaining rigorous human oversight. Both use human-in-the-loop approaches to increase content development efficiency, improve consistency, and optimize SME time without replacing expert judgment. Speakers from the first organization will present an in-house tool that assists item writers with distractor creation, reference suggestions, rationale drafting, and rapid item generation. They will share pilot results and productivity metrics across four disciplines. The second organization will describe its use of in-house AI tools to generate formative assessment content and rationales for SME review and revision, highlighting lessons learned, implementation challenges, and outcomes. Panelists will compare their approaches, discuss practical considerations, and provide recommendations for assessment practitioners, followed by audience Q&A.
Pamela Kaliski, American Board of Internal Medicine
Ally Kulesher, National Board of Medical Examiners
Kristina Oberle, National Board of Medical Examiners