When my colleagues and I published a systematic review of durable learning in nursing education in 2024, we looked at 51 studies. The answer was the same across nearly all of them: retrieval practice, case-based reasoning, simulation, spaced repetition. Techniques that make a student do something with a concept instead of just receiving it. None of this was new. Educators have known the evidence for active learning for decades. The harder finding came in the research that followed.
We asked nursing students what actually worked in their classrooms. They confirmed it. Case studies, group discussions, simulation. Lecture was described as ineffective more often than not. Assigned readings went unread. Quizzes produced anxiety more reliably than understanding. Then we asked faculty the same question, and they agreed with the students. So we asked what stopped them from teaching the way they already knew worked. The answer, almost every time, was time. Most faculty also told us they couldn’t reliably tell whether what they were currently doing was working at all.
That’s not a finding about nursing education. It’s a finding about every discipline that requires more content and more feedback than one person can produce in a normal semester. The methods that work are too labor-intensive to run at scale, so most faculty don’t run them. That gap, between what the evidence says and what a single faculty member can actually deliver, is the problem the debate about AI in higher education keeps missing.
I’ve practiced as a nurse practitioner in cardiology since 1994. I direct the acute care nurse practitioner program at Saint Joseph’s University. And I’m a co-author on those studies. Two years ago I started building AI tools for my own students instead of waiting for someone else to build them. I wrote the pedagogy myself and deployed it in my own courses.
Here’s what changed. Generating a clinical case variation used to take me an hour: a new patient presentation, different labs, a different complicating history. Now it takes thirty seconds. I can run spaced retrieval across an entire curriculum and give every student a formative clinical reasoning session without being in the room. In an asynchronous online program, where I’m never actually in the room to begin with, that’s not a convenience. It’s the difference between the technique existing on paper and happening at all. The techniques that were too expensive to run at scale are affordable now, because the thing that made them expensive, the hours required to produce enough content at enough variation, is gone.
That difference shows up at the bedside, not just on the exam. A student who has reasoned through a dozen versions of acute dyspnea, heart failure one time, pulmonary embolism the next, a COPD exacerbation after that, each with different labs and a different medication history, doesn’t just score better on the board question. She walks into a room ready to build a differential instead of retrieve a memorized one. Faculty get something too. One of the studies found that most nursing faculty couldn’t measure whether their teaching was working. When a formative reasoning session shows the gaps before the final exam, that stops being true. The feedback arrives while there’s still time to act on it.
I haven’t seen the major edtech platforms building for that. They’re building for the existing structure: summarize the reading, generate a quiz from the upload, answer a student’s question at 11pm. Faster than what came before it. Not pedagogically different. Building something that’s actually different means knowing what retrieval practice requires, what makes a case generate reasoning instead of pattern-matching, what kind of feedback changes performance and what kind arrives too late to matter. Knowing the technology and knowing the learning science don’t often live in the same person. When they don’t, the tool optimizes for the wrong thing.
I’m not arguing every faculty member should build their own AI. I built mine because I’m wired for it, and I had a specific problem I wanted solved my own way. Most faculty aren’t wired that way and don’t need to be. Tools that generate case variation, run retrieval at scale, deliver feedback without a faculty member in the loop already exist, and more are coming. The question isn’t whether to build. It’s whether the people who understand what learning actually requires are in the room when these tools get built. Right now, mostly, they aren’t. Engineers building AI for higher education understand language models. They don’t always know that retrieval practice without enough variation produces recognition instead of recall, or that feedback arriving after the exam changes nothing.
It’s part of why NursingEdAI exists. Faculty need the hours back before they can even think about redesigning what they teach.
The debate about whether students will use AI to cheat will settle itself. Institutions adapt, policy catches up, the field finds its footing. It always does. The debate that won’t settle on its own is whether AI closes the gap between what the evidence says works and what one faculty member can deliver in a normal semester, or whether it just makes the same lecture arrive faster. That was never a question about the students. It was always a question about us.