On September 11, 2025, I gave an invited talk for Purdue’s School of Engineering Education Research Seminar: Bridging the Gap: Promoting Mindful and Reflective Use of Generative AI in Undergraduate Engineering and Computing Education. The work was a collaboration with Andres Bejarano, Rhianna Kuperus, and Bárbara Fagundes. We discussed the AI-Lab, a short instructional intervention that gives students a structured opportunity to experiment with generative AI and think through its role in their learning.

Attendees submitted questions before the seminar, and I added answers to the remaining questions in the final slide deck afterward. This post collects those responses and expands the reasoning behind them. It is an edited written companion rather than a transcript. The discussion below concerns the work and questions presented in September 2025; I have also made the limits of the evidence more explicit. The complete slides, including the question appendix, are available here.

What we studied

The AI-Lab combines preparation before class, roughly 30–40 minutes of a lecture, and a follow-up assignment. Students work with course material, examine generated answers, discuss mistakes and strategies with classmates, and reflect on their own use. Instructors choose problems that expose limitations students can investigate using what they are learning in the course.

The AI-Lab workflow, showing preparation before class, guided work with generative AI during the lecture, and follow-up reflection.
The AI-Lab instructional sequence, reproduced from slide 5 of the seminar presentation.

Our evaluation covered Spring and Fall 2024, with 831 students who completed both surveys and remained enrolled. The courses included data structures for CS and data science students, competitive programming, and first-year engineering. We paired pre- and post-intervention survey responses and combined the quantitative analysis with focus groups held five to seven weeks later.

We did not detect a significant overall increase in reported use on homework and programming projects. Students did report changes in comfort, openness, and debugging-related use. Focus group participants described more deliberate prompting, greater attention to incorrect outputs, and decisions about when assistance was interfering with their own thinking. Those accounts help explain why frequency alone gives an incomplete picture of tool use. They do not establish that students learned more, became less dependent, or stopped cheating.

There were also differences between courses in students’ reported desire to use AI. I offered possible explanations during the talk, including different starting habits and course demands. These remain hypotheses: the study does not establish that one major’s students were more dependent or more cautious than another’s. Similarly, I use “AI literacy” here to include evaluating outputs, communicating effectively, and deciding when assistance serves a learning goal. The focus groups give us evidence about students’ reasoning on those issues, rather than a comprehensive measurement of their AI literacy.

1. How did you handle possible dishonesty in self-reported AI use?

This is a real limitation. A student who thinks an honest answer might look like an admission of misconduct has a reason to underreport use. We considered students’ reports of their behavior alongside their perceptions and reflections, but having multiple kinds of questions does not remove that incentive.

Pairing responses lets us examine reported change within students. It cannot tell us whether the reports are accurate, or whether the intervention itself changed how comfortable students felt reporting their use. The focus groups add useful detail, though participants may also present their behavior favorably. I therefore would not use stable reported frequency as evidence that academic dishonesty was absent. Our results describe what students reported, and the explanations they gave us, within those limitations. Slides, p. 26.

2. How do we balance the ethical issues, and is AI like a calculator?

The calculator comparison is useful up to a point. Both raise questions about which skills students need to practice themselves and which parts of a task can reasonably be delegated. The comparison becomes less helpful when it treats all assistance as interchangeable. An LLM can propose an interpretation of a problem, choose an approach, produce a solution, and explain it. That reaches into much more of the work we often want students to practice.

I would start with a specific assignment and learning objective. What is the student supposed to become able to do, and what assistance is compatible with that? There is no single answer to “the ethical dilemma” without specifying the concern. Academic integrity, access, ownership of work, and dependence raise related but different questions. Students need explicit guidance about the choices their course asks them to make. Slides, p. 26.

3. What does mindful, deliberate engagement look like?

A concrete example is asking for help that preserves the next piece of reasoning for the student: “Here is my problem and what I have tried. Help me find the issue by asking guiding questions,” or “Suggest how I could start, then let me work through it.” These are examples of the kind of use I encouraged, rather than quotations from participants.

In the focus groups, students described supplying more context, refining requests, checking whether answers made sense, and sometimes reducing their use because they felt it was taking over their thinking. Those decisions are what I mean by deliberate engagement. A carefully written prompt is only one part of it. A student also needs to ask whether the interaction helped them understand, whether they can explain the result, and whether they could make progress on a related problem independently. Slides, p. 26; participant excerpts on pp. 10–14.

4. Should CS courses return to handwritten assignments?

My answer at the seminar was that I did not expect a broad return to handwritten programming homework. That was my assessment, not a statement of departmental policy. Large enrollments make manual grading difficult, and checking handwritten code introduces its own problems: a solution can be logically close to correct while containing small errors that are tedious to identify without running it.

There are still good reasons to use handwritten work for selected tasks. Tracing an algorithm, explaining an invariant, or developing a small example can reveal useful information about understanding. I was also seeing greater reliance on in-person examinations. The choice should follow the skill being assessed. Moving all the weight to exams would leave us with less room to assess sustained design, revision, and collaboration, which are also things we want students to learn. Slides, p. 27.

5. How could we move beyond self-reports?

Prompt histories, code edits, debugging traces, and observations of students working could all help. For example, a sequence of edits might show whether a student investigated an error, repeatedly requested replacements, or accepted an answer without checking it. Pairing those traces with an explanation from the student would help us interpret what happened.

Two examples I mentioned in the talk were Prather and colleagues’ The Widening Gap (2024), which combined observation, interviews, and eye tracking, and Ghimire and Edwards’ Coding with AI (2024), which analyzed keystrokes alongside students’ prompts and AI replies. These illustrate different ways to study the process behind a submitted program.

Collecting that evidence across courses is difficult. Students use different editors and tools, move between environments, and approach programming in different ways. A recorded interaction captures only part of that process. Instrumentation also needs consent and sensible limits on what is collected. The next step I would find most useful is a smaller study connecting observed behavior to subsequent independent work. More logs alone would not settle whether a student learned; we need evidence that connects the interaction to the capability we care about. Slides, p. 28.

6. How does this differ from earlier educational technologies?

Concerns about shortcuts and dependence have a long history. The part that concerned me in this seminar was the breadth of work students could delegate. A single conversational interface can help with interpreting a prompt, planning a solution, writing code, and composing an explanation. That makes it easier to produce an apparently complete submission while doing less of the reasoning an assignment was intended to develop.

For computing education, that brings us back to why we assign implementation work. Building a data structure can teach students to reason about representations, invariants, efficiency, and failure cases. We should be able to explain which of those capabilities an assignment develops. Implementing a red-black tree is one route to learning representations and algorithmic reasoning, rather than the ultimate objective itself. We should be clear about the connection between the work students do and the understanding we expect them to carry forward. Slides, p. 29.

7. How can teaching preserve problem-solving and algorithmic thinking?

The AI-Lab is one attempt to give students practice making decisions about assistance before those decisions become habits. Students can examine a generated answer, explain its problems, compare approaches, and discuss when they would use the tool again. Reflection becomes more useful when it refers to a concrete interaction the student has just had.

I would also preserve opportunities for students to work independently and demonstrate what they understand. An assisted assignment and an independent task can serve different learning and assessment purposes within the same course. The exact balance needs to fit the subject and students. We do not yet have a universal recipe, and our study cannot supply one. What it gives us is evidence about students’ responses to a practical intervention and questions worth pursuing about their later learning. Slides, p. 30.

8. Should colleges replace conventional homework with instruction on using AI?

I would keep meaningful homework and teach students how to make informed choices about AI within the course. Getting an answer through a tool does not necessarily exercise the reasoning that an assignment is meant to develop. At the same time, an assignment can deliberately ask students to critique an output or use assistance under stated conditions.

My preference is for instructors to make these decisions with their course objectives in view. Institutions can support that work through resources, training, and clear policies, while leaving room for differences between subjects and assignments. A blanket requirement to use AI would be as difficult to justify across every task as a blanket assumption that its use serves no educational purpose. The practical question is what the student must do, explain, and understand by the end. Slides, p. 31.

9. Will improving models make an intervention about limitations obsolete?

At the seminar, I answered yes to the first part of the question: I expected tools to become better at interpreting intent, recognizing missing information, and asking users for clarification. The audience was also right that this is a moving target. A problem that exposes a weakness in one tool may be handled well by another version. The original slides included predictions about future capabilities and timelines. Those were expectations expressed in 2025, not findings of the study, and I would not treat the dates as a basis for instructional design.

Some examples will need to change. The underlying questions can remain useful: What did I ask? What assumptions did the answer make? How can I check it? What part of this task do I need to learn to do myself? Even a correct answer may arrive before a student has practiced the intended reasoning. Instructors should revisit the activities as tools change, while retaining opportunities for students to examine how assistance affects their own work. Slides, p. 31.

10. Is there evidence for the “Junior Year Wall”?

I used that phrase for a concern: students might complete early courses with substantial assistance and then struggle when later courses require foundational skills they have not developed. Our AI-Lab study did not follow students through a degree, so it does not demonstrate that such a wall exists or measure its size.

I face the same temptation in my own coursework: when I need to be able to take a problem from start to finish independently, I have to avoid delegating that reasoning to a tool. The concern was informed by observations and by the familiar risk of completing work without practicing the underlying skills. That can motivate an intervention, but it cannot establish a longitudinal effect. Testing the idea would require following students over time, measuring their independent capabilities, and considering other explanations for later difficulties. My position is that we can teach students to evaluate their use now while doing the research needed to understand those longer-term outcomes. We should be equally clear about the reason for acting and the evidence still missing. Slides, p. 32.

Related papers and teaching materials

Download the full seminar presentation (PDF, 32 pages, including speaker notes and the Q&A appendix).

The original AI-Lab framework paper, with Andres Bejarano and Chirayu Garg, describes the instructional approach. The multi-course evaluation manuscript, with Andres Bejarano, Rhianna Kuperus, and Bárbara Fagundes, contains the study discussed here; it has been revised since the seminar.

For instructors, I have shared the workshop handout, lecture notes and pre-lab materials, and Spring 2025 AI-Lab lecture slides. These are starting points to adapt to a course’s objectives and the tools students can access.

This work was supported by the Purdue Innovation Hub and conducted under Purdue IRB protocol IRB-2023-1079. Thank you to the ENE seminar audience for the questions and discussion, and to the students who participated in the research.