Purdue researchers create an AI-based VR system that turns natural language into plant science analytics

09-18-2026

VR BioTalk lets users analyze data without any programming skills, simply by moving their head and issuing verbal commands, such as "show me plants taller than 10% of the average height and show the graph of its distribution." The users can generate hypotheses and analyze large datasets through simple dialogue.

VR BioTalk lets users analyze data without any programming skills, simply by moving their head and issuing verbal commands, such as "show me plants taller than 10% of the average height and show the graph of its distribution." The users can generate hypotheses and analyze large datasets through simple dialogue.

 

Modern sensors gather vast amounts of data that plant scientists can use to understand how crops grow. Field-gathered data are large and unstructured, so they often go unused. A single automated field scanner can produce 10 terabytes of data a day and more than a petabyte over a growing season; a volume so large that analyzing the data usually requires advanced programming skills that most biologists don’t have, or biological knowledge that is not common for computer scientists. Now, a team of researchers from Purdue University’s Department of Computer Science and the University of Arizona has built a way to sidestep coding entirely: put on a VR headset, walk through a virtual version of your field, and simply ask questions.

Sponsored by the NSF, the team introduced the system, called VR BioTalk, which is a hands-free, immersive, voice-controlled visual analytics tool for plant phenotyping data. Wearing an off-the-shelf VR headset, a user can stand inside a full 3D reconstruction of a scanned field and issue plain-language commands, for example, “Show me all leaves smaller than the average and calculate their leaf area index,” and watch the system carry out the request and display the answer in seconds.

“Our goal was to remove the programming knowledge requirement that stands between plant scientists and their own data,” said Jorge Vazquez, a Ph.D. student in Computer Science at Purdue and the paper’s first author. “Instead of writing a script and waiting, you put on the headset, look at the plant data, and talk. The system analyzes what you are asking, does the computation, and shows you the result while you explore the field.”

“The challenge we are trying to solve is the domain gap,” said Professor Bedrich Benes, the lead of the project. “Efficient analysis of large phenotyping datasets requires computational pipelines and programming that most life scientists don't have access to. Meanwhile, computer scientists who can build those pipelines often cannot formulate the biological hypotheses that make the data worth analyzing in the first place. VR-BioTalk aims to let the plant scientist stay in the driver’s seat.”

The system works in three stages. First, raw 3D point clouds of the field are processed to isolate individual plants and extract traits such as height, radius, leaf area, leaf area index, number of leaves, and stem dimensions. Second, when a user speaks, VR-BioTalk converts their voice to text with a speech-recognition model and then uses an AI-based language model to translate that text into a specific analytics command. Third, a visualization engine executes the command and updates the 3D scene, highlighting plants, computing statistics, or generating charts that float alongside the field.

“The point-cloud 3D data acquired with scanners is particularly difficult to visualize,” says Voicu Popescu, the co-principal investigator on the project leading the visualization effort. “The points lack connectivity, so the surfaces have to be reconstructed. We have done this by leveraging the coherent trajectory of the scanner, which tells you which 3D points are connected to which.”

The work grew out of data collected by the University of Arizona’s Field Scanner, a large gantry-mounted phenotyping platform that scans experimental sorghum plots with high-resolution cameras and laser scanners, but is applicable to any field data, for example from drones.

“We collect far more data than we can analyze. Every season the scanner captures things nobody looks at, not because they aren’t interesting but because getting at them takes too much time. So, when the turnaround on a question is a week, you only ask the questions you are already fairly certain about. When it’s a few seconds, you start asking the ones we were really curious about, and that is where the interesting finds come from,” said Duke Pauli, co-principal investigator at the University of Arizona who oversees the Field Scanner. 

To test whether real users could work with the system, the team ran a study with participants drawn from Agronomy, Environmental Science, Plant Biology, and Computer Science. Participants used VR-BioTalk to answer questions about relationships among plant traits and then completed standardized surveys measuring usability, cognitive load, and their sense of presence. The results indicated the system was highly usable, engaging, and easy to use, even for participants with no programming background. The system also understands different accents. 

“VR-BioTalk has the potential to support research, but also education,” said Professor Alejandra Magana, who leads the educational effort. “We see an opportunity to make complex scientific data more accessible to learners by allowing them to interact with data naturally, through conversation, exploration, and visualization. Instead of making programming expertise a prerequisite for engaging with large datasets, students can focus on developing the scientific reasoning skills needed to ask questions, recognize patterns, test ideas, and generate hypotheses. Ongoing studies will help us improve human-computer interaction and manage cognitive load.”

The team’s next step is building true digital twins - virtual plants that can simulate plant responses to different environments that were never actually experienced by the plant. In the next phase, the petabytes of data sitting untouched on servers could become readily accessible, and a tool like VR BioTalk will be key to unlocking the data’s potential.

In addition to Vazquez, Benes, Popescu, Pauli and Magana, the team includes Purdue’s Shuwen Yang, a Ph.D. student in Computer Science, and Yiqun Zhang, a Ph.D. student in the School of Applied and Creative Computing. University of Arizona co-authors include Jeffrey Demieville of the School of Plant Sciences, Brennan Huppenthal of the Department of Computer Science, and Nirav Merchant, director of the university’s Data Science Institute. This work was supported by the National Science Foundation, the USDA National Institute of Food and Agriculture, and the U.S. Department of Energy’s Biological and Environmental Research program.

 

About the Department of Computer Science at Purdue University

Founded in 1962, the Department of Computer Science was created to be an innovative base of knowledge in the emerging field of computing as the first degree-awarding program in the United States. The department continues to advance the computer science industry through research. U.S. News & World Report ranks the department No. 14 and No. 15 overall in undergraduate and graduate computer science, respectively. Graduates of the program are able to solve complex and challenging problems in many fields. Our consistent success in an ever-changing landscape is reflected in the record undergraduate enrollment, increased faculty hiring, innovative research projects, and the creation of new academic programs. Learn more at cs.purdue.edu.  

Last Updated: Sep 22, 2026 4:01 PM