Summer Science Event Generates Enthusiasm For Food Science, Ag, and Computers

Categories
Services
Written by
Kevin Silverstein

Summer Science Fun!


Last  year a dedicated team of nearly 30 scientists from CFANS, MSI, UMGC and Macalester College teamed up to create and deliver a fun-filled weeklong summer science event for middle school students. Another year has come and gone, and so too has this year’s delivery of this event. But this wasn’t simply a rinse and repeat in this second year. We analyzed reactions to last year’s delivery and met monthly all year long with the goal of improving what was already a very successful endeavor.

 

What were some of the changes for the second year of this program? In a major change, we carved out 20-30 minutes each morning for Tex Ostvig, leader of CFANS’s Office of Inclusive Excellence, to provide inspirational instruction to the students on leadership development activities. Students responded remarkably well to reflective exercises on how to be a good listener, what are my core values, what are the different forms of group governance and more. These highly interactive activities served to boost students' self image and really anchored them each day as they moved to each new space for instruction. In a program that changed daily by design, Tex was a constant presence each day. And once this initial activity was done, students were emotionally ready to engage with the scientific activities that followed. Check out pictures of Tex with the students in each of the 5 days in the gallery below.

 

We also had a new sponsor for this year’s event, Forever Green Initiative. Many thanks to them for a $10,000 grant that allowed us to provide 7 students-in-need scholarships to attend, snacks, lab supplies, porta-potties, and other incidentals, along with carry-over funds from PepsiCo and Cargill from last year.

 

And once again, many thanks to all the dedicated volunteers that planned this event and delivered an incredible experience for these kids!! Your nimbleness, especially when it rained necessitating plans B and C, were amazing. There is no question in my mind that we have convinced several kids to consider college, consider science, and consider the U of M as a place they want to be in the future. Way to go!

 

Check out a sampling of the photos from this year’s event below.

 

Tex highlights: Leadership activities.

 

Students in a classroom participating in their first leadership activity skills-building exercise
Students in small groups participating in a leadership-building activity
Students in a classroom participating in a leadership activity focused on listening skills
Students listen to a presentation on leadership focused on governing styles
Group photo of participants holding University of Minnesota folders, certificates of their completion of the camp


Day 1. Food Science and Nutrition featuring fresh ice cream and cheese (voted most memorable!)

Students in grades 6-8 wearing personal protective equipment in a food lab
Student wearing personal protective equipment cutting out a cookie shape in dough made with Kernza
Students in grades 6-8 wearing personal protective equipment getting a tour of the University of Minnesota Food Science lab
A group of students in the food science laboratory interactively performing an experiment with red cabbage leaves in different pH environments


Day 2. DNA extraction, sequencing and Mutant fruit flies

Three students examining the contents of their DNA extraction kits
Students working in groups to carry out DNA extraction
A student loads DNA onto a pen-drive-sized sequencer as the other students watch
Students shine red light onto optogenetic mutant flies that perform different actions depending on their mutation profiles


Day 3. Wheat breeding, Plant Pathology, Kernza and the Conservatory 

Students in the process of threshing wheat, blowing away the chaff
Students simulating genetic crosses with ping pong balls negotiating a first-generation cross
Demonstration of the incredible length of the roots of the perennial Kernza in comparison to the annual wheat using a life-sized image on a scroll
Students tour the conservatory at the plant growth facility


Day 4. Dairy barn activities and a walk with the calves

Students weighing animal feed in preparation for feeding the cows a balanced diet
A tour of the St. Paul Campus dairy barn
Two students herding a dairy cow out the dairy barn
Lining up 3-month old dairy cow calves at the St. Paul Campus dairy barn for a “show”


Day 5. Supercomputers and computing activities

Students get a tour of the Supercomputing Institute at the University of Minnesota
An event leader presenting on using computers to link information
Small group of students sitting at tables using what their learned all week for final camp activity
Larger view of students sitting at tables using what their learned all week for final camp activity

 

 

 

 

​This activity supported in part by: MnDRIVE Global Food Ventures, University of Minnesota​

WinterTurf hackathon and 2024–2025 sensing update

Categories
Written by
Ann Piotrowski, Majid Farhadloo, and Bryan Runck

As the WinterTurf 2024–2025 data collection season comes to a close, the WinterTurf sensing nodes are being removed to make way for spring maintenance and regular golf course operations. This winter marked our largest data collection effort yet, with 75 sensing nodes deployed across the northern hemisphere at golf courses and research sites. These nodes collected 702,905 data packets and 16,713,955 sensor readings, capturing the daily changes that influence turf health during the harshest months of the year.

By refining our technology and pursuing collaborative research, we aim to equip superintendents with the knowledge they need to protect and maintain their greens throughout the winter.

Technological advancements in WinterTurf sensing

 

This season, our team introduced sensing improvements to our data collection and sensor monitoring efforts. Our new Command Execution (CommandExe) support tool for our v3 data logger enables a remote command function to check connectivity and fine-tune functionality. Additionally, our updated dashboards provide real-time diagnostics, improving our daily monitoring capabilities.

 

With every season comes challenges. Some courses experienced poor cellular signal or quality, not allowing the node to send data in real time and limiting our remote access for diagnostics. To address this issue, our system is designed to store all data locally on an internal microSD card, which we can download once the node comes back to the lab in the spring. Another challenge in winter is the limited sunlight – our system relies on incoming solar energy with a battery backup. During the darkest months, some nodes still require manual battery charging by course superintendents, ensuring continued operation in very low-light conditions.

 

Exploring data through a multidisciplinary hackathon

 

Recently, we had an exciting two-day intensive hackathon event that included researchers, data scientists, and turfgrass experts. The meeting aimed to generate research questions and uncover patterns at a fast-paced tempo using our growing and extensive dataset. The group explored questions such as:

 

  • How do CO2 accumulation rates differ between these three cover conditions: ice, impermeable covers, and impermeable covers and ice (Figure 1)?
  • Which combination of fall practices correlates most strongly with reduced winterkill damage?
  • How do light intensity levels under different covers correlate with turfgrass recovery rates?

 

A bar graph showing weekly average CO2 levels under various winter turf cover types including impermeable covers and ice.

Figure 1. Exploratory bar graph showing weekly average CO2 levels under cover types: impermeable covers, ice, both ice and impermeable covers, or other cover type. Credit: Majid Farhadloo.

 

While the hackathon was mainly exploratory, it identified new directions for future research. The collaborative meeting highlighted the value of multidisciplinary analysis in understanding complex environmental data.

 

A banner image representing WinterTurf data collection efforts on golf courses across the northern hemisphere during winter.

A banner image representing WinterTurf data collection efforts on golf courses across the northern hemisphere during winter.

 

Final thoughts

 

Our goal remains the same: to provide golf course managers with research-based knowledge and tools for winter turf management. By refining our technology and pursuing collaborative research, we aim to equip superintendents with the knowledge they need to protect and maintain their greens throughout the winter. As we reflect on another successful season, we look forward to further advancements. Stay tuned for more updates as we continue to dig into the data.

 

 

 

​This activity supported in part by: MnDRIVE Global Food Ventures, University of Minnesota

Research Working Towards Big Data Analysis

Written by
Jesse Erdmann

Researchers that are transitioning from doing analysis locally on their laptop to remote or large scale analysis often fall into a few common traps. Moving from local analysis isn’t necessarily difficult, but there are some good habits to develop that will make your transition more productive.

Establish a small test set

The set should be representative of the variety of data in the full data set if possible, but doesn’t necessarily need to produce results similar to that of the full analysis. There are two key wins that come from having a small test set predefined.

Development loop speed 

There is a temptation, especially for those who are used to working on smaller datasets, to write code and then run it against the complete dataset. However, code is almost never perfect on the first, second, or even third attempt.  If running a test takes several minutes or more that can force you to focus your attention elsewhere while you wait, reducing your efficiency and causing even longer delays. It is best to have a small, even if nonsensical, test set that can run in 10 seconds or less. Rapid iteration is key, especially during early development phases.

Validating changes before large runs

The point of this set is to establish a set of unit tests that exercise the code which can be used to verify core functionality. There is little worse than starting a long running task, getting most of the way through the task, and then encountering an error due to a typo or other simple oversight that requires starting the process over again. Having tests that check for basic functionality at the edges of expected input will save a lot of time by preventing waiting for executions that can never successfully complete.

Additionally, as new errors are encountered based on real data make sure to extend both the code and the tests to appropriately handle the new cases that were improperly handled.  This does not necessarily mean to solve data problems during execution, but when faults or exceptions occur due to type errors, etc, catch the error and log it in a way that can be presented to the user as part of a list of data cleaning tasks to perform before the next attempt. In a large dataset, only returning the first encountered error is a sure way to make the task take far too long. Instead, log each case as they are encountered, while ensuring that the program keeps going all of the way to the end generating an easy-to-understand error log.

Execution environment

One of the keys to ensuring reproducible behavior is tracking which external libraries are used in a program and more specifically which versions. In modern software development we are very dependent on others and their contributions to make our own development processes tractable. As with everything there are pros and cons to the way software is currently being developed. Further challenges and opportunities will arise as AI generated code, or AI assisted development becomes more common.

For now, ensure familiarity with the concept of semantic versioning. In brief, the first number is the major version, the next version is the minor version, and the third version is the patch version. For most purposes, when setting up an execution environment a good rule of thumb is to use the minor version of a library as the one to base an environment on. This should allow for fixes to be applied at the patch level, but more substantive changes can be adopted as needed.

Required libraries

In this case, required libraries only refer to the libraries that are directly imported and invoked by the code under development. Each required library may also include subsequent libraries, but trying to enumerate these or their versions will make building an execution environment much more challenging. The tradeoff is that these dependencies of dependencies can introduce unexpected changes.  

Packaging systems

Once a requirements list has been established, most programming environments have tools that can be used to create an environment that includes the contents of the requirements list. In Python, this can be achieved with the pip command. However, it is important to note that some libraries will be C or C++ based and require an appropriate compiler to build.

This is where a tool like conda comes into play. Where a tool like pip will install dependencies from their source code, conda maintains repositories with prebuilt binaries instead. Conda also supports multiple languages. The tradeoff is that conda adds a significant amount of disk usage to an environment.

If the intent is to rebuild an environment on every system where the code will be executed this is not much of a problem. If, however, ensuring the exact same version of required libraries are available and the environment itself will be packaged for distribution the additional gigabytes of storage can be more of a disadvantage. 

Container Images w/Docker, Apptainer, or Kubernetes

Over the last decade Docker images and other container infrastructure providers have become more popular as a way to provide a lightweight way to build a frozen, all inclusive, execution environment. This allows a researcher to deploy an identical execution environment and code from their laptop to a High Performance Computing or Cloud Computing environment.

Many CyberInfratructure providers such as the Minnesota Supercomputing Institute and the NSF’s ACCESS provide facilities for running containers from provided images using Apptainer. In cloud computing Kubernetes or other tools might be more readily available. 
 

Putting it all together

With some planning, a researcher could develop a good test suite to ensure that their code produces expected results, create a reproducible environment to the degree of their choice, and get the best of both rapid development locally using a minimal data set as well as the ability to then run the full data set on appropriately scaled hardware elsewhere. Each case will have unique circumstances that are difficult to provide a standard solution for. However, hopefully this introduction to some of the possibilities will help researchers choose a path that will help them develop a robust approach to working with big data.

 

 

 

 

​This activity supported in part by MnDRIVE Global Food Ventures, University of Minnesota

CyberTraining for Agri-Food Scientists

Services
Written by
Kevin Silverstein

GEMS’s new NSF award

Training ag and food scientists how to make the most out of supercomputers

Do you write code (maybe in R, maybe in Python)? Maybe you’re an agri-food scientist in a lab, or in a similar field in industry. Perhaps you’ve tried to write some code to get your work done, figuring you can run it on your laptop. But when you churn on that data you find after two weeks of keeping your laptop open to run the code that it’s probably not going to finish any time soon… Epiphany. You need to do this in a smarter way! Well you’re in luck – GEMS was just awarded a grant from NSF titled “Cyber Training: Pilot -- Breaking the Compute Barrier, Upskilling Agri-Food Researchers to Utilize HPC Resources” and you are exactly the audience we wish to reach with the courses we are developing.

The Team

The course material is being developed by an interdisciplinary team including:

  • Instructor Joe Axberg, who works on supercomputing infrastructure by day but is also a part-time instructor after hours in the IT Infrastructure Program in the College of Continuing and Professional Studies.
  • Project mastermind, Ali Joglekar, an agricultural economist by trade with real-world experience working with smallholder farmers in East Africa
  • Kevin Silverstein, a cofounder of GEMS Informatics and expert bioinformaticist with specialties in molecular plant-microbe interactions and large-scale informatics problems
  • Jesse Erdmann, systems architect and guiding force behind much of the inner workings of GEMS infrastructure.
  • Ben Lynch, Director of MSI, with a long track record of innovating in the scientific computing domain

Scope

The proposal identifies 7 separate course modules that each contain a significant number of concepts, representing a broad swath of content to help up-skill agri-food researchers to use HPC:

  1. Introduction to High-Performance Computing (HPC) for Agri-Food Researchers
  2. Introduction to Cloud Computing for Agri-Food Researchers
  3. Hands-On: Use HPC to Analyze More Data and Faster
  4. Leveling-Up I: Expanding HPC Skills for Agri-Food Research
  5. Computer Science for the Agri-Food Researcher
  6. Leveling-Up II: Advanced HPC Concepts
  7. High Performance Workstations and Servers for Agri-Food Researchers

It is expected that the number of modules proposed and the content within each module may change as we build out the course and modules during the development phase(s). The goal of our iterative course development process is to refine this broad set of concepts and craft cohesive modules each with a duration of 2-6 hours of lecture materials (depending on the module). Some modules will span multiple weeks. Several rounds of purposeful learner feedback is integral to our course development process. 

The course modules will also be “stackable” allowing learners to have some ability to “pick and choose” modules in line with their competencies and interests.

Figure showing timeline from course development starting August 2023, versioning and full course deliverylivery

Timeline for delivery

As the diagram indicates, we are planning to offer each HPC for Agri-Food Researchers course over a 10-12 week period of time. The proposed course structure, in line with our other GEMS Learning offerings, includes asynchronous and synchronous instructional elements. Lecture materials will be delivered via videos and readings, which students are expected to review ahead of the weekly 1.5 hour class session. This instructor-led class time is primarily dedicated to synthesizing and reviewing the content being taught. To solidify learning objectives, students will be expected to engage in regular assessments and hands-on learning exercises outside of the weekly class session. We estimate that students will be responsible for 6+ hours/week of self-directed learning activities, depending on their skill-levels.

You can influence our content!

The good news is we are still in development, so you can influence these course offerings. Please fill out this survey now to make your voice heard!

 

 

 

​This activity supported in part by MnDRIVE Global Food Ventures, University of Minnesota

Become a Data Scientist for Digital Agri

Services
Written by
Kevin A. T. Silverstein

What skills should you be building for a career as a data scientist in Digital Agriculture?

Everything we do has a spatial and temporal component. These days GIS skills are a big plus!

I often have students come to me saying “I’m passionate about the agri-food sector and I want to do data science. What skills do I need to land a satisfying job, and ultimately a career, in this area?” So I figured I’d take the opportunity to share what I typically say to them here, for the benefit of all those out there with similar goals and interests. Key items are in bold.

First off, there is so much unstructured data out there, in disparate formats and locations that you aren’t going to get far without a programming language under your belt. And in this field, that really boils down to Python and/or R. Sure, Chat-GPT can write code, but trust me, it’s not there yet.

Next, everything we do has a spatial and temporal component. These days GIS skills are a big plus! You don’t have to be a GIS expert, but you should know what a coordinate reference system and datum are, understand the limitations and practicalities of aggregation and disaggregation to different levels of resolution, and be facile with vector and raster manipulations.

Data science applied to any domain involves statistics and modeling. Moving beyond point estimates with p-values and understanding Bayesian statistics will get you far. And an understanding of databases (e.g., relational, graph, or columnar) can also be useful.

Finally, many datasets are incredibly large, so analyses often can’t be performed on your laptop. So familiarity with doing analyses on High-Performance Computing (HPC) infrastructure can be critical. These systems have mechanisms in place for you to schedule jobs to be run across multiple processors in tandem with other people’s jobs, and there are conventions and rules of etiquette for that.

Depending on the positions you have in your wish list, you may want to be sure to have either a Masters level degree or Ph.D. Masters should suffice if you’re happy to have someone else identify and devise the scope of the problems you work on. If you want to do pure R&D and define your own problems, a Ph.D. will likely be necessary.

If you're lacking some of these skills or qualifications, there are numerous data science degrees offered across the country, as well as short instructional modules such as those included in GEMS Learning.

 

 

 

 

​This activity supported in part by MnDRIVE Global Food Ventures, University of Minnesota

Building a Platform for Agri-Food Informatics

Categories
Written by
Joe Axberg

GEMS at the Intersection of Technology and Agriculture

Hi, I’m Joe Axberg, an Application Engineer at the Minnesota Supercomputing Institute (MSI).  I am part of the Development and Operations team that builds and maintains our GEMS applications and services. It is through that backend “IT” lens that I will be writing my first post on our blog site.  Others on the GEMS team will be blogging about all the great things being done with GEMS - in other words - how GEMS is being used and how it is making a difference.  In this post though, I’d like to talk a little bit about the technology that powers GEMS.

Here at the University of Minnesota, GEMS is a collaboration between the College of Food, Agricultural, and Natural Resource Sciences (CFANS) and Research Computing, including the Minnesota Supercomputer Institute (MSI) and U Spatial.  The MSI is responsible for providing a variety of high performance computing resources to researchers across the University and beyond. The kind of high performance computing power needed to solve the big problems and crunch the big data.

Utilizing high performance computing can be challenging, complex, and intimidating.  In the agri-food research space, adoption of high performance computing has historically not been very high. Agri-food researchers are not computer scientists - but neither should we need them to be.   Applications and tools should exist that lower the barrier to access high performance computing to the agri-food researcher.

Hence the development of the GEMS Informatics and GEMS Learning Platforms.

What is Informatics? 

A quick Google search revealed this definition:

“the science of processing data for storage and retrieval”

“...processing data for storage and retrieval” - that certainly does sum up what the GEMS platform does - in a perhaps over-simplified way.  At GEMS, we expand on that definition: “turning…data into actionable information for farmers, scientists, governments or companies…” (Check it out at our About Page)

These days it is all about the data and the amount of data is immense - especially in the agri-food sector.  Challenges abound in the form of sustainability, changing climate, distribution, and more.  The data is out there and there is a lot of it.

The volume and complexity of this data often means that traditional methods for gathering, cleaning, organizing, storing, and analyzing data run out of steam. 

The aim of the GEMS Informatics platform is to make the job of the agri-food researcher easier.  An application that allows the data to become actionable more quickly.  Present within GEMS are applications and services that can handle the entire lifecycle of agri-food data.

Coupled with the GEMS platform is GEMS Learning.  Provided are a series of training modules and courses for upskilling agri-food researchers in the use of technology.  GEMS Learning is constantly developing and looking at new training opportunities for agric-food researchers in the area of high performance computing.

What Powers GEMS Informatics?

At risk of sounding cliche’, it really is about the people. Please visit the GEMS website to learn more about the passionate group of researchers, professors, technologists, and administrators who make GEMS possible.

On the technical side, the GEMS platform is a thoroughly modern platform developed using some of the latest technologies and techniques.  The various components of the GEMS platform leverage both private and public cloud services. Docker containers form the basic infrastructure of the platform.  Modern web development technologies are used to implement the platform.

 

 

 

 

​This activity supported in part by MnDRIVE Global Food Ventures, University of Minnesota