Summer Science for Youth – Expanding outreach with Bread Science

Written by
Kevin Silverstein

Who would have ever guessed that 6 years after the pandemic middle-schoolers would still be crazed about sourdough!! 

Sure enough, on February 9 at 8 am, registration for our new “Science of Bread Making” 4-day summer science event went online with 16 slots open, and within 30 minutes we were full with a 17-person waiting list. So we opened another 16 slots, creating an AM group and a PM group. The pressure was on for our team of 30 volunteers to create an unforgettable summer science experience June 15-18, 2026 for this crowd – the first offering of this new event on top of the third delivery of the Food Ag & U experience that followed the week afterwards.

Summer Science Event Generates Enthusiasm For Food Science, Ag, and Computers

Categories
Services
Written by
Kevin Silverstein

Summer Science Fun!


Last  year a dedicated team of nearly 30 scientists from CFANS, MSI, UMGC and Macalester College teamed up to create and deliver a fun-filled weeklong summer science event for middle school students. Another year has come and gone, and so too has this year’s delivery of this event. But this wasn’t simply a rinse and repeat in this second year. We analyzed reactions to last year’s delivery and met monthly all year long with the goal of improving what was already a very successful endeavor.

 

What were some of the changes for the second year of this program? In a major change, we carved out 20-30 minutes each morning for Tex Ostvig, leader of CFANS’s Office of Inclusive Excellence, to provide inspirational instruction to the students on leadership development activities. Students responded remarkably well to reflective exercises on how to be a good listener, what are my core values, what are the different forms of group governance and more. These highly interactive activities served to boost students' self image and really anchored them each day as they moved to each new space for instruction. In a program that changed daily by design, Tex was a constant presence each day. And once this initial activity was done, students were emotionally ready to engage with the scientific activities that followed. Check out pictures of Tex with the students in each of the 5 days in the gallery below.

 

We also had a new sponsor for this year’s event, Forever Green Initiative. Many thanks to them for a $10,000 grant that allowed us to provide 7 students-in-need scholarships to attend, snacks, lab supplies, porta-potties, and other incidentals, along with carry-over funds from PepsiCo and Cargill from last year.

 

And once again, many thanks to all the dedicated volunteers that planned this event and delivered an incredible experience for these kids!! Your nimbleness, especially when it rained necessitating plans B and C, were amazing. There is no question in my mind that we have convinced several kids to consider college, consider science, and consider the U of M as a place they want to be in the future. Way to go!

 

Check out a sampling of the photos from this year’s event below.

 

Tex highlights: Leadership activities.

 

Students in a classroom participating in their first leadership activity skills-building exercise
Students in small groups participating in a leadership-building activity
Students in a classroom participating in a leadership activity focused on listening skills
Students listen to a presentation on leadership focused on governing styles
Group photo of participants holding University of Minnesota folders, certificates of their completion of the camp


Day 1. Food Science and Nutrition featuring fresh ice cream and cheese (voted most memorable!)

Students in grades 6-8 wearing personal protective equipment in a food lab
Student wearing personal protective equipment cutting out a cookie shape in dough made with Kernza
Students in grades 6-8 wearing personal protective equipment getting a tour of the University of Minnesota Food Science lab
A group of students in the food science laboratory interactively performing an experiment with red cabbage leaves in different pH environments


Day 2. DNA extraction, sequencing and Mutant fruit flies

Three students examining the contents of their DNA extraction kits
Students working in groups to carry out DNA extraction
A student loads DNA onto a pen-drive-sized sequencer as the other students watch
Students shine red light onto optogenetic mutant flies that perform different actions depending on their mutation profiles


Day 3. Wheat breeding, Plant Pathology, Kernza and the Conservatory 

Students in the process of threshing wheat, blowing away the chaff
Students simulating genetic crosses with ping pong balls negotiating a first-generation cross
Demonstration of the incredible length of the roots of the perennial Kernza in comparison to the annual wheat using a life-sized image on a scroll
Students tour the conservatory at the plant growth facility


Day 4. Dairy barn activities and a walk with the calves

Students weighing animal feed in preparation for feeding the cows a balanced diet
A tour of the St. Paul Campus dairy barn
Two students herding a dairy cow out the dairy barn
Lining up 3-month old dairy cow calves at the St. Paul Campus dairy barn for a “show”


Day 5. Supercomputers and computing activities

Students get a tour of the Supercomputing Institute at the University of Minnesota
An event leader presenting on using computers to link information
Small group of students sitting at tables using what their learned all week for final camp activity
Larger view of students sitting at tables using what their learned all week for final camp activity

 

 

 

 

​This activity supported in part by: MnDRIVE Global Food Ventures, University of Minnesota​

Pedigree analysis tools illuminate ancestry of over 8.5M wheat varieties in CIMMYT’s international nursery

Categories
Written by
Kevin Silverstein

The International Wheat and Maize Improvement Center (CIMMYT) is the world’s primary source of breeding material for wheat and corn (maize). Founded in 1943 and vaulted to international recognition in the 1960’s and 70’s, partly through the work of U Minnesota alum and Nobel prize Laureate Norman Borlaug, CIMMYT was the original model for the centers that now comprise the CGIAR. Over the years, CIMMYT has accumulated a large database of pedigree, trait and passport information for millions of wheat genotypes that include finished varieties, germplasm accessions, and (advanced) breeding lines. This information is critical for breeding programs to track the inheritance of desired traits (e.g., yield, disease resistance, drought tolerance, baking quality) for selections that are made for subsequent generations. The coefficient of parentage (COP), also known as the inbreeding coefficient, is a particularly useful metric when considering any two genotypes as potential parents for breeding. Formally this quantity indicates the likelihood that for any gene, the copies that occur in both genotypes are descended from a common ancestor.


GEMS Informatics worked with CIMMYT to analyze a subset of their large database, the 2012-2024 spring and durum wheat international nursery dataset. This dataset included 8.5 Million genotypes, 11.8 Million genotype aliases (e.g., internal cross names, commercial release names, abbreviations), and 18 Million genotypic relationships. We set out to accomplish 3 tasks:

 

  1. Build a python code repository with a scalable infrastructure. GEMS Informatics designed an approach using SQLAlchemy and SQLite that can accommodate 100’s of millions of genotypes and their relationships. This effort was successful and took 6 months. All CIMMYT’s data can be ingested in just 1 hour on a contemporary laptop.
  2. Resolve naming discrepancies identified in the CIMMYT data. GEMS staff analyzed all common_name and cross_name designations among the 8.5 million genotypes and putatively identified 544 pairs of genotypes in CIMMYT’s genebank that may be duplicative, and hence require consolidation in their database. It is a testament to the care that CIMMYT staff have employed that there were only 544 “typos” among the names for these 8.5 million genotypes. Examples include common typographical errors (e.g., C0723595 and CO723595; II53.546 and 1153.546), punctuation variants (e.g., 4715D(5B) and 47-1-5D (5B); DARTS-IMPERIAL and DART´S IMPERIAL), compound word variants (e.g., PLAN ALTO and PLANALTO; YANG MAI 6 and YANGMAI 6), language variants (ALGERIAN and ALGERIEN; FEDERATION and FEDERACION) and misspellings (e.g., AEGILOP UMBELLULATA and ARGILOPS UMBELLULATA; ATALANTA and ATLANTA; AUBAKOMUGI and AOBAKOMUGHI).
  3. Provide harmonized pedigree datasets and query capabilities via an API to CIMMYT. All of the following questions can now be answered via API queries to the database: What are the parents of any wheat genotype? What are the pedigree entries at any arbitrary level? (level 1 = parents; level 2 = grandparents; level 3 = great grandparents; …) What is the full recursive pedigree for any genotype? (ideally traced back recursively to landraces, if possible) What are the known aliases for any wheat genotype? What is the matrix of pairwise COP values for any pair or list of genotypes?

Future Work

Cleaning and harmonizing wheat pedigrees worldwide. Previously PedTools, developed by GEMS in 2017 could support modest pedigree sizes involving thousands of genotypes, initially targeted to wheat pedigrees for US and Canadian varieties. This work enabled  the GEMS team to further enhance their PedTools infrastructure so that it can scale to collections with genotype counts numbering 10 million - 100 million. This makes it suitable to expand to CIMMYT’s full wheat genebanks as well as publicly accessible repos like GrainGenes and GRIN. Further, PedTools has the ability to match genotypes across organizations so it may serve to unify the wheat pedigree collections across countries, CIMMYT and public repositories, mapping accessions at each center to each other. 

Additional crops: The original incarnation of PedTools was used to harmonize ~10,000 soybean pedigrees for a UMN soybean breeding project. Our soybean breeding collaborator is currently digitizing decades of old printed variety breeding information. So we plan to revisit that harmonizing effort with the new version of PedTools. Soybean breeders don’t use the Purdey notation (variety 1 / variety 2 // variety 3) for pedigrees, but instead utilize an arithmetic notation (e.g., ((variety 1 x variety 2) x variety 3)). With this in mind the underlying architecture of PedTools has been designed to accommodate a plugin of any custom set of rules to parse variety names and pedigrees for specific crop communities. In this manner, in the future we can write a new parser to ingest a new format of pedigrees and the rest of the PedTools machinery remains unchanged since the internal representation of varieties and their relationships is the same. 
With this potential for expandability, we plan to engage pedigree data curators for other crops at various CGIAR Centers and elsewhere to help standardize their pedigrees and thus streamline and accelerate trait discovery and varietal development efforts.

 


Photo credit: A. Morgounov/CIMMYT.


Funding. The activities described here were conducted with support from the Government of Mexico and Minnesota State Government MnDRIVE funding made available to GEMS Informatics.

 

Sensing Below-Ground Environments to Better Predict Potato Disease Threats

Written by
Senait D. Senay and Philip Pardey

Potatoes are a pervasive staple and specialty crop the world over, but so too are the pests and diseases that affect potato yields, tuber quality and farmer profitability. However, like the productive part of the crop itself, many potato diseases develop below ground, incurring costly damage well before the farmer becomes aware of the problem. Getting a better handle on the spatial extent, depth and temporal variation of soil temperature, moisture and other environmental variables that affect the development of potato diseases is key to modeling the field-scale risks posed by these threats. This is especially so if the aim is to model disease development in ways that provide farmers with actionable (real-time) information to mitigate or manage the crop production and profitability outcomes of these diseases.


Verticillium wilt is a long standing scourge of potato farmers. In related work, we estimate this particular soil borne fungi is a threat to almost 72% of the world’s potato growing area. V. wilt infections first become evident above ground when the plant’s lower leaves wither and die. Symptoms progress upwards until the entire plant yellows and wilts. The disease causes early senescence of the plant, which results in economically significant yield losses and tuber discoloration. In some instances, costly fumigation can be an effective mitigation strategy, while long rotations (3 years or more) with other crops can reduce the inoculum load of this long-lived disease at a particular site.

 
Creating fit-for-purpose biotic threat models that reveal the potential risks associated with V. wilt and other crop diseases at field scale and beyond is a core research focus of the GEMS Biotic Threat Analytics Lab. Pest risk prediction models and timely access to the targeted information products they enable helps farmers and others prioritize disease intervention on local (and neighboring) farms, informs a host of post-farm supply-chain decisions that rely on prospective crop production outcomes, feeds valuable information into early warning systems, and informs crop breeding strategies.

Digging Deeper into Above- and Below-Ground Environmental Data


Appropriately scaled environmental data both above and below ground data are required to informatively model the field-level risks posed by V. wilt (and other crop pests and diseases). While there are  several relevant gridded environmental datasets to hand, most are at coarser resolutions that extend well beyond the area extent of a typical potato field or farm. Moreover, these datasets often lack relevant below ground variables (e.g., soil moisture and temperature, at variable depths) that in combination with other variables are required to develop and deploy actionable pest prediction models of soil-born biotic threats. To rectify these two shortcomings, we turned to our GEMS Sensing team to provide real-time sensing of the needed environmental data. 
 

To best align our environmental sensing efforts with incidence and severity information on V. wilt, we also paired up with Dr. Ashish Ranjan’s Lab in the University of Minnesota’s (UMN) Department of Plant Pathology. Ashish conducts extensive V. wilt trials at UMN’s potato disease nursery located at the U’s Sand Plains Research Center in Becker, Minnesota. 


Siting Sensors to Reap the Biggest Predictive Bang for the Buck!


In 2023 we ran a test deployment of two GEMS sensing systems in the V. wilt resistance screening blocks at Becker, MN. Each system was configured with 3 above ground sensors (temperature, barometric pressure, and relative humidity) and 5 below ground sensors (soil moisture, temperature, permittivity, bulk soil electrical conductivity, and porosity). The above ground sensors were deployed in 3 replicates, and the below ground sensors at 3 depths. Our statistical assessment of these real-time data indicated that one set of above ground sensors coupled with below ground sensors at two depths yielded the optimal sensor configuration. 
 

GEMS Sensor, above ground sensing node


For the 2024 growing season we scaled up our sensing efforts to 17 sensing stations, each with 3 above ground sensors and 5 below ground sensors. Fifteen sensor systems were deployed in the research plots where select potato varieties are screened for V. wilt by the Ashish Lab, plus 2 sensing systems for benchmarking in the (disease free) potato breeding plots at Becker managed by Dr. Laura Shannon in UMN’s Department of Horticultural Science. 


The precise placement of each sensing system was informed by an environmental profiling exercise prior to field deployment. First we digitized the boundaries of each of the 16 blocks used in the V. wilt screening nursery then overlaid that on gridded data we accessed from GEMS Exchange on 10 variables of potential relevance for disease risk modeling; including elevation, slope, available water storage (AWS) and soil organic carbon stock estimate (both at 3 depths throughout the rootzone). Our aim was to sense as much environmental variation from within the study area as possible in the process of generating our targeted below (and above) ground environmental variables.
 

Gridded environmental data layers used to inform sensor deployment


The deployed location of each sensor is marked by the red dot in image #3, where in this instance each disease nursery block is overlaid on just one (i.e., elevation) of the 10 environmental variables we used to select a site for each sensor. 
 

Locations identified for sensor placement based on the environmental variability analysis work done on the study area.


The wealth of high-resolution, real-time (every 15 minutes) environmental data generated by this deployment is now being analyzed and integrated with correspondingly geo-tagged V wilt field data from the Ashish Lab. Field-scale predictive pest models are also being prototyped drawing directly on these novel, environment-linked-to-disease data sets to both develop and ground truth our modeling results. Working with our industry partners, PepsiCo, we look forward to further refining and then geographically scaling up the deployment of these predictive models to provide real-time, fit-for-purpose insights into dealing with this (and other) pesky potato diseases.     

 

Verticillium wilt Image credit: Utah State University

 

 

​This activity supported in part by MnDRIVE Global Food Ventures, University of Minnesota

Breeding Better Cassava for Climate Resilience

Categories
Written by
Thomas Kono, Sean Festemaker, Kevin Silverstein, Nathan Carlson, Phil Pardey

Cassava, a root crop, is a critical staple food crop planted on 32 million acres worldwide. It is widely grown throughout sub-Saharan Africa, notably in Nigeria, DR Congo, Ghana, Angola, and Mozambique, but with large acreages in Vietnam, Brazil, Indonesia and India. Besides being a critical source of calories and dietary fiber for many poorer households throughout Africa it is also a very versatile crop, serving as an important source of animal feed and starch with uses in foods, glues, biodegradable products and drugs. It has a global market value of $48.7 billion and is also a priority crop for the Vision for Adapted Crops and Soils (VACS) program, led by the U.S. State Department, which aims to create resilient food systems in Africa by growing nutritious, climate-adapted crops in healthy soils.

Unlocking Cassava’s Potential for Food Security and Climate Resilience

Conventional breeding is a painstaking laborious process that takes many crop generations to develop the first in a stream of varieties that adapt to ever-evolving market and climate conditions. To speed up this process, the International Center for Tropical Agriculture (CIAT) sequenced thousands of cassava varieties to reveal genetic markers of adaptive and harmful traits. GEMS colleagues Nathan Carlson, Tom Kono, and Kevin Silverstein in the Minnesota Supercomputing Institute (MSI), developed a queryable genomics database to enhance cassava breeding and improvement efforts at CIAT and elsewhere. To do so they drew on the whole genome resequencing data spanning 3,673 accessions of cassava provided by CIAT and identified short DNA sequence variants among them–totalying over 9 million sequence variants! More specifically, the MSI team identified nonsynonymous variants, a subset of the sequence variants that change the amino acid sequence of the plants’ proteins from the reference genome sequence. The functional impact of the nonsynonymous variants was then predicted using a sequence constraint model called BAD_Mutations (Chun and Fay 2009, Kono et al. 2018) to identify sequence variants with potential impact on cassava trait variation.

cassava root

Tackling Deleterious Mutations

CIAT breeders, led by Sean Fenstemaker, are excited at the possibilities these data provide for them. Deleterious mutations can significantly reduce crop yield and quality. CIAT’s breeding program now incorporates BAD_Mutations, an innovative SNP annotation tool designed to identify harmful genetic variants in cassava. This tool employs a likelihood ratio test based on alignments of publicly available angiosperm genomes, allowing for improved detection of deleterious mutations. 

Why Use BAD_Mutations?

Jonathon Newby, Cassava Program Leader, CIAT noted that  “While smallholder cassava farmers are faced with a range of new threats, there are also many untapped opportunities for this formerly neglected crop to address food security and nutrition,and still be a globally competitive product in industrial and food application. The cassava variant database addresses challenges and explores new opportunities for cassava breeding to unlock this potential. By comparing genetic variants with their ancestral origins, it provides insights into diversity and traits conserved in plants, aiding in the identification of key genetic variations. The database also supports molecular marker development, parent selection, and breeding strategy refinement.”

Enhancing Breeding Strategies

BAD_Mutations helps identify and select against genetic variants that negatively impact traits of interest. By using this tool, breeders can enhance phenotypic variation, leading to the development of robust and high-yielding cassava varieties. This genomic precision is vital for adapting to changing environmental conditions and meeting market demands. Additionally, breeders may use BAD_Mutations as a strategy for in silico validation of trait-linked markers, further ensuring the accuracy and effectiveness of molecular breeding efforts.

CIAT is using BAD_Mutations in tandem with other advanced technologies such as flower-inducing and doubled haploid techniques. These methods, combined with the University of Minnesota’s genomic tools, facilitate backcrossing-based trait introgression and systematic exploration of heterosis, significantly improving breeding efficiency of CIAT and its partners.

While cassava genetics was the focus of this project, the resulting queryable database framework has much broader applications within agriculture. Identification of genetic variants of potentially large effect is a technique that is useful for general crop and animal improvement, especially for complex traits (e.g., yield) which are typically under the influence of many genetic loci. Construction of an efficient, query-ready database allows for the genetic variation data to be rapidly assessed with standard input and output formats, making it easier for researchers to interpret the data.

This project demonstrates how cross-institution collaborations can accelerate applied research efforts. By partnering, CIAT and UMN crunched through a very large set of genomics data (requiring continuous compute cycles on hundreds of supercomputer processors for a month) into a format that can easily be used by geneticists and breeders to improve an important staple crop.

Global Collaboration for Food Security

These advanced genomic tools are pivotal in enhancing crop resilience and productivity by addressing the challenge of deleterious genetic mutations. CIAT invites researchers and global partners to collaborate in using this new cassava variant effect database. There is a real urgency to accelerate crop breeding to address global food security and poverty reduction concerns in the face of consequential changes in climate worldwide. Novel partnerships that pool complementary resources are key to making significant strides that result in timely and climate-resilient improvements in cassava and other crops.

 

 

 

 

 

​This activity supported in part by MnDRIVE Global Food Ventures, University of Minnesota
 

Food Agriculture & U Summer Science Camp

Written by
Kevin Silverstein

Fun for everyone at interdisciplinary Food Agriculture and U summer science camp

Imagine going to an agriculturally-themed summer camp where you got to see and taste protein bars, puffed cereal, ice cream and cheese being made at an industrial grade facility; did experiments where you pulled iron out of fortified breakfast cereal; extracted and sequenced DNA from your food; used a supercomputer to decipher that DNA; simulate a wheat breeding experiment and identify diseased plants in the field; stick your arm inside a dairy cow’s stomach, and lead your own dairy calf in a field by the barns! Wow, that’s a lot to experience in one week. But that’s precisely what 15 brave, inquisitive and incredibly bright 11, 12 and 13-year olds did at the inaugural launch of the ‘Food, Agriculture and U’ summer science camp, held in collaboration with the University of Minnesota’s Youth Programs June 24-28, 2024.

 

The camp was initially conceived by Kevin Silverstein at the Minnesota Supercomputing Institute and at GEMS Informatics, George Annor at the CFANS Department of Food Science and Nutrition, and Getiria Onsongo, at Macalester College and GEMS. But they soon found that people from all over the University loved the idea and joined the effort, volunteering their time. In all, more than 25 professionals ended up making significant contributions to the camp, nearly all interacting directly with the kids on one or more of the 5 days. Special thanks to co-organizing leaders Shea Anderson and Daryl Gohl at the UMGC, Emily Conley and Becca Hall at Agronomy & Plant Genetics and Plant Pathology, and Tony Seykora and Isaac Salfer in Animal Science for their extensive planning and staff recruitment.

 

In recruiting students for the camp, significant effort was made to reach the Native American community. An outstanding liaison at each of two institutions, Migizi in Minneapolis and the American Indian Magnet School in St. Paul, helped us recruit students that normally miss out on these opportunities. We also attended two powwows and appealed to families there directly.

 

Corporate sponsors were also amenable to the concept. We are very grateful for contributions from PepsiCo, Cargill, and Fairbault Foods for allowing us to provide a free camp experience for 6 of our 15 students. And they also allowed us to purchase healthy snacks (which the kids greatly appreciated and kept talking about) for all participants each day. Thanks also go to NSF and ACCESS for an allocation that enabled the students to decode their DNA sequence on the Nations' supercomputing infrastructure, and to the Digital Science Initiative for helping us to manage donations. 

 

Please check out the gallery which has 4 photos from each day, highlighting the diversity of activities. Miraculously, even though this was the first time giving this camp, all 5 days went forward smoothly and were a great hit. There are a few tweaks to make next year for sure, but the response was overwhelmingly positive and we are all delighted!

Day 1

Exploring Food Science and Nutrition

 

Camp students in Food Science Lab
Camp students in Food Science Lab

 

Day 2

DNA sequencing, extraction and mutants

 

plant DNA extraction in a lab
students looking at plant DNA properties

 

camp students reviewing optogenetics
plant DNA extraction in a lab

 

Day 3

Linking ag data via supercomputers

 

computer exercise
students touring the Minnesota Supercomputer

 

database review
students in conference room

 

Day 4

Visiting agricultural fields on campus with breeders and plant doctors

 

student wheat threshing
students in farm field

 

students in a field
students in a field

 

Day 5

Out in the barns with the dairy cattle

 

 

students with animal feed

 

 

 

 

 

 

 

 

 

group of students walking a dairy calf

 

student with dairy cow

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

​This activity supported in part by MnDRIVE Global Food Ventures, University of Minnesota

Relaunching GEMS Informatics Exchange APIs

Services
Written by
Kevin Silverstein and Phil Pardey

APIs: Now well-documented and much easier to use

One of the big frustrations in using computing to solve large, multidisciplinary challenges is managing data sets from different disciplines. Often, you have to go to each individual site and download the entire dataset. Then you have to parse out the subset of data fields you want within the geographies, spatial resolutions and time periods you care about. It is still the case that a few groups provide their data in the form of an Application Programmer Interface (API), where the data are served in a structured form with clear metadata documentation. Data can be sliced and diced how you like, selecting subsets of geography, time, and variables of interest. Once you sign up and obtain an API key, it just takes a few lines of code in Python or R to establish a connection and query at will!

GEMS has been building out a portfolio of APIs since 2021 across a range of useful datasets seeking to span the full Genetics x Environment x Management x Socioeconomic data landscape. Those who tried GEMS Exchange before will know that we used to have a middle layer managed by RapidAPI. Users found that cumbersome and confusing, so we are now using our own Apache APISIX server within our own web pages to serve you your key and monitor usage. We’re confident that your experience will be super easy this time around. Let’s get you started!

First, check out which APIs might interest you at our GEMS Exchange page. To obtain your API key simply click here for key. (Note you will need to have a Globus.org account, which is free – or you can connect via your academic institution, Google account, or ORCID). Once you know which APIs interest you, explore our collection of Jupyter notebooks in Github that give you practical guidance on how to use them. Many of the APIs we offer use the GEMS Grid which help ensure they are interoperable. And the GEMS Grid itself has recently been made open source, so you can place your own data sets on the Grid and interoperate with the community.

We are always happy to hear of useful datasets that could be added to the GEMS gridded collection in Exchange, so by all means reach out with suggestions or queries here.

 

 

 

 

​This activity supported in part by MnDRIVE Global Food Ventures, University of Minnesota

The GEMS Informatics Grid Goes Open Source

Written by
Kevin Silverstein

We are delighted to announce that we have just released the GEMS Grid code library, where the code is under the open source Apache 2.0 license, which allows anyone to use the code for commercial or non-commercial purposes – you simply need to provide attribution to GEMS Informatics when you use or modify it.

Just before we at GEMS Informatics started developing Application Programmer Interfaces (APIs) in agriculture for GEMS Exchange, GEMS geospatial expert Jeffery Thompson worked with others in the GEMS team and colleagues at NSIDC to develop the GEMS Grid, a hierarchical discrete global gridding system. This Grid has allowed us to provide data sets at different resolutions ranging from 36 km to 1 m, and still have them remain functionally interoperable. The interoperability is possible because we have written the code to allow users to project data onto the grid, aggregate data to coarser resolutions, and, notably, also disaggregate data to finer resolutions. The latter operation is ordinarily a difficult problem, but is made easier, as I discuss below, since we enable the users of our code to thoughtfully address it in a standardized, replicable way.

Many problems in agriculture (e.g., understanding the spatial location of crop production) require equal area parcels of land to do proper calculations. Working with strict lat-lon coordinates won’t suffice as areas near the equator are significantly different in size as areas near the poles. The GEMS Grid preserves equal-area assumptions as it divides land, so you can do these calculations with confidence, and preserve aggregation-disaggregation consistency in the data, even if you are not a GIS expert.

Pictorial description of the 5 options for disaggregation on the GEMS Grid

So let’s look at the 5 options for disaggregation that GEMS geospatial developer Olena Boiko included in the GEMS grid toolbox, schematically described in the figure she developed above.

Option 1. Value transference. 
In this case, if you were to subdivide a 3 km2 resolution grid cell into 9 x 1 km2 cells, this option would be appropriate for any value that is deemed roughly constant throughout the area applied. Examples would be rainfall in inches or grain yield in bushels / acre.

Option 2. Even value division. 
Sometimes the quantity measured in a cell represents a cumulative value for the area in which it is reported. In this case, if the parent cell is homogenous, then splitting it up into 9 equal-area pieces would require that you divide the value in each equivalent cell by a factor of 9. Examples where this selection makes sense include grain production in bushels, crop acreage, and population.

Option 3. Value transference with a mask. 
This one is similar to Option 1 except we are no longer making the assumption that the distribution of values in the parent cell is spatially homogeneous. For example suppose you were measuring grain yield, but you knew that 3 of your nine cells had buildings occupying them (see white areas in the Figure). In this case you only transfer your values to 6 remaining cells (colored peach) that have arable land. Cells are binary with this option (i.e., either allowed a value or not).

Option 4. Even value division with a mask. 
Analogously, you can mask out cells in the value division case when you know that your daughters cells are not all equal. This is just like the case in Option 3, except you divide your parent-cell value evenly by the number of viable daughter cells. In this pictorial example, there are 6 viable daughter cells, so each gets a value of 900/6 = 150. This would be appropriate if you were computing grain production in bushels and you had a total value that needed to be split up across the 6 arable daughter parcels.

Option 5. Flexible division with a mask. 
This scenario is the most flexible, and allows the user to create a master mask with arbitrary weights at each daughter cell. It allows you to block off daughter cells entirely, and prescribe the relative weights of all remaining daughter cells. This is ideal for situations where you are allocating crop distributions and you want to avoid certain land use features (e.g., lakes, forests, housing) and probabilistically distribute the remaining crop areas (e.g., with higher probability near soils with a high SSURGO National Commodity Crop Productivity Index).

I’m confident these flexible disaggregation tools will provide much easier, more accurate, and replicable solutions for your particular spatial analytic problem. So please give them a try!

 

 

 

 

​This activity supported in part by MnDRIVE Global Food Ventures, University of Minnesota

CyberTraining for Agri-Food Scientists

Services
Written by
Kevin Silverstein

GEMS’s new NSF award

Training ag and food scientists how to make the most out of supercomputers

Do you write code (maybe in R, maybe in Python)? Maybe you’re an agri-food scientist in a lab, or in a similar field in industry. Perhaps you’ve tried to write some code to get your work done, figuring you can run it on your laptop. But when you churn on that data you find after two weeks of keeping your laptop open to run the code that it’s probably not going to finish any time soon… Epiphany. You need to do this in a smarter way! Well you’re in luck – GEMS was just awarded a grant from NSF titled “Cyber Training: Pilot -- Breaking the Compute Barrier, Upskilling Agri-Food Researchers to Utilize HPC Resources” and you are exactly the audience we wish to reach with the courses we are developing.

The Team

The course material is being developed by an interdisciplinary team including:

  • Instructor Joe Axberg, who works on supercomputing infrastructure by day but is also a part-time instructor after hours in the IT Infrastructure Program in the College of Continuing and Professional Studies.
  • Project mastermind, Ali Joglekar, an agricultural economist by trade with real-world experience working with smallholder farmers in East Africa
  • Kevin Silverstein, a cofounder of GEMS Informatics and expert bioinformaticist with specialties in molecular plant-microbe interactions and large-scale informatics problems
  • Jesse Erdmann, systems architect and guiding force behind much of the inner workings of GEMS infrastructure.
  • Ben Lynch, Director of MSI, with a long track record of innovating in the scientific computing domain

Scope

The proposal identifies 7 separate course modules that each contain a significant number of concepts, representing a broad swath of content to help up-skill agri-food researchers to use HPC:

  1. Introduction to High-Performance Computing (HPC) for Agri-Food Researchers
  2. Introduction to Cloud Computing for Agri-Food Researchers
  3. Hands-On: Use HPC to Analyze More Data and Faster
  4. Leveling-Up I: Expanding HPC Skills for Agri-Food Research
  5. Computer Science for the Agri-Food Researcher
  6. Leveling-Up II: Advanced HPC Concepts
  7. High Performance Workstations and Servers for Agri-Food Researchers

It is expected that the number of modules proposed and the content within each module may change as we build out the course and modules during the development phase(s). The goal of our iterative course development process is to refine this broad set of concepts and craft cohesive modules each with a duration of 2-6 hours of lecture materials (depending on the module). Some modules will span multiple weeks. Several rounds of purposeful learner feedback is integral to our course development process. 

The course modules will also be “stackable” allowing learners to have some ability to “pick and choose” modules in line with their competencies and interests.

Figure showing timeline from course development starting August 2023, versioning and full course deliverylivery

Timeline for delivery

As the diagram indicates, we are planning to offer each HPC for Agri-Food Researchers course over a 10-12 week period of time. The proposed course structure, in line with our other GEMS Learning offerings, includes asynchronous and synchronous instructional elements. Lecture materials will be delivered via videos and readings, which students are expected to review ahead of the weekly 1.5 hour class session. This instructor-led class time is primarily dedicated to synthesizing and reviewing the content being taught. To solidify learning objectives, students will be expected to engage in regular assessments and hands-on learning exercises outside of the weekly class session. We estimate that students will be responsible for 6+ hours/week of self-directed learning activities, depending on their skill-levels.

You can influence our content!

The good news is we are still in development, so you can influence these course offerings. Please fill out this survey now to make your voice heard!

 

 

 

​This activity supported in part by MnDRIVE Global Food Ventures, University of Minnesota

Become a Data Scientist for Digital Agri

Services
Written by
Kevin A. T. Silverstein

What skills should you be building for a career as a data scientist in Digital Agriculture?

Everything we do has a spatial and temporal component. These days GIS skills are a big plus!

I often have students come to me saying “I’m passionate about the agri-food sector and I want to do data science. What skills do I need to land a satisfying job, and ultimately a career, in this area?” So I figured I’d take the opportunity to share what I typically say to them here, for the benefit of all those out there with similar goals and interests. Key items are in bold.

First off, there is so much unstructured data out there, in disparate formats and locations that you aren’t going to get far without a programming language under your belt. And in this field, that really boils down to Python and/or R. Sure, Chat-GPT can write code, but trust me, it’s not there yet.

Next, everything we do has a spatial and temporal component. These days GIS skills are a big plus! You don’t have to be a GIS expert, but you should know what a coordinate reference system and datum are, understand the limitations and practicalities of aggregation and disaggregation to different levels of resolution, and be facile with vector and raster manipulations.

Data science applied to any domain involves statistics and modeling. Moving beyond point estimates with p-values and understanding Bayesian statistics will get you far. And an understanding of databases (e.g., relational, graph, or columnar) can also be useful.

Finally, many datasets are incredibly large, so analyses often can’t be performed on your laptop. So familiarity with doing analyses on High-Performance Computing (HPC) infrastructure can be critical. These systems have mechanisms in place for you to schedule jobs to be run across multiple processors in tandem with other people’s jobs, and there are conventions and rules of etiquette for that.

Depending on the positions you have in your wish list, you may want to be sure to have either a Masters level degree or Ph.D. Masters should suffice if you’re happy to have someone else identify and devise the scope of the problems you work on. If you want to do pure R&D and define your own problems, a Ph.D. will likely be necessary.

If you're lacking some of these skills or qualifications, there are numerous data science degrees offered across the country, as well as short instructional modules such as those included in GEMS Learning.

 

 

 

 

​This activity supported in part by MnDRIVE Global Food Ventures, University of Minnesota