Home PublicationsData Innovators5 Q’s with Sebastian Schultheiss, CEO of Computomics

5 Q’s with Sebastian Schultheiss, CEO of Computomics

by David Kertai

The Center for Data Innovation recently spoke with Sebastian Schultheiss, CEO of Computomics, a Germany-based company that uses machine learning to help plant growers predict crop performance and make informed breeding decisions. Schultheiss explained how Computomics combines genetics, field trials, and environmental data to identify promising plants, evaluate predictions, and accelerate crop development.

David Kertai: What sets Computomics’ approach apart from other companies in the plant breeding industry?

Sebastian Schultheiss: Most genomic prediction systems, which use genetic information to estimate how well a plant will perform in a given environment, rely on linear mixed machine learning models. These statistical models estimate how individual genetic markers, which are identifiable locations in a plant’s DNA, contribute to a trait such as yield. Although this approach works, it can miss complex relationships between genetic markers and interactions between a plant’s genetics and its growing environment.

To address this issue, our platform, xSeedScore®, uses nonlinear machine learning, an approach that identifies complex patterns and relationships that do not follow a simple, straight-line pattern, such as how combinations of genetic markers influence crop yield or how a plant’s genetics interact with drought or temperature to affect its growth. We also treat the growing environment as a central part of the prediction rather than something to average out. Water availability, temperature, soil conditions, and the growing season all affect crop performance.

By accounting for these factors, our models can better predict how a plant might perform in a location or climate that the breeding program has not tested before. This helps breeders identify locations and growing conditions where promising plants are most likely to thrive, allowing them to focus field trials on environments that will provide the most useful information about crop performance.

Kertai: What data does Computomics use to predict how a crop will perform?

Schultheiss: We combine three main types of data. First, we use genotype data, which describes a plant’s genetic makeup. This includes genetic marker data or whole-genome sequencing, which reads an organism’s entire DNA, as well as pedigree information that records a plant’s ancestry and helps us understand relationships between breeding lines.

Second, we use phenotype data, which records the observable characteristics of a plant. This includes historical field-trial measurements of yield, disease resistance, crop quality, and other traits that breeders want to improve. We draw on measurements collected across different years and locations whenever those records are available.

Third, we use environmental and management data, including weather patterns, soil characteristics, water availability, and farming practices at each trial site and during each growing season. This information is especially important because a yield measurement without information about the conditions that produced it tells us little about how that plant might perform elsewhere. For example, a crop may produce less grain because of drought rather than because of its genetics. Environmental records help us distinguish these effects and make more useful predictions.

Kertai: How do you know when a model’s prediction is reliable enough to guide a growing decision?

Schultheiss: We test our models under conditions that reflect the decisions growers actually face. A model can appear accurate if researchers randomly withhold a few plants from a dataset because the model may have already learned from closely related plants grown under similar conditions. Instead, we test whether the model can predict plants the breeding program has never tested or performed reliably in a growing season it has not seen before.

To do this, we withhold entire growing locations or years of data and check whether the model can still identify the strongest candidates. These approaches, known as leave-one-environment-out and leave-one-year-out validation, test whether predictions remain useful when the model encounters conditions it did not see during training.

We also examine calibration, which measures whether a model’s confidence matches how often its predictions prove correct. A breeder selecting the top ten percent of a population does not need the model to rank every plant perfectly. The machine learning model needs to identify promising candidates reliably while flagging predictions with greater uncertainty. This allows breeders to test uncertain candidates in the field rather than discard them prematurely. 

Kertai: How does Computomics protect proprietary plant data while still generating useful predictions?

Schultheiss: A breeding company’s germplasm, meaning its collection of genetic material used to develop new varieties and often reflecting years of its own breeding work, represents years or even decades of investment. Protecting that material is therefore a priority from the start.

We process customer data on our own infrastructure under data-processing agreements and keep each customer’s data separate. We train models for individual customers and usually tailor them to specific breeding programs within each company. We never pool genetic material across clients. This separation also ensures that our predictions remain useful to each customer. For example, a model trained on genetic material from multiple companies might recommend crossing two promising breeding lines even though a customer has no access or legal right to use one of them.

Additionally, we also improve our models using information that does not expose a customer’s proprietary genetics, such as public reference genomes, environmental data, and modeling techniques we develop ourselves. Every breeding program can benefit from these improvements while its genetic data remains separate from other customers’ data.

Kertai: What is the biggest misconception about using machine learning models in plant breeding?

Schultheiss: There are three common misconceptions. The first is that machine learning outright replaces field trials. In reality, our predictions depend on data from plants that breeders have grown and measured. Field trials provide the evidence our models need to learn how genetics and growing conditions affect crop performance.

The second is that machine learning-based genomic prediction changes a plant’s DNA. It does not. Genomic prediction analyzes genetic variation that already exists within a breeding population and estimates which plants are most promising for future breeding. Growers still develop new plants by crossing selected parents, while our models help them decide which candidates deserve further testing. Our models can narrow millions of potential candidates to a manageable group that breeders can test more thoroughly.

The third misconception is that breeding programs need perfectly organized data before they can benefit from machine learning models. In practice, we can often work with imperfect records, including missing measurements, inconsistent trait definitions, and renamed field sites. We can account for some gaps and inconsistencies when analyzing historical data, but breeders cannot recover environmental information they never recorded, such as weather conditions at a specific site during a particular growing season. Programs can therefore begin generating useful predictions from the data they have while improving their data collection practices over time.

You may also like

Show Buttons
Hide Buttons