Perspectives

Practical insights and opinions from agriculture and natural resources experts—brought to you by the Center for Sustaining Agriculture and Natural Resources.

Agricultural Data: Big or Small?

By Kirti Rajagopalan, Biological Systems Engineering and CSANR Faculty Leadership Team

Agriculture may be a big-data domain, but many specific AI applications depend on relatively few high-quality observations. Making AI useful in these settings means combining limited data with existing agricultural knowledge and using uncertainty to guide what we collect next.

Agriculture is often characterized as a “big data” domain in the world of artificial intelligence (AI). In some contexts, that is certainly true. Satellite imagery, genomics, sensors, and modern agricultural equipment can generate enormous amounts of data. But for many real agricultural AI problems, the data are not particularly big at all.

This disconnect has become increasingly apparent to me through our recent projects and conversations with computer scientists involved in the AgAID AI Institute. As AI applications proliferate, it is easy to take the availability of big data for granted, but that assumption does not fit many agricultural applications.

“Big data” depends on perspective

What we mean by “big data” depends very much on who we are and where we are standing.

In a recent project using satellite imagery to predict tillage practices, it took us three years to collect ~600 high-quality ground-truth observations. From an agricultural research perspective, that is a substantial dataset. Each observation required significant effort to collect. We can link those observations to an enormous amount of satellite imagery and derive many additional variables, but ultimately we still have about 600 examples where we know what we think happened on the ground.

The Washington Soil Health Initiative’s State of the Soils work provides another example. Collecting soil-health measurements from over 1,000 locations across Washington was a massive undertaking involving field sampling, laboratory analyses, coordination with producers, and substantial investments of time and resources. From the perspective of agricultural research, 1,000 well-characterized locations is a large and valuable dataset.

From the perspective of modern AI, 600 or 1,000 observations is small.

Small data in practice

The tillage mapping project offers one example of how researchers are working with limited, costly agricultural data and using uncertainty to guide new data collection.

That difference in perspective matters. Many of the highly visible recent AI advancements were made in a different data world with extraordinary amounts of data. Computer-vision models can learn from millions of images. Large language models learn from enormous collections of text. Many of our expectations about what AI can do—and many of the methods we now bring to other disciplines—have emerged from this world of massive datasets.

Then we bring those expectations and methods to an agricultural problem where three years of work produced 600 ground-truth observations. There is a disconnect. And that disconnect becomes larger as our questions become more specific.

The response cannot always be to collect more data. Of course, we should continue collecting agricultural data. Larger, better, and more diverse datasets will be enormously valuable. But many agricultural observations are expensive and slow to produce. Another growing season takes another year. Another soil-health observation requires another sample and laboratory analysis. Another ground-truth tillage observation requires determining what actually happened in that field.

AI for the realities of agricultural data

If small datasets are the reality for many agricultural problems, then AI needs to meet agriculture where it is. What does that mean in practice?

Agricultural science has history and memory. AI should use it.

Agricultural data do not exist in a vacuum. A field has memory. Its condition today reflects previous crops, tillage, weather, fertilizer applications, residue management, irrigation, and other decisions and processes operating over years or decades.

Agricultural science has history and memory too. We already know a great deal about soils, crops, water, weather, plant physiology, and management. We have decades of experiments, process-based models, physical relationships, and knowledge accumulated by scientists, agricultural professionals, and producers.

Yet in applying AI to agricultural problems, we can too easily approach model development as though none of this history exists. We give a model 600 observations and ask it to discover the relationships for itself.

That may make sense when millions of training examples are available. It makes much less sense when each observation is expensive and our entire dataset took years to assemble.

AI for small agricultural datasets needs better ways to combine what we observe with what we already know. That might mean incorporating scientific knowledge into model structure, using existing models and relationships, transferring information learned in other places or datasets, or constraining models using what we know to be biologically or physically reasonable.

The specific methods will differ among applications. The broader principle is simple:

We should not ask small datasets to rediscover agricultural knowledge that already exists.

Uncertainty and diversity need to drive how we develop and evaluate models and collect data

When data are limited, knowing what a model does not know becomes especially important. This requires uncertainty at the level of each individual prediction—for example, an estimate of uncertainty for the tillage prediction made for each field. It is not enough to know how well a model performs overall. We also need to know whether the model has high confidence in its prediction for one field and very little confidence in its prediction for another.

In “small data” settings, uncertainty therefore cannot be an afterthought—a statistic calculated after the model has been built or simply a measure of average model performance. Uncertainty for individual predictions needs to be integrated into how we develop, evaluate, and use AI models.

That uncertainty can also help us decide where our next observations would be most valuable. If obtaining another 100 ground-truth tillage observations is expensive, fields where predictions have high uncertainty are logical places to look for new information.

But uncertainty alone is not enough. We also need to collect data that are both informative and diverse, representing different crops, soils, production systems, management practices, environments, and years. Together, uncertainty and diversity answer two complementary questions: Where does the model know the least? What new observations would broaden what the model knows?

Rather than treating data collection and AI as separate activities—first build a dataset and then train a model—we can make them part of an iterative process:

collect data → build the model → quantify uncertainty for each prediction → identify uncertain and underrepresented conditions → strategically collect new data → improve the model.

This approach is especially important when every additional observation is costly. The goal is not simply to make a small dataset bigger. It is to make each new observation add as much new information as possible.

Get Perspective

Receive emails with CSANR news and Perspectives.

Moving agricultural AI forward 

As AI applications proliferate, we need to pay more attention to a basic reality that can easily be overlooked: for many of the specific agricultural questions we want AI to help answer, relevant high-quality observations will remain limited and very costly to obtain.

Small data should therefore not be viewed simply as a temporary limitation that will disappear as agriculture collects more data. For many agricultural questions, it is the setting in which AI will need to operate. Meeting that challenge requires movement in both directions. Agricultural AI needs to make better use of approaches that already exist for learning from limited data, incorporating prior knowledge, quantifying uncertainty, transferring knowledge from models trained on larger datasets, and guiding new data collection. At the same time, continued advances in AI are needed to make these approaches more powerful, accessible, and practical for complex agricultural settings.

There is good reason for optimism. The AgAID Institute and other AI institutes focused on advancing AI for agricultural applications have brought AI researchers and agricultural scientists together, creating an important space for the initial development and testing of ideas that better reflect the realities of agricultural data. But this work needs to be taken further. We need to move these ideas from individual projects and demonstrations toward AI approaches that can be more broadly adopted across agricultural research and applications.

The opportunity is not simply to make agricultural datasets bigger, but to continue advancing—and adopting—AI approaches that can make the most of the data and agricultural knowledge we actually have.

Kirti Rajagopalan
Associate Professor
Biological Systems Engineering
Washington State University
kirtir@wsu.edu

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *