Agricultural Data: Big or Small?

Limited observations require agricultural AI to combine existing knowledge, uncertainty, and more strategic data collection.

Comparison of iterative model development. The traditional process cycles between collecting data and building or improving a model. The guided process adds a third step: characterize prediction uncertainty and underrepresented conditions, then use those findings to collect new data strategically.

Agricultural datasets may contain extensive satellite imagery, sensor measurements, or other variables while still having relatively few high-quality observations for training and evaluating AI models. This Perspectives article examines that mismatch using agricultural examples, including a tillage-mapping project with approximately 600 ground-truth observations. Kirti Rajagopalan argues that AI models working with limited agricultural data should incorporate existing scientific knowledge rather than attempting to rediscover known relationships. Researchers should also quantify uncertainty for individual predictions and identify underrepresented crops, soils, production systems, environments, and years. Together, uncertainty and data diversity can guide researchers toward the observations most likely to improve a model. The goal is not merely to make agricultural datasets larger, but to develop AI approaches that use limited data and existing agricultural knowledge more effectively.

This publication is part of an archive and may not meet current digital accessibility standards. CSANR is working to improve digital accessibility of all materials. If you need this content in an alternative format, please contact csanr@wsu.edu.

Authors

Rajagopalan, K.

Related Products

Related Project

Year Published

2026

Area of Focus

Agricultural Technology

Topics

Community Engaged Research and Production Systems

Collaborator

Funding Source