News

Beyond the model: Evaluating AI agricultural advisory systems so they work in the field

Agricultural advisory services are increasingly adopting generative AI (gen AI) systems, including tools based on large language models (LLMs) such as chatbots, to provide farmers with tailored information on everything from how to manage pests to changes in commodity prices.

Man, left foreground, holds phone, screen facing camera. Second man sits behind counter. Shelves with supplies in background.
  • Artificial Intelligence
  • advisory services

By Josué KpodoMay 18, 2026

Key takeaways

 

Evaluating AI farmer advisory systems requires assessing real-world usefulness, usability, and trust—factors that matter as much as technical scores.

Assessments should happen at three levels: models, systems, and processes all shape farmer outcomes.

Equity should be measured, not assumed; inclusive benchmarking reveals who benefits—and who is left out.

Agricultural advisory services are increasingly adopting generative AI (gen AI) systems, including tools based on large language models (LLMs) such as chatbots, to provide farmers with tailored information on everything from how to manage pests to changes in commodity prices. However, developers of these tools face many challenges, including the need to function seamlessly in local languages and in different geographical areas and contexts.

Given these issues, developers must reliably ensure that AI tools perform safely and effectively in diverse real-world settings. A key approach is the benchmarking process. An AI benchmark is a standardized evaluation framework for comparing models across specific tasks (e.g., answering questions) and metrics (e.g., accuracy). For LLMs, a benchmark serves as a common testbed for assessing capabilities such as reasoning, factual knowledge, following instructions, robustness, and safety.

Currently, most LLM benchmarking efforts focus on achieving high scores in technical statistics such as accuracy and correctness. While useful, this approach represents only a fraction of what makes agricultural advisory successful. Advisory systems must not only be agronomically sound and context-specific, but also easy to understand for users with varying literacy levels. Factors like usability, linguistic diversity, and trust are often ignored in model-centric evaluations, yet they dictate whether a farmer actually uses the service.

The consensus among development practitioners is shifting: benchmarking must move beyond isolated model tests toward collaborative approaches that consider model behavior, system performance, and governance to ensure more trustworthy and reliable advisory tools for farmers.

Recognizing the gaps and potential for change, the AGX AI community—which aims to accelerate the development of responsible AI for small-scale producers in Africa and Asia—is bringing together researchers and practitioners to assess how best to evaluate these systems. On November 6, 2025, some AGX AI community members met during the IFPRI webinar session, Benchmarking LLMs for Agricultural Advisory: Insights from a Global Community of Practice. This post reviews key insights from that meeting.

Read More Here