Independent evaluation

What researchers and scientists say.

GroundTruth runs on the same retrieval and citation infrastructure behind ExtensionBot, the Extension Foundation's public agricultural chatbot. Here's what independent researchers, a multi-industry AI benchmark, and Extension weed scientists found when they tested ExtensionBot against real questions.

Multi-industry AI benchmark

Ranked among the top agriculture-specific models in the first AgriBench results

The Extension Foundation is a founding member of the AgriBench Consortium, housed at the University of Illinois Center for Digital Agriculture alongside Bayer Crop Science, John Deere, Microsoft, and Digital Green, with support from the Gates Foundation and USDA-NIFA. The consortium's first published leaderboard scored ExtensionBot's answers blind against 26 other chatbots, including general-purpose models like ChatGPT and Gemini.

#2 of 6agriculture-specific chatbots
#6 of 27all chatbots, incl. ChatGPT & Gemini
#1 / #4ag-specific / overall, in horticulture
What was measured

Narrative answer quality across agricultural topics, scored blind by the AgriBench Consortium against 26 other chatbots — nothing about the questions was tailored to ExtensionBot.

What wasn't measured

This round of AgriBench didn't score citations or hallucination rate — the two things ExtensionBot is built around. That's the strength this benchmark didn't get to credit.

Why it matters

Even without credit for sourcing, it placed ahead of most agriculture-specific tools and competitive with general-purpose models built at far larger scale.

"These results validate the competitiveness of ExtensionBot relative to major commercial and foundation-model systems, particularly in areas where Extension has strong, research-based content." — Extension Foundation, Connect, December 2025
Field-expert review · weed science

The only chatbot that consistently showed its sources

Four practicing Extension weed scientists tested ExtensionBot against ChatGPT, DeepSeek, and Google Gemini on 24 realistic farmer questions spanning integrated weed management, herbicide selection, resistance, and weed biology.

80%top score of 4 bots, herbicide resistance
better on herbicide-specific management
Only botthat consistently linked its sources
What was measured

24 real-world weed management questions — some simple, some multi-step, some with the misspellings a real farmer would type — scored excellent to poor by four Extension weed scientists.

Where it led

Herbicide resistance and herbicide-specific management — the two categories where a confidently wrong answer costs a grower an entire season.

What went uncredited elsewhere

ChatGPT and DeepSeek rarely cited sources at all. When researchers checked an uncited claim from DeepSeek on herbicide cross-resistance, it was simply wrong.

"Only ExtensionBot consistently linked to the sources of its information, allowing our testers to learn more about a topic, as well as verify the bot's accuracy." — Emily Unglesbee, Grow IWM, March 2026
Peer-reviewed research · Animal Frontiers, Oxford University Press

Asked the same question 1,000 times, ChatGPT gave 1,000 different answers

Researchers from Texas A&M AgriLife Extension and Oklahoma State University tested three OpenAI models on a basic stocking-rate question — asked identically 1,000 times per model, for two different counties. Depending on which GPT version answered, accuracy on the exact same question swung from under 1% to over 99%.

0.2–99.8%GPT accuracy range, same question
360,000+Extension documents in the corpus
Editablewrong answers get corrected, not repeated
The test

"Is 5 acres per cow/calf pair a sufficient stocking rate?" — asked 1,000 times per model, for a semi-arid Texas county and a wetter Florida county with very different correct answers.

The finding

GPT-4 and GPT-4o essentially inverted each other's answers between model versions — both delivered with the same fluent confidence, only one actually consistent with reality.

The contrast

ExtensionBot answered correctly for the region it had data on, and plainly lacked coverage for the region it didn't — rather than guessing either way.

"Specialized AI platforms trained on region-specific extension publications can offer more accurate and context-relevant advice to producers compared to generalized AI tools like ChatGPT." — Prestegaard-Wilson & Vitale, Animal Frontiers, January 2025
In the press
"This is not an effort to replace people at all, but a way to broaden our reach and be able to handle more volume of questions than we could handle otherwise." — Jayson Lusk, Dean of Agriculture, Oklahoma State University, Brownfield Ag News, May 2024

ExtensionBot is backed by a partnership of more than 20 land-grant universities, the Extension Foundation, and the USDA — the same network GroundTruth's retrieval layer draws from.

Want to see how it holds up against your own use case?

A scoped pilot is the fastest way to know if this fits — measured against your own questions, not a published benchmark.

Start a pilot