T R U F F A I R E
← Blog
Agriculture6 min read

Confidence Scores: When to Trust a Diagnosis

A confidence score is not self-esteem. It states how far the evidence narrowed the field — and it should change what you do next.

T

Truffaire

20 August 2026

Every automated crop diagnosis comes with a confidence figure, and most people treat it as decoration. The condition is named, so the condition is the answer, and the percentage beside it is something the software prints.

That reading loses the most useful piece of information in the report. Confidence is not the system expressing self-esteem. It is a statement about how far the evidence narrowed the field of possible conditions — and a diagnosis that narrowed it partway calls for a different action than one that narrowed it fully.

Used properly, the score determines what you do next. Ignored, it makes a tentative answer and a certain one look identical.

What the number is measuring

The system is not choosing between "right" and "wrong". It is distributing plausibility across the conditions consistent with what it can see.

High confidence means the evidence was distinctive enough that one condition fits substantially better than the alternatives. Moderate confidence means several remain plausible and one leads. Low confidence means the evidence did not separate them much at all.

The important consequence: low confidence is rarely a statement that the crop has an exotic disease. It is usually a statement that the evidence was thin — a photograph too far away, a single plant part, symptoms in an early stage that look like several things at once.

Which is good news, because thin evidence is the one input the farmer controls, as set out in what a diagnostic photograph needs to show.

Why an honest low score is worth more than a confident guess

There is a design choice underneath this, and it is worth naming.

A system can be built to always return a single confident-looking answer. It will be right often and it will be wrong invisibly, because nothing in its output distinguishes the cases where it was guessing from the cases where it was sure.

A system that reports uncertainty gives up the appearance of authority and gains something more useful: the ability to say reshoot before you spend. In an agricultural context where the downstream action costs money, labour and a treatment window, that distinction is the difference between a tool and a liability.

This is the same argument made in a different domain in what "AI-powered" should mean in business software — that a system's willingness to state its limits is a quality signal rather than a weakness.

What each band should trigger

The bands are conventions rather than physical constants, and the point is that each should lead somewhere different.

High confidence. The evidence was distinctive. Proceed to the economic question — whether the treatment costs less than the expected loss, which is the subject of deciding on treatment before you spend. Confidence in the identification does not by itself justify spending.

Moderate confidence. One condition leads, others remain plausible. The correct first response is almost always to improve the evidence — more plant parts, closer images, both leaf surfaces — rather than to act on the leading answer. Where the treatments for the plausible conditions differ materially, this step is not optional.

Low confidence. The evidence did not separate the field. Reshoot properly, and if a second attempt with good capture still returns low confidence, the case has moved beyond what a photograph can resolve. That is when a laboratory referral is the right next step rather than a failure.

The asymmetry that governs all three: reshooting costs minutes, and treating the wrong condition costs the input, the labour and the window in which the actual problem continues to spread.

Where confidence and severity are frequently confused

These are two different numbers and conflating them produces bad decisions in both directions.

Confidence is how sure the system is about what it is.

Severity is how far advanced it is, which determines how quickly you must act.

A high-confidence, low-severity result means you know what it is and have time. A low-confidence, high-severity result is the uncomfortable one — something is progressing quickly and the identification is uncertain. That combination argues for urgency in improving the evidence, including a same-day referral, rather than for treating on a guess because the situation feels pressing.

Acting fast on an uncertain diagnosis is the most expensive available response to that scenario, and it is the most common.

What the score cannot account for

Confidence describes the fit between the evidence and the known conditions. Several things sit outside it.

A condition the system has not been built to recognise may return a plausible-looking match to something it does know. Two conditions present simultaneously — common in the field, and genuinely difficult — may produce a moderate score that does not indicate why. And environmental damage that mimics disease can present convincingly.

None of these are arguments against the score. They are arguments for treating it as one input alongside what the farmer already knows about the field, the weather and what happened last season — which is the standing position in how ARCORA works.

Frequently asked questions

What percentage is high enough to act on?

There is no universal threshold, because the stakes differ. Where the treatments for the plausible alternatives are similar and inexpensive, a moderate score may be actionable. Where they diverge or the input is costly, improve the evidence first.

Why does the same plant sometimes give different scores?

Because the evidence differs. A closer image, a different plant part or better light changes what the system can see. That is the mechanism working correctly, not inconsistency.

Does low confidence mean the system failed?

No. It means the evidence did not separate the possibilities, which is a fact about the photograph far more often than about the disease.

Should I treat while waiting for a lab result?

That is an agronomic judgement involving severity, spread rate and cost, and it is one to take with local expertise rather than from a report. The report's job is to tell you the identification is unresolved.

Do scores improve over time?

Systems improve as they encounter more verified field cases, which is precisely why the record built with FPO partners matters — the reasoning behind one FPO, one plant.

Where to start

Next time a diagnosis returns anything below high confidence, do not act on it. Reshoot — closer, more plant parts, both leaf surfaces — and compare the two reports.

The change in the score between the two is the clearest demonstration available of what the number actually measures, and it will change how you read every report afterwards.

The other figures on the report — severity, economic impact, the treatment protocol — are covered in reading a crop diagnosis report.

To discuss ARCORA deployment through an FPO, get in touch.

More in Agriculture