Skip to content
CiderHQ
Search

Cider science

How a trained sensory panel works

How do experts decide what a cider tastes like?

In short

By treating a group of trained people as a measuring instrument. Assessors are screened for basic sensory acuity, trained until they use the same words for the same sensations, and calibrated against physical reference standards so that a score of 4 for bitterness means the same thing to each of them.

They then assess samples blind, in randomised order, in individual booths under controlled lighting and temperature, with replication so that each assessor’s consistency can be checked statistically.

What a panel establishes is intensity and difference. It does not establish whether a cider is good, whether people will buy it, or why a compound is present — those are separate questions requiring consumer panels, market research or chemistry.

Building the instrument

Screening comes first. Candidates are tested for the ability to detect the basic tastes at low concentration, to rank intensities correctly, and to discriminate aromas. This removes people with specific insensitivities and, equally importantly, identifies which assessors are reliable rather than which are enthusiastic.

Training then builds a shared lexicon. The panel works through a set of reference standards — a solution of a known concentration of a bitter compound, a spiked sample carrying a specific ester, a deliberately oxidised control — and agrees on the term for each and on where it sits on the scale. This is the step that makes the panel an instrument rather than a group of opinions, and it takes considerable time.

Calibration is then maintained. Panels drift, individual assessors drift, and both are checked by periodically re-presenting known samples and examining whether the scores still agree. A panel that has not been recalibrated is producing data of unknown quality.

The two things panels are used for

The triangle test deserves particular mention because it is the workhorse of the field and because its logic is instructive. Three samples are presented, two identical and one different, and the assessor identifies the odd one. Guessing gives a one-in-three success rate, so the statistics simply ask whether the observed success rate exceeds chance by enough to be convincing. It is a robust test precisely because it asks for a discrimination rather than a description.

Descriptive analysis and difference testing answer different questions.
MethodQuestion answeredWhat it produces
Descriptive analysisHow intense is each attribute in this sample?A profile across the lexicon, with statistics across assessors and replicates
Triangle testCan these two samples be told apart at all?A yes or no with a stated confidence, from the proportion of correct identifications
Duo-trio and paired comparisonWhich of these differs, or which is more intense on a named attribute?A directional difference with a confidence level
Threshold determinationAt what concentration does this compound become detectable?A group threshold, with the spread between individuals

The controls, and why each exists

What a panel cannot establish

A trained panel cannot tell you whether a cider is good. Preference is not an attribute of the drink, and trained assessors are explicitly trained away from expressing it, because a person calibrated to detect a compound at low concentration is by construction unrepresentative of ordinary drinkers. Preference is measured on separate, untrained consumer panels, with large numbers and no training.

It cannot establish cause. A panel that finds two ciders differ in phenolic character has established a difference, not the reason for it; that requires analysis, and often it requires an experiment in which one variable was deliberately changed.

It cannot tell you what an ordinary drinker will notice. A panel’s thresholds are lower than a consumer population’s, so a difference that a panel detects reliably may be invisible in the market.

And it cannot settle disputes about authenticity, tradition or value. Those are not sensory questions, and a panel result presented as though it answered one is being misused.

Competition judging is a different activityCompetition judges assess against a written style specification and do render a verdict, which is what they are for. That is a legitimate exercise and it is not descriptive analysis: the judges are usually not calibrated against reference standards, the samples are usually not replicated, and the output is a ranking rather than a measurement.

Also answered on this page

Questions this page covers, so you can tell at a glance whether it is the one you want.

Related

What people ask next

Questions readers ask about the things this page mentions. Each one goes to the section that answers it rather than to a page written to receive the question.

Sources

What this page rests on. Where a source is marked as registered rather than read, CiderHQ is recording that the body is authoritative on the subject without claiming to have worked through the document itself. See our evidence policy for what each state means.