What we have measured, and what we have not.
An instrument that reads tasting prose into sixteen flavour characters, where every reading carries the words that produced it. This page is the honest state of it: the corpus it runs on, what has been measured about it, what has not been, and the calibration programme proposed to change that.
What we have built
Figures read from the database, refreshed hourly.
- Bottles in the catalogue
- 493
- Tasting notes held, with provenance per note
- 470
- Readings in the catalogue
- 3,708
- Evidence contributions behind them
- 6,023
- Readings a person has ruled on by hand
- 185
- Readings awaiting a ruling
- 0
Every reading carries which note, which organisation and which phrase produced it. A reading with no quotable phrase is refused rather than stored.
Most readings are confirmed by the reader itself rather than by hand. That is a measured position, not a shortcut: across every proposal a person has ruled on, they agreed with the machine 96.8% of the time, and a review step agreed with that often is re-typing an answer rather than adding judgement. The same rate is printed on the method paper, which is versioned and dated; the figure above is computed from the database as this page renders. One case is excluded and stays excluded: a reading that asserts a character is ABSENT is never confirmed on a single source without a person saying so.
What we have measured
Measured 10 September 2026. These are dated figures about the system as it was that day, each with its denominator; they are not re-derived silently.
The reader is not deterministic
Fourteen real notes, five identical readings each, replicated on a second seventy calls. One character in five is found on some readings and not others (20%, then 18% on the replication). Of the characters found every time, intensity moved on 41–45% and the quoted phrase moved on 43–48%. A single reading misses about one character in ten that the note supports. What it changes: every reading stored before this was one draw recorded as though it were the reading. Notes are now read several times, the union kept, and how often each character came back recorded beside it.
Stated confidence predicts recurrence — and is itself a draw
The reader’s own confidence figure, checked against measured recurrence for the first time: AUC 0.798, scored on the single value the queue stores. Every character under 0.50 flickered; from 0.70 up, 0–7% did. All of the discriminating power sits below 0.60. And the figure itself moves on 93% of stable characters between readings. What it changes: it works as a sort key for a reviewer’s attention, never as a fact about a reading, and never as a floor — a floor would remove exactly the rows the number can speak about.
The scale has five positions and the evidence uses two
Of 8,321 intensity readings, 85% are a 2 or a 3; 0 and 4 together are 4.3%. The same shape appears in the confidence column: 466 individual readings took twelve distinct values, and five round numbers covered 85% of them. What it changes: more rings would be invented precision. Whether five positions are already too many is a question for a tasting panel, not a design choice.
The vocabulary is uncontrolled
Across 4,452 distinct phrases in 471 notes, 28% of the frequent short phrases land in more than one character — “toffee” thirty times in one, twenty-one in another — concentrated in three adjacent cask characters. The reader responds to the words beside a phrase rather than to the note’s overall lean, shown under controlled conditions, so the boundary is real and describable. The lexicon that would settle it holds 66 terms, 3% of the words in use, and feeds nothing. What it changes: it has to be authored from tasting, with boundary rules stated, rather than grown.
What we have not
The instrument has never been compared to a person tasting the same whisky, nor to the chemistry of the liquid. Everything above measured it against itself, which is the first time anything was measured at all — but self-consistency is not accuracy.
Seven magnitudes in the system are in effect and set by judgement: the source weights, the section weights and boost, the drift cap, the consensus count, the aggregated ceiling and the matchable threshold. Four priors — the mapping from chemistry and declared cask to perceived character — are deliberately empty, and an automated test fails if anyone fills them without a measurement behind it. They are slots waiting for the study below.
The agreement rate above is thinner than it looks, and it is the number most likely to be quoted back at us. It rests on 185 rulings, not on a sample drawn at random: they are the readings a person happened to reach, which are plausibly the easier ones. And a raw percentage does not correct for the agreement you would expect by chance — the standard corrections for that are well established and have not been applied here. It is re-measured against every ruling made from here, so a fall would be visible rather than inferred; until there are more of them it is an indication and not a reliability figure.
The programme
Proposed to Food Science & Technology at Nanyang Technological University, with Heriot-Watt’s International Centre for Brewing and Distilling and the Scotch Whisky Research Institute, September 2026. Not funded. Not underway.
- The sensory anchor set. Forty-eight whiskies chosen by declared type, a trained panel of eight, blind. Per-character reliability at panel and single-taster level, pre-registered. The first ground truth the instrument did not produce itself.
- The chemistry arm. The same whiskies by GC-MS, with olfactometry on a subset. Which characters are predictable from the compound profile, at what precision — and which are not, reported as a finding rather than forced to fit. The four empty priors get their first values, from data.
- The reader against the panel. The panel’s own notes, read by the instrument, compared to the panel’s scores. Whether prose can be read into measurement at all.
- Calibration and publication. Every constant calibrated against the above, or held empty with the reason. The reliability figures published, including the characters where the taxonomy failed.
Named now so it is not argued for later: a character the panel cannot measure reliably means the sixteen are fifteen; tasters using three positions means the scale collapses; a Singapore panel and a Scottish panel disagreeing systematically means the taxonomy does not travel unchanged. Each of those is the programme producing an answer.