Suppose you want to teach a language model to read product listings — to look at a shirt and extract, reliably, that it is a V-neck, cotton, in Spanish. Fine, one shirt. Now do it for thousands of product types, hundreds of attributes per type, in several languages, and you need training labels for all of it. The combinations run to millions of annotations, and at that scale, Amazon’s researchers write in a new paper, human labeling becomes “prohibitively costly.” This is the recurring problem of the AI economy: the robots are cheap, but teaching them is expensive, because the teachers are people.
The obvious shortcut is to have language models generate the labels synthetically — recent work, cited as Negri et al., 2025, has shown this is possible — but then you have a quality-control problem wearing a solutions costume. Who checks the checkers? Amazon’s answer, presented in its SynthAVE work, is a benchmark for attribute-value verification spanning 12,726 products across 229 product types, 792 attributes and 4 languages: Spanish, French, Italian and German.
The validation mechanism is the fun part. Amazon built what it calls a multi-LLM arena. Every sample is evaluated by 21 judge configurations — seven model families, each given three different prompts — and the final label is set by majority vote. When the arena disagrees with the synthetic label, a human expert adjudicates. When it agrees, the case is checked only as part of a stratified audit sample. Human effort is deliberately concentrated where it can actually change the answer; the rest of the time, the crowd of models governs itself.
The results are striking, and slightly counterintuitive. The individual judges barely agree with one another — Fleiss’ kappa of 0.76 — but that moderate disagreement is by design, because the judges were selected for diversity. Their combined majority vote agrees with human experts at Cohen’s kappa of 0.92, or 95.0% agreement, and Amazon estimates the resulting label quality at 97.9%. Diverse models with middling individual judgment aggregate, apparently, into a highly reliable committee.
It is the wisdom of crowds, rebuilt out of the things that ate the crowd’s jobs. The factory that once paid millions of human judgments now runs on twenty-one artificial ones and a vote, with people promoted upstairs to handle the appeals. One suspects the humans mostly rule in favor of the machines.

