A country average leaves people out
When researchers ask whose values an AI model reflects, they often compare countries. A new study argues that the border is too blunt a line.
The authors compared answers from 10 commercial language models with 50,116 people in the 11th European Social Survey. The survey covered 29 European countries and Israel, along with details such as education, income, occupation, religion and age.
The models did not sit equally close to every group. On average, their survey answers were closer to those of people with more education, greater financial security and more white-collar work. Religious identity also produced some of the widest gaps.
This does not mean a model holds a coherent set of human values. It means its selected answers landed nearer to some groups than others on this particular survey.
How the comparison worked
The researchers selected 47 questions about values and opinions. They asked each model every question 20 times in English, using the original response scales and no system prompt, then took the majority answer.
The set covered models from OpenAI, Anthropic, DeepSeek and Mistral. Eight questions were left out of the cross-model comparison because at least one model would not give a usable answer.
For each person and question, the team calculated the distance between the person’s answer and the model’s answer on the survey scale. Those distances became an alignment score from zero to one, then were averaged across questions.
It is a neat way to compare a large population with several models. It is also narrow. A tick on a survey scale is easier to measure than the way a model responds in a real conversation.
The social gaps were not small enough to ignore
Financial security showed a consistent pattern. The widest difference between people who felt most and least comfortable on their household income was 0.0385 on the study’s zero-to-one scale.
Education moved in the same direction: people with higher qualifications were generally closer to the models’ answers. Unemployed respondents tended to score lower, while white-collar occupations scored higher on average.
Religious denomination produced the largest spread among the demographic variables. The paper reports lower average alignment for Muslim and Eastern Orthodox respondents than for Protestants, with a 0.051-point gap between the outer groups.
These are associations inside the survey data. They do not establish why the gaps exist. Training material, post-training choices, the English prompt and the survey itself may all contribute.
Country still mattered
Looking below the national level did not make country disappear. On the broader set of value-laden questions, country alone explained between 7.3% and 27.8% of individual score variation, depending on the model. That was at least as much as all 15 demographic variables together.
Reweighting countries so their demographic mix looked more alike did not remove most of the difference between them. The best predictions came from combining country and social variables, where the study’s boosted models explained between 27.8% and 42.8% of the variation.
The balance changed on a narrower set of 21 questions about basic human values. For several model families, demographics then mattered more than country. What researchers choose to call a value changes the result.
That is the useful part. A pluralistic AI system cannot be evaluated against one national average and assumed to represent everyone inside it.
What is confirmed, found and still open
Confirmed: the three-author paper was submitted on 7 August 2026 and accepted at AIES 2026. It uses the ESS Round 11 dataset, 10 models, 15 demographic variables and repeated survey prompts. The authors released the paper under CC BY 4.0.
The research finding: agreement varied across countries and social groups. Country and demographics were complementary, and the relative weight of each depended on the questions being measured.
Still open: whether the pattern survives other prompts, other languages, newer model versions and more natural interactions. The study used majority answers, did not vary the prompt and treated multiple-choice responses as a proxy for values.
Some survey fields were incomplete too: income decile was missing for 20.8% of respondents and occupation for 12.3%. This is one careful measurement of a difficult problem, not a final map of who AI represents.
Sources
- Wightman, Bied and De Bie - People Are Not Just Their CountriesPrimary paper record submitted 7 August 2026 and accepted at AIES 2026. Source for authorship, scope, licence and the headline findings.
- Wightman, Bied and De Bie - Full paperFull primary manuscript. Source for the model set, prompting method, group results, variance analysis, missing data and limitations.
- European Social Survey - ESS11 integrated file, edition 4.1Primary dataset record for the survey wave used in the analysis. The paper reports 50,116 respondents and applies ESS post-stratification weights where possible.



