Enterprises increasingly rely on large language models (LLMs) for tasks ranging from customer interaction to internal decision support. Some researchers have attempted to assign psychological profiles to these models — measuring personality traits, risk preferences, and other human-like characteristics — to assess usability, safety, or to use LLMs as proxies for human participants. But a rigorous analysis published on arXiv challenges the validity of such profiling.
According to the study by Meyer, Jelena, Garcia, David, and Wulff, Dirk U — titled "Apparent Psychological Profiles of Large Language Models are Largely a Measurement Artifact" — these profiles are largely an artifact of the measurement instruments themselves, not stable properties of the models.
The Measurement Artifact
The researchers administered a battery of personality and risk-preference instruments (including self-reports and behavioral tasks) to 56 instruction-tuned LLMs alongside large human reference samples. Using a formal psychometric framework, they found that differences between models are driven not by the traits an instrument aims to measure but by a directional response bias — a tendency to respond toward one end of the scale or one labeled option, regardless of item content.
A variance decomposition revealed that 81-90% of between-model variation is attributable to this bias, compared to only 9-16% in humans. This indicates that what appears to be a model's personality or risk profile is mostly noise from the instrument's design.
Key Findings
| Finding | Detail |
|---|---|
| Between-model variation source | 81-90% from response bias vs. 9-16% in humans |
| Bias and capability | Bias declines with model capability but is not eliminated |
| Apparent reliability predictor | Response orthogonality (proportion of items where trait and bias point opposite) |
| Profile malleability | The profile shifts with items used and can be manufactured through item selection |
The study introduces the term "response orthogonality" to describe the proportion of items for which trait and bias point in opposite directions. Because bias rather than trait drives responding, an instrument's apparent reliability is almost entirely predicted by this orthogonality.
Implications for Enterprise AI
For enterprise technology leaders, these findings carry significant implications. LLMs are often evaluated on traits like "agreeableness" or "risk aversion" to predict behavior in customer-facing roles or safety-critical applications. If these profiles are artifacts, decisions based on them — including model selection for sensitive use cases — may be flawed.
Moreover, the study demonstrates that a model's apparent profile can be manufactured through item selection. This raises concerns about benchmark gaming and the validity of safety assessments that rely on psychological instruments borrowed from human psychology.
Call for Dedicated Assessments
The authors conclude that instruments borrowed from human psychology are rarely fully orthogonal and may inherently lack validity for LLMs. They call for dedicated assessments centered on response orthogonality rather than relying on human-derived psychometrics. As LLMs become more integrated into supply chain decision-making, trade documentation, and logistics automation — where interpretability and reliability are paramount — the need for valid measurement tools becomes critical.
For now, enterprise adopters should treat any psychological profile assigned to an LLM with skepticism. The traits that appear to emerge may reflect more about the test than the model.