iGEN
Visit IGEN World Explore IGEN Expo
EXPLORE UPGRADE PLANS
BREAKING
US Stock Markets Rally as Chip Stock Gains Lift Nasdaq, S&P 500 and Dow SEBI Clarifies Unlisted Share Sale Rules: 200-Buyer Private Deal Limit GeM completes 10 years as India's trusted digital public procurement platform Moody's Assigns First-Time Baa2 Rating to RBL Bank, One Notch Above India's Sovereign Sebi Bars Zee's Subhash Chandra, Punit Goenka From Market for One Year Zepto Defers IPO by Two to Three Quarters After Tepid Investor Response Tim Cook: India Among Apple's Best Global Markets as June Quarter Records Revenue Domestic funds reach record 21% stake in Indian companies as FPI ownership drops to 17% Cybercriminals widen net as assessees rush to meet I-T return filing deadline Bloomberg Delays India's Sovereign Bond Index Inclusion as Market Reforms Need Further Testing US Stock Markets Rally as Chip Stock Gains Lift Nasdaq, S&P 500 and Dow SEBI Clarifies Unlisted Share Sale Rules: 200-Buyer Private Deal Limit GeM completes 10 years as India's trusted digital public procurement platform Moody's Assigns First-Time Baa2 Rating to RBL Bank, One Notch Above India's Sovereign Sebi Bars Zee's Subhash Chandra, Punit Goenka From Market for One Year Zepto Defers IPO by Two to Three Quarters After Tepid Investor Response Tim Cook: India Among Apple's Best Global Markets as June Quarter Records Revenue Domestic funds reach record 21% stake in Indian companies as FPI ownership drops to 17% Cybercriminals widen net as assessees rush to meet I-T return filing deadline Bloomberg Delays India's Sovereign Bond Index Inclusion as Market Reforms Need Further Testing
Home ›› Technology ›› Ai ›› Llms ›› LLM Psychological Profiles Are Mostly Measurement Artifacts, Study Finds

LLM Psychological Profiles Are Mostly Measurement Artifacts, Study Finds

A new study from arXiv demonstrates that psychological profiles assigned to large language models (LLMs) are largely measurement artifacts, driven by a directional response bias rather than stable traits. The findings call into question the validity of using human psychological instruments on LLMs, with significant implications for AI safety assessments and research proxies.

iG
iGEN Editorial
June 20, 2026
LLM Psychological Profiles Are Mostly Measurement Artifacts, Study Finds

Enterprises increasingly rely on large language models (LLMs) for tasks ranging from customer interaction to internal decision support. Some researchers have attempted to assign psychological profiles to these models — measuring personality traits, risk preferences, and other human-like characteristics — to assess usability, safety, or to use LLMs as proxies for human participants. But a rigorous analysis published on arXiv challenges the validity of such profiling.

According to the study by Meyer, Jelena, Garcia, David, and Wulff, Dirk U — titled "Apparent Psychological Profiles of Large Language Models are Largely a Measurement Artifact" — these profiles are largely an artifact of the measurement instruments themselves, not stable properties of the models.

The Measurement Artifact

The researchers administered a battery of personality and risk-preference instruments (including self-reports and behavioral tasks) to 56 instruction-tuned LLMs alongside large human reference samples. Using a formal psychometric framework, they found that differences between models are driven not by the traits an instrument aims to measure but by a directional response bias — a tendency to respond toward one end of the scale or one labeled option, regardless of item content.

A variance decomposition revealed that 81-90% of between-model variation is attributable to this bias, compared to only 9-16% in humans. This indicates that what appears to be a model's personality or risk profile is mostly noise from the instrument's design.

Key Findings

Finding Detail
Between-model variation source 81-90% from response bias vs. 9-16% in humans
Bias and capability Bias declines with model capability but is not eliminated
Apparent reliability predictor Response orthogonality (proportion of items where trait and bias point opposite)
Profile malleability The profile shifts with items used and can be manufactured through item selection

The study introduces the term "response orthogonality" to describe the proportion of items for which trait and bias point in opposite directions. Because bias rather than trait drives responding, an instrument's apparent reliability is almost entirely predicted by this orthogonality.

Implications for Enterprise AI

For enterprise technology leaders, these findings carry significant implications. LLMs are often evaluated on traits like "agreeableness" or "risk aversion" to predict behavior in customer-facing roles or safety-critical applications. If these profiles are artifacts, decisions based on them — including model selection for sensitive use cases — may be flawed.

Moreover, the study demonstrates that a model's apparent profile can be manufactured through item selection. This raises concerns about benchmark gaming and the validity of safety assessments that rely on psychological instruments borrowed from human psychology.

Call for Dedicated Assessments

The authors conclude that instruments borrowed from human psychology are rarely fully orthogonal and may inherently lack validity for LLMs. They call for dedicated assessments centered on response orthogonality rather than relying on human-derived psychometrics. As LLMs become more integrated into supply chain decision-making, trade documentation, and logistics automation — where interpretability and reliability are paramount — the need for valid measurement tools becomes critical.

For now, enterprise adopters should treat any psychological profile assigned to an LLM with skepticism. The traits that appear to emerge may reflect more about the test than the model.


Sources:

Keep Reading

Recommended Stories

Reinforcement-Aware Knowledge Distillation Boosts LLM Reasoning Efficiency Technology

Reinforcement-Aware Knowledge Distillation Boosts LLM Reasoning Efficiency

Researchers propose RL-aware distillation (RLAD) to address distribution mismatch and objective interference in knowledge distillation for LLM reasoning. The method uses Trust Region Ratio Distillation (TRRD) to selectively imitate teacher policies during reinforcement learning. RLAD outperforms offline distillation, standard GRPO, and KL-based on-policy distillation across logic and math benchmarks.

June 21, 2026
Large Language Models Can Read Compressed Text That Humans Cannot, Researchers Find Technology

Large Language Models Can Read Compressed Text That Humans Cannot, Researchers Find

A new research paper introduces BabelTele, a compact, non-human-readable text format that large language models can still interpret with high semantic fidelity. The approach compresses text to 27.9% of its original length while preserving 99.5% of meaning, potentially reducing context overhead and costs in enterprise AI deployments.

June 20, 2026
AI Economist Agent: New Framework Uses RAG, Knowledge Graphs and LLMs for Grounded Economic Analysis Technology

AI Economist Agent: New Framework Uses RAG, Knowledge Graphs and LLMs for Grounded Economic Analysis

Researchers propose an AI economist agent framework that combines retrieval-augmented generation (RAG), knowledge graphs, and LLM-based agents to ground economic analysis in data and theory. Tested on U.S. inflation persistence and bank stress-test scenarios, the approach improves economic coherence and traceability of generated reports.

June 20, 2026
New Method LUCID Detects Hallucinations in LLM-Based Knowledge Graph Reasoning Technology

New Method LUCID Detects Hallucinations in LLM-Based Knowledge Graph Reasoning

Researchers introduce LUCID, the first hallucination detection method designed for large language model-based knowledge graph reasoning. By jointly leveraging attention scores, KG semantics, and structural information via a graph neural network, LUCID achieves state-of-the-art performance across nine datasets compared to 15 baselines. The method addresses a critical gap where existing detection techniques overlook structural information in knowledge graphs.

June 20, 2026