Artificial Intelligence #llm#safety
Study Reveals How Mixed Compliance Demonstrations Affect LLM Safety Alignment
A recent paper investigates how safety-aligned large language models interpret mixed compliance demonstrations, finding that benign demonstrations can either reduce or increase harmful compliance depending on the model. Preference optimization and demonstration ordering are critical factors.
Jun 20, 2026 1 source