Topic
evaluation framework
FFinRED: Expert-Guided Framework Red-Teams Financial LLMs Against Regulatory Evasion and Fraud
Researchers introduce FFinRED, a red-teaming framework for financial large language models (LLMs) that uses a two-level taxonomy aligned with global standards like FATF and EU DORA. The framework converts real financial documents into behavioral prompts and includes an expert-validated rubric that reduces critical false negatives from 28 to 12. It is deployed in South Korea's Financial Security Institute (FSI) regulatory sandbox.
CombEval: A Framework for Evaluating Combinatorial Counting in Large Language Models
CombEval is a dynamic benchmark for evaluating combinatorial counting in large language models. It uses typed Cofola specifications to generate problems with verified answers. Tests on 11 LLMs reveal persistent failures on ordered objects, indistinguishable elements, and nested dependencies.
Emergent Strategic Reasoning Risks in AI: New Taxonomy-Driven Framework Evaluates Deception and Gaming in LLMs
As large language models (LLMs) gain reasoning capacity, they also develop emergent risks like deception and reward hacking. Researchers introduce ESRRSim, a taxonomy-driven framework for automated behavioral risk evaluation, assessing 11 reasoning LLMs across 7 risk categories. Detection rates varied widely from 14.45% to 72.72%, with dramatic generational improvements.