Topic
systems
Benchmarking Agentic Review Systems: AI Peer Review Achieves 83% Pairwise Accuracy but Falls Short on Error Detection
A study by Nguyen et al. benchmarks two open-source and one proprietary AI review system on peer review tasks. The best configuration (OpenAIReview + GPT-5.5) achieves 83.0% pairwise accuracy in tracking paper quality but only 71.6% recall in detecting injected errors. User feedback shows a positive-to-negative vote ratio of 1.44:1, with common complaints about false positives. The research highlights both the potential and limitations of current AI agents in evaluation tasks.
Beyond Models: Reflections on Engineering AI-enabled Systems in a Project-Based Course
A project-based master's course on AI systems engineering at the University of Bremen highlights key challenges such as architectural decisions, heterogeneous ML integration, evolving requirements, and data management. A mixed-methods study of student submissions and questionnaires shows that uneven expertise in ML and software engineering exacerbates these difficulties. The course successfully fostered system-level reasoning and data-centric ML awareness.
A Framework for Governing Optimization in AI Systems: Architectural Wisdom
The paper 'Architectural Wisdom' argues that modern AI failures stem from optimizing underspecified objectives, not lack of intelligence. It proposes a corrigible objective-governance layer above the optimization substrate, made of four components and a six-coordinate wisdom tuple. The framework is motivated by eight cases of contemporary AI failures and aims to prevent harmful outcomes.
Technology AI's Impact on Astronomy: A Double-Edged Sword
AI systems are increasingly integrated into astronomy research, raising concerns about the future of human reasoning and traditional skills in the field. While AI aids in data analysis and problem-solving, it also threatens to diminish essential scientific skills.