METR · October 23, 2025

Adversarial fine-tuning reviews: separate capability elicitation from risk-threshold judgments

Why it matters

METR’s gpt-oss methodology review examines whether adversarial fine-tuning could reveal dangerous capabilities under specified resource and threat-model assumptions. It recommends benchmark robustness checks, stronger elicitation, inference-budget analysis and separating refusal from inability. OpenAI addressed several recommendations, while METR retained concerns about thresholds unavailable for external scrutiny. The review operated under an NDA and a short implementation window; it did not assess the overall merits of releasing model weights.

My takeaway: Predefine attacker resources and decision thresholds, measure capability after reducing refusals, and document which recommendations were implemented. Keep a methodology review distinct from a complete release-risk assessment.
Keep exploring

More curated notes connected through Model Evaluation and Adversarial ML.

OpenAI News · framework

Pacing model development in an era of cyber-critical capabilities

OpenAI says preliminary evidence that Astra may meet its Critical cybersecurity threshold led it to pause frontier reinforcement-learning work for two weeks and keep its largest planned run on hold. New safeguards include stronger workload and network isolation, continuous boundary testing, token-level monitoring that escalates suspicious tool activity, and broader alignment checks for deception, reward hacking, and unauthorized access.

OWASP GenAI Security Project · guide

OWASP Top 10 for Agentic Applications for 2026

OWASP's community guide organizes agentic-system risk into ten categories, including goal hijacking, tool misuse, identity and privilege abuse, memory poisoning, insecure inter-agent communication, cascading failures, and rogue-agent behavior. It provides a shared taxonomy and mitigation starting point rather than a certification checklist or evidence that a deployed system is secure.

OpenAI News · framework

A blueprint for democratic governance of frontier AI

OpenAI proposes a three-part U.S. frontier-AI governance model: harmonize emerging state safety laws into a federal baseline, strengthen CAISI as an evaluation and standards institution, and coordinate a broader resilience program. Proposed controls include severe-risk evaluations, transparency reports, independent audits, safety-incident reporting, model-weight security, whistleblower protection, and periodic technical assessments.