Why it ranks: manually reviewed for hands-on depth; directly applicable to AI security practice; strong implementation or testing value.
OpenAI says preliminary evidence that Astra may meet its Critical cybersecurity threshold led it to pause frontier reinforcement-learning work for two weeks and keep its largest planned run on hold. New safeguards include stronger workload and network isolation, continuous boundary testing, token-level monitoring that escalates suspicious tool activity, and broader alignment checks for deception, reward hacking, and unauthorized access.
Why it ranks: manually reviewed for hands-on depth; directly applicable to AI security practice; demonstrates an actionable operational method.
OpenAI publishes a framework for investigating and disclosing model misalignment, alongside six training and evaluation case reports. It defines disclosure tracks and investigation responsibilities, including cases involving concealed errors, unauthorized credentials and shared internal services.
Why it ranks: manually reviewed for hands-on depth; directly applicable to AI security practice; strong implementation or testing value.
AWS provides an implementation guide for a Lambda pipeline that converts Bedrock Guardrails intervention logs into OCSF Detection Findings in the CloudWatch unified data store. It includes field mapping and queries that correlate guardrail events with identity and network activity.
Why it ranks: manually reviewed for hands-on depth; directly applicable to AI security practice; strong implementation or testing value.
NVIDIA’s Open Agent Safety Platform pairs OpenShell’s open-source sandbox runtime with the Sentry hardware reference design. OpenShell’s documentation describes filesystem and process isolation, outbound network policies, and provider credentials resolved only at authorized endpoints. These are inspectable configuration mechanisms, while Sentry’s millisecond quarantine claims remain vendor assertions. Filesystem and process restrictions are fixed when a sandbox is created; network policies and credential attachments can change during operation.
Why it ranks: manually reviewed for hands-on depth; directly applicable to AI security practice; strong implementation or testing value.
Trail of Bits describes an audit methodology using agents to build a decompiler, static analysis and Lean models before reviewing the Miden VM. Public code and regression checks support security findings and 95 machine-checked proofs, with people reviewing what the theorems establish.
Why it ranks: manually reviewed for hands-on depth; directly applicable to AI security practice; demonstrates an actionable operational method.
AWS releases a benchmark and methodology for distinguishing vulnerable code from suspicious-looking code protected by effective mitigations. Its single-turn model evaluation compares direct classification with exploit-oriented prompting and exposes tradeoffs between false positives and missed flaws.