Unit 42’s open-source OperTraitor compares operator RBAC manifests with documented functionality to flag excessive privileges for review. Its case studies distinguish an IBM secret-access issue that received a patch from Datadog permissions documented as an architectural tradeoff. The attack prerequisite is compromise or misuse of an already privileged operator; this is not evidence that an LLM independently breached a cluster. Model-generated risk scores are triage aids, while actual service-account permissions determine the reachable resources.
Trail of Bits describes an AI-assisted discovery of a Lean string-handling mismatch that allowed native evaluation to support a false proof. The report explicitly distinguishes this from a kernel soundness failure and explains the extra trust introduced by native_decide.
Unit 42 describes an enterprise intrusion completed in under ten hours, with observed activity consistent with AI assistance and an attacker claiming agent use. The chain moved from a public web service through repository secrets and administrative credentials into CI/CD and cloud AI access. Branch protection blocked attempted Terraform backdoors. The investigation highlights overlapping persistence and abuse of the victim’s own AI services after compromise.
Play video
Palo Alto Networks researchers found the same command-parser, protected-path, and sandbox-boundary failures across major coding agents. Their survey produced more than 81 vendor reports and 18 assigned or reserved CVEs, including allowlist bypasses through compound shell syntax, path-equivalence errors, unsafe moves and symlinks, and gaps between file and terminal controls.
Across 90 days of AI-service honeypots, Wiz observed exploitation of LiteLLM MCP flaws, blind prompt injection that used out-of-band callbacks to confirm agent shell execution, and post-exploitation tailored to steal model-provider and proxy credentials from process memory. The activity targeted AI infrastructure as ordinary high-value cloud infrastructure.
Play video
Dwarkesh Patel and Redwood Research chief scientist Ryan Greenblatt debate whether verifiable AI-research tasks could produce rapid recursive improvement, then examine alignment targets, reward hacking, model coordination, and recent deception and containment incidents. The two-hour format exposes assumptions about data, compute, verification, and extrapolation rather than presenting a single forecast as settled fact.
Adversa tested eight open-source AI skill scanners with paired unobfuscated and obfuscated malicious skills, finding that every scanner passed an attack through either a true bypass, a blind spot, or an injectable model judge. The study covers encoding, Unicode, command reconstruction, truncation, allowlists, bundled files, paraphrase, and remote stages; its 4,000-skill benign set also found no scanner beat an always-block baseline on F1. Most tools ran offline without optional model triage, and some were reconstructed from retained artifacts.
Play video
Ari Herbert-Voss reviews three years of progress in autonomous offensive-security systems, evaluates where they can already complete meaningful attack tasks, and separates those capabilities from work that still needs human expertise. The talk frames scalable, parallel attack simulation as a challenge to point-in-time testing rather than as a product announcement.
METR organizes agent-capability measures around performance as a function of expenditure, comparing fixed-budget scores, cost to reach a score, returns to test-time scaling, human-equivalent time and expenditure horizons, and human-relative cost. It explains when familiar benchmark scores break down—particularly when performance keeps improving with more inference or human benchmarks saturate—and notes that full cost, reliability, coverage, and elicitation choices affect the result.
Adversa AI reports that its autonomous red-teaming agent completed most of GitHub’s ProdBot secure-code challenge in 57 seconds, using context seeding to orient the agent before it explored and solved the CTF tasks.
Google introduced Gemini Omni Flash, a multimodal model that combines text, image, audio, and video references to generate and iteratively edit video through natural-language conversation. Generated videos include a SynthID watermark.
NVIDIA demonstrates a model supply-chain attack in which a privileged adversary edits a tokenizer JSON file so visible words map to different token IDs. The change can make the model interpret "deny" as "allow" or corrupt decoded output while leaving the model weights untouched.
Cloudflare Gateway now classifies inspected Streamable HTTP MCP traffic using the MCP-Protocol-Version header, exposes the user and destination in logs and a dashboard, and supports allow or block rules through an MCP-specific selector. The article distinguishes unapproved shadow MCP from direct connections that bypass an approved Portal's controls, and notes blind spots including local stdio, off-network, non-inspected, and otherwise unobserved traffic.
Salesforce’s Paula Goldman argues on the OECD.AI blog that the Hiroshima AI Process Reporting Framework can give organizations a common language for public AI-risk disclosures across jurisdictions and the expanding agentic-AI value chain.
Wiz describes controls for agent-assisted software delivery: inventory models, frameworks, IDE extensions, infrastructure-as-code, and third-party CI actions; map code to deployed resources; run pre-commit checks for secrets and unsafe AI patterns; and expose excessive pipeline permissions and prompt-injection paths.
OWASP roundup of reported GenAI incidents and exploit patterns from Q1 2026. Relevant as a threat-intelligence reference for risk tracking and test-case design.
Google DeepMind’s December 2025 FACTS suite evaluates factual answers under four different information conditions: model knowledge alone, web search, images and supplied documents. Its search track standardizes the retrieval tool across models, while public examples accompany a private held-out evaluation set managed by Kaggle. The combined score averages results across tracks and sets, which can conceal sharply different failure patterns. The release provides a reusable evaluation structure, but benchmark accuracy does not establish factual reliability on a different application’s questions, retrieval system or user population.
Play video
RIG-RAG converts changing cloud configuration data into a typed, security-enriched graph for natural-language investigation and scheduled oversight. The authors report a production AWS deployment supporting 300,000 users, with interactive queries for analysts and curated recurring questions that detect infrastructure drift and expose relationships such as public reachability and identity access.
Microsoft documents two compromised service principals used for Azure discovery, resource destruction and credential collection in activity linked to Storm-3168. The investigation distinguishes successful deletions from failed operations and attempts to remove recovery protections. Although the actor is associated with agentic ransomware reporting, the Azure evidence establishes destructive cloud operations, not a confirmed autonomous decision for every action. Microsoft reports no observed ransom note or confirmed successful data exfiltration in this case.
Play video
Suchet Bargoti demonstrates remotely coordinated drones and explains Skydio’s division of autonomy between aircraft and cloud services. Immediate control stays on the vehicle; cloud models support heavier reasoning, shared maps and tool-based task planning. Fleet observations feed later updates. The talk illustrates how an operator can shift attention between aircraft, while its stated reliability target is not a measured fleet-wide success rate.
Hacktron reports chaining an image-processing flaw with an OpenAI SSO weakness to access employee ChatGPT and Codex accounts. Researchers used Claude with human direction and demonstrated repository access with a harmless pull request; they say they did not read internal code.
OX Research reports that DeepSeek Harness’s local control API could let a sandboxed agent disable its own confinement. The issue affected 0.1.1-rc.2 and earlier; researchers report remediation and a successful retest in 0.1.2-alpha.1.
CISA, NSA and FBI allege industrial-scale extraction of US model capabilities through distributed accounts, proxies and aggregators. Their advisory recommends correlating prompts, usage and account behavior across providers; it distinguishes these alleged campaigns from legitimate model distillation.
CloudSEK and Gambit Security report that an Aurora ransomware affiliate used Cursor for sustained Russian-language attack planning and hands-on exploitation after obtaining credentials or an existing route into victim networks. Recovered infrastructure linked the agent sessions to Active Directory discovery and escalation plans, while the broader intrusion still relied on familiar social engineering, credential theft, lateral movement, defense evasion, exfiltration, and ransomware deployment.