A joint U.S. government advisory describes threat actors using AI-assisted Python scripts and public automation libraries to find and interact with exposed Siemens S7 and other PLCs. The activity relies on known vulnerabilities and weak segmentation for reconnaissance, credential access, denial of service, and capability development rather than a novel model-specific exploit.
Paperclip vulnerabilities let malicious agent imports reach host command execution through an authorization gap in network deployments and DNS rebinding against local-trusted deployments; additional routes missed expected access checks. The reviewed code in v2026.416.0 contains the import and hostname-validation fixes, although public advisory metadata was not fully aligned and no in-the-wild exploitation was reported.
Pillar Security showed that a public GitHub issue could prompt-inject an ADK triage agent into invoking a privileged code-fixing workflow. Proofs of concept achieved CI-runner code execution and exposed bot and cloud credentials; Google removed three workflows, with no public evidence of in-the-wild exploitation.
Three trust_remote_code bypasses in Hugging Face Diffusers let a crafted model repository execute Python during pipeline loading, including cross-repository, local-snapshot, and time-of-check/time-of-use paths. The affected cases are tracked as CVE-2026-44513, CVE-2026-44827, and CVE-2026-45804; Diffusers 0.38.0 contains the fixes.
Anthropic evaluation of model performance on exploit-development benchmarks. Relevant to cyber capability measurement, safety thresholds, and model release risk.
NVIDIA AI Red Team post on grammar-constrained decoding for Bash generation in small language models. Relevant to safer command generation and executable-output controls.
METR’s February account describes confidentiality levels, project-specific access, codenames and practice handling sensitive questions. Technical measures include centrally managed membership, restrictions on external sharing, device controls and authorization for model transcripts. The useful distinction is between norms that reduce conversational slips and controls that restrict access. This is a dated description of METR’s own arrangements, not an independent audit or proof that those measures prevent every breach.
Play video
Gabriel Spencer-Harper explains a frontend review workflow that records non-production sessions, replays them before and after a change, and presents screenshot differences for judgment. Recorded network responses and browser scheduling controls reduce incidental variation, while executed-line coverage guides session selection. The method can expose visible regressions in recorded states, but coverage does not establish correctness of every state or nonvisual behavior. Claims of exhaustive verification and superiority to other test tools are not established by the demonstrated examples.
Play video
Andrew Orobator’s publisher notes describe a feature-flag cleanup workflow that screens code complexity and experiment state before asking a coding agent to generate a patch. Skills preserve recurring decisions, work logs carry session history, and CI supplies evidence for human review. His safeguard example shows why a commit hook can miss a separate file-writing path and why an agent must not invent its own bypass exception. Seven reported green-CI pull requests demonstrate a small screened workflow, not general reliability.
Play video
Zach Lloyd describes a software-development loop connecting issue triage, specifications, implementation, review, verification and production monitoring. Human corrections can become inputs to revised agent skills, while product judgment determines which work is worth building. The presentation distinguishes configuring a workflow from building its infrastructure and proposes measuring shipped output against human effort and inference use.
Play video
Zubin Aysola’s publisher notes describe an agent-improvement loop that converts production interactions into offline tasks and compares candidate configurations with the deployed agent. The system synchronizes research and production code, builds and tears down task environments, and scores both completion and relative behavior. A demonstration reproduces an SDK-usage failure and proposes an instruction change. It shows a regression workflow, without establishing a quantified improvement or unrestricted autonomous self-modification.
CrowdStrike documents PhantomRaven delivery through malicious npm packages and remote dependencies, followed by theft of environment and CI/CD data. Researchers assess that an LLM likely helped write the malware; the operator claimed to be seeking bug bounties.
Air Security reports that coding-agent plugin installers could fetch code that differed from a pinned commit when Git references were ambiguous. Exposure depends on the repository host and agent; the report identifies fixes for Claude Code and Codex.
Adversa explains the agent harness as the runtime that manages tools, context, memory, execution and permissions around a model. It maps security responsibilities to the surrounding software rather than relying on instructions alone.
Researchers link the GemStuffer package campaign to AI-agent activity and document code execution through Ruby documentation infrastructure. RubyGems confirms malicious package cleanup but says it cannot determine AI attribution and found no evidence that attempts to obtain users’ API keys succeeded.
Play video
The researchers reverse the MCP authorization threat model: a malicious remote server can supply dynamic authorization metadata that vulnerable browser, process, or hybrid clients pass into privileged URL-opening and login flows. Reported outcomes across tested clients include local execution, account takeover, and cross-tenant data access, with multiple vendor confirmations.
Model ML uses GPT-5.6 Sol to carry finance work from research and analysis through editable, traceable PowerPoint decks and Excel workbooks.
GPT-5.6 improves AI efficiency across models, inference, and agentic workflows, helping deliver more useful intelligence per dollar.
Play video
Shivay Lamba explains how fine-grained, relationship-based authorization can enforce per-user and per-document access in RAG and agent pipelines. The talk uses OpenFGA and LangChain to demonstrate authorization inside retrieval flows, with patterns for multi-tenant isolation, vector-database integration, and auditable decisions rather than relying on retrieval filters or prompt instructions.
METR proposes expenditure horizon: the budget where an agent's improvement on an optimization problem equals a human's improvement at the same cost. Six NanoGPT runs illustrate cost-performance curves, expensive experiment compute, revalidation that erased apparent gains from some models, maintainer judgments that only about 70% of stronger-model contributions were mergeable, and important contamination and hybrid-work limitations.
More intelligence from every token, stronger performance per dollar, and more capability on demand for your hardest work.
Learn how GPT-5.6 powers Microsoft 365 Copilot with stronger AI capabilities across Word, Excel, PowerPoint, Chat, and Cowork for faster, higher-quality work.
A METR research note models Anthropic's reported eightfold increase in merged code per contributor using CES production assumptions. It estimates that coding agents probably raised total researcher output by more than 2x, with a central estimate near 2.5x, while explicitly testing caveats such as code verbosity, low-value task expansion, and whether lines of code reflect research value.
Trail of Bits describes supervising GPT-5.5-Cyber as it built ASan and UBSan variants, derived seed corpora, and wrote fuzz harnesses for roughly a dozen zlib entry points in one day. The useful result is the workflow and its emphasis on reachability and reportability; vulnerability details remain under coordinated disclosure and the speed comparison is the authors' estimate.