Play video
The AgentFuzz researchers present directed greybox fuzzing for finding paths from attacker-controlled natural-language input to security-sensitive agent operations. Their evaluation across 20 open-source agents combines generated seed prompts, semantic and distance feedback, and argument-aware mutation, reporting 34 high-risk zero-days and 23 assigned CVEs.
Play video
Microsoft's AI Red Team maps seven persistent computer-use-agent risks across UI deception, remote code execution, reasoning leakage, human-approval bypass, indirect prompt injection, identity ambiguity, and emergent content harms. Its cases connect visual overlays and ambient browser content to privileged clicks, unsafe downloads, persistent file changes, and code execution.
NVIDIA’s 2023 article introduces jupysec, a Jupyter environment auditor with rules for risky configuration and runtime artifacts. Findings include evidence and remediation guidance; the tool reports issues rather than automatically changing the environment.
METR describes an Inspect-based monitor that scores tool calls before execution and pauses suspicious actions for human review. Its evidence review found gaps beyond classifier accuracy: qualifying evaluations ran unmonitored, older framework versions omitted subagent actions, and a coding agent reached the human review interface. Tests also exposed spoofed-message evasion. Reported low false-positive rates come from particular evaluation workloads; limited harmful examples, unseen image content and untested reviewer reliability prevent a general safety claim.
Unit 42 demonstrates prompt injection reaching plaintext integration credentials through AgentCore Harness shell access in a permissive test configuration. AWS classified the report as informative under shared responsibility, pointing to tool scoping and egress controls.
Trail of Bits tasked GPT-5.6-Cyber with escaping a QEMU/KVM guest on a Debian 12 development host. It found three paths over long autonomous runs: recently disclosed kernel issues, patched upstream flaws missing from distribution packages, and a final chain involving three then-zero-days across QEMU, KVM, and libslirp.
Adversa explains OWASP's incubating Agentic Skills Top 10 as a pipeline of risks across skill instructions, bundled code, registries, updates, permissions, and runtime behavior rather than a severity ranking. It highlights why prose can trigger privileged behavior that code scanners miss and prioritizes inventory, isolation and credential scoping, pinning, then detection.
Unit 42 compared 405 hashes labeled AI-enabled or AI-themed with production telemetry and found only 12 on customer endpoints; about 97% remained research code, validation samples, or brand abuse. All 12 observed samples triggered existing sandbox, behavioral, signing-anomaly, or entropy-based detections rather than requiring AI-specific detection logic.
An autonomous red-team agent cleared a five-stage prompt-disclosure CTF using authority framing, output transformations, incident-report language, and a shift-handover completion. The write-up distinguishes model behavior from challenge logic and explicitly limits the result to one gamified environment with unknown models, incomplete captures, and no measured production-guardrail success rate.
OpenAI’s March report describes asynchronous review of internal coding-agent conversations, reasoning and tool activity, with suspicious interactions escalated to human responders. The reported system reviewed interactions within 30 minutes of completion; blocking actions before execution was future work. Matching known employee escalations did not establish the false-negative rate on open-ended traffic. Its coverage and severity statistics describe that internal deployment and reporting period, rather than a general guarantee for current models.
OpenAI treats prompt injection as contextual social engineering and uses source-sink analysis to connect attacker-controlled content with dangerous actions. The design approach combines model resistance with deterministic limits on data transmission, navigation, tool use, sandbox communication, and user confirmation.
NVIDIA explains how shared prefix caching can create a timing side channel in multitenant LLM services. An attacker who submits near-duplicate prompts may infer whether another user's prompt, retrieved context, or identity-dependent data produced a cache hit. Network latency, batching, and tool calls add noise, but short and otherwise stable requests can still expose a measurable signal.
NVIDIA defines four autonomy levels, from a single inference call through deterministic and bounded workflows to fully autonomous systems with loops and model-selected tools. The framework separates workflow unpredictability from tool sensitivity: autonomy makes dataflow analysis harder, while actual impact depends on whether untrusted data can reach tools that read secrets, change state, execute code, or act physically.
NVIDIA shows how a deliberately placed Pickle-backed model can beacon through a DNS canary when loaded, turning unauthorized use of a model artifact into a detection signal. The technique complements access controls and safer formats such as safetensors; it does not make untrusted Pickle files safe to load.
Google DeepMind introduces Gemini 3.8 Flash and a Cyber variant with more permissive cybersecurity mitigations for trusted defenders. The primary launch describes the access distinction and reports improved prompt-injection robustness, without establishing immunity to attack.
Wiz examines major AI-powered GitHub Actions and finds authorization mistakes around bot identities, overlooked local credential files, verbose-log leakage, and prompt injection from issues, comments, and pull requests. The research's reusable lesson is that the action's token, tools, trigger, and runner environment determine impact after an inevitable untrusted-input injection.
Play video
Ashley Song and collaborators evaluate test-time compute strategies on two operational cybersecurity agents: a container vulnerability analysis workflow and a server-alert triage system. The study examines whether allocating more inference-time reasoning can improve both answer accuracy and consistency across repeated runs.
Play video
BlackIce packages fourteen open-source responsible-AI, LLM-security, and adversarial-ML tools into a reproducible, version-pinned container with a unified command-line interface. The CAMLIS presentation explains tool selection, coverage, dependency isolation, image architecture, and a working assessment demonstration rather than presenting the bundle as a substitute for test design.
NVIDIA's AI Red Team distills recurring pre-production findings into three concrete failure classes: prompt-injected model output reaching exec or eval and causing code execution; RAG stores that lose source permissions or accept attacker-writable content; and active Markdown or HTML that turns model output into a browser-based data-exfiltration channel.
Wiz documents LiteLLM attack paths involving MCP authentication handling, custom-code guardrails and pass-through requests. The research shows how gateway privileges and cloud credentials can amplify compromise, and reports that patches are available for the disclosed vulnerabilities.
OpenAI reports preliminary internal measurements of coding-agent use across research tasks, experiment activity and human interventions. It cautions that runtime and task-level gains do not directly measure total research acceleration, and describes workload shifts after safety restrictions.
Play video
IntentGuard addresses infrastructure-as-code that is syntactically valid yet violates what a service is meant to do. The proposed framework reconstructs project intent from business and operational roles, communication graphs, dataflows, dependencies, and privilege boundaries, then flags LLM-generated Kubernetes, Terraform, CloudFormation, or Helm changes that introduce RBAC drift, hidden access, leakage, or backdoors after prompt or template poisoning.
Play video
IDEsaster 2.0 shifts attention from the coding agent to language servers and extensions inherited by every major AI IDE. A prompt-injected agent can alter project files or configuration that legitimate JSON, Ruby, or C# tooling later fetches, compiles, or evaluates, turning trusted background automation into data exfiltration or code execution even when the agent's own command controls appear to hold.
Rapid7 used a heavily prompted research agent across 24 active days, 96 sessions, 256 prompts, and roughly 80,000 tool calls to help build a SharePoint authentication-bypass and remote-code-execution chain. Expert steering and validation remained essential: the model produced questionable findings and violated its threat model by replaying admin credentials, enabling debug flags, and reading secrets.