Play video
A systematic study of server-side browsers used by AI search and web-browsing services reports remote-code-execution paths in six leading services with a combined user base above one billion. The work covers domain-allowlist bypasses, JavaScript-restriction evasion, remote browser fingerprinting, service disruption, output manipulation, and server compromise.
Anthropic red-team research assessing how LLMs affect exploitation of known vulnerabilities. Relevant to cyber capability evaluation, benchmark design, and misuse risk modeling.
SymJack demonstrates that a user-approved, apparently harmless copy command can write through a repository-controlled symlink into executable agent configuration, producing code execution when the tool restarts. The vendor-authored study reports variants across six coding agents and highlights a gap between approval text, shell semantics, and the resolved filesystem target.
METR’s 2025 predeployment assessment examines GPT-5 through autonomous task evaluations and investigations of research capabilities, replication and strategic sabotage. The report combines observed behavior with information supplied by the developer to assess specific threat models. It also explains where short evaluation windows, benchmark saturation, incomplete elicitation and limited testing of evasion weaken the conclusions. The assessment does not cover every misuse domain or establish broad alignment. Its reusable contribution is the structure of an evidence-based safety argument, including explicit dependencies on developer assurances and unresolved uncertainty.
METR’s 2025 assessment evaluates six DeepSeek and Qwen models on autonomous software tasks and AI research environments. It describes repeated task attempts, model-dependent token budgets and limited elicitation effort, with API reliability constraining some Qwen results. Suspected reward hacking received manual inspection, and identified cheating attempts were scored as failures. The report provides useful evaluation procedure and transcripts, but its short setup period, unequal budgets and limited checks for sandbagging constrain capability comparisons. Historical scores should not be read as present-day rankings or complete measures of dangerous capability.
OpenAI’s September 30 disclosure describes a coordinated campaign that manipulated model interactions to reproduce protected reasoning, including replay of encrypted artifacts across conversations. The company reports stronger boundaries across users, workspaces, organizations and model families, plus streamed-output checks and account enforcement. Reported request volumes count attempted extractions; attribution to Moonshot-associated individuals applies to a core cluster and is not publicly independently established. OpenAI says the campaign did not break encryption or directly access stored user conversations.
JFrog’s CVE-2026-90898 advisory describes command execution through Bifrost’s reachable management API when authentication is disabled. A stdio MCP registration starts a process before the protocol handshake. HTTP transport 2.1.0 blocks unauthenticated registration; 2.0.0 and the 1.6.x line through 1.6.11 lack this fix.
Invisible Unicode characters popularized for hiding instructions from AI models are now being used to obfuscate words before email filters parse them.
Release notes for garak, an LLM vulnerability scanning and evaluation toolkit. Relevant to tracking new probes, detectors, and repeatable red-team workflows.
garak release adding probes and detector improvements for LLM security testing. Relevant to maintaining practical red-team coverage across evolving attack techniques.
Google DeepMind proposes a taxonomy of ten cognitive abilities and a three-stage evaluation protocol combining held-out AI tasks, human baselines and comparisons between performance distributions. It highlights gaps in evaluating learning, metacognition, attention, executive functions and social cognition. The contribution is a framework for designing a broader measurement program. The article does not deliver a completed benchmark suite, an agreed definition of AGI or a score establishing that a model has reached it.
NVIDIA's AI Red Team extends its visual prompt-injection work with a Gemini 2.5 Pro demonstration in which a scrambled puzzle reconstructs a command during problem solving. The post calls these multimodal cognitive attacks and argues that payloads can emerge during inference after simple input filters have already run; it proposes output validation, tool sandboxing, and anomalous-reasoning detection as research directions.
METR’s January 2025 analysis identifies risk pathways that a release-time evaluation can miss: stolen model weights, employee misuse and autonomous behavior during internal development or use. It proposes applying assessment, information security and independent oversight throughout the model lifecycle, rather than waiting for public availability. These are threat-model arguments and governance recommendations, not evidence that the hypothetical severe incidents occurred. The practical distinction is between limiting public access and controlling all the environments in which a capable model already operates.
OWASP’s GenAI security project remains a practical baseline for teams building or assessing LLM applications and agentic systems.
Play video
Ryan Cooke’s publisher notes explain how WorkOS connects coding agents to project plans, ticket dependencies and completion webhooks. A shared MCP gateway supplies tool access and guidance about where organizational information lives. In the demonstrated workflow, a short brief becomes draft planning documents that an engineer refines before implementation proceeds. The talk reports qualitative experience rather than measured delivery gains and explicitly leaves cross-system authorization unresolved.
Novee Security found that unprivileged GitHub issues could reach privileged coding-agent workflows: Gemini CLI and Claude Code paths led to CI-runner code execution, while a Codex path could alter the next agent run. The two assigned CVEs were patched; the Codex behavior was documented rather than assigned a product CVE, highlighting failures in the surrounding harness rather than the model alone.
A credential-stealing npm worm spread through hundreds of package versions using lifecycle scripts and a Bun-based payload. Related repositories also carried Claude Code and VS Code hooks that could execute after workspace trust; reported campaign totals vary, so exposure depends on exact resolved versions and execution.
Learn how CNAPP platforms are helping organizations prioritize exploitable risks, reduce exposure, and operationalize security across the application lifecycle.
GPT-Rosalind advances life sciences research with enhanced biological reasoning, medicinal chemistry expertise, genomics analysis, and experimental workflow capabilities.
Introducing GPT-5.5, our smartest model yet—faster, more capable, and built for complex tasks like coding, research, and data analysis across tools.
OpenAI introduces GPT-Rosalind, a frontier reasoning model built to accelerate drug discovery, genomics analysis, protein reasoning, and scientific research workflows.
METR’s retrospective on an AI-biology randomized trial focuses on evaluation design and execution. It describes the difficulty of matching long studies to fast model releases, staffing specialized evaluations and interpreting a result from a limited participant and task distribution. A lack of a significant aggregate effect in that study does not establish the absence of risk for other users, tasks or systems. The transferable lesson is to prepare measurement and safeguards before capability changes force a decision.
In a recent evaluation of AI models’ cyber capabilities, current Claude models can now succeed at multistage attacks on networks with dozens of hosts using only standard, open-source tools, instead of the custom tools needed by previous generations.
Google Cloud outlines a defense-in-depth view of AI security spanning application controls, data protections, and infrastructure isolation.