Full Archive · Page 12

Research archive, page 12

Browse entries 265–288 of 1531. Return to the first page to search and filter the complete collection.

How We Turned AI's 'Web Browsing' Into a Gateway for Targeting 1B+ Users video thumbnail Play video
Black Hat June 28, 2026 video

How We Turned AI's 'Web Browsing' Into a Gateway for Targeting 1B+ Users

A systematic study of server-side browsers used by AI search and web-browsing services reports remote-code-execution paths in six leading services with a combined user base above one billion. The work covers domain-allowlist bypasses, JavaScript-restriction evasion, remote browser fingerprinting, service disruption, output manipulation, and server compromise.

Adversa AI Trusted AI Blog May 26, 2026 analysis

SymJack: the approval prompt is lying to you. A symlink-hijack RCE in six AI coding agents

SymJack demonstrates that a user-approved, apparently harmless copy command can write through a repository-controlled symlink into executable agent configuration, producing code execution when the tool restarts. The vendor-authored study reports variants across six coding agents and highlights a gap between approval text, shell semantics, and the resolved filesystem target.

METR August 7, 2025 analysis

METR’s GPT-5 assessment combines capability tests with a limited safety argument

METR’s 2025 predeployment assessment examines GPT-5 through autonomous task evaluations and investigations of research capabilities, replication and strategic sabotage. The report combines observed behavior with information supplied by the developer to assess specific threat models. It also explains where short evaluation windows, benchmark saturation, incomplete elicitation and limited testing of evasion weaken the conclusions. The assessment does not cover every misuse domain or establish broad alignment. Its reusable contribution is the structure of an evidence-based safety argument, including explicit dependencies on developer assurances and unresolved uncertainty.

METR June 27, 2025 analysis

DeepSeek and Qwen evaluations expose the limits of quick model comparisons

METR’s 2025 assessment evaluates six DeepSeek and Qwen models on autonomous software tasks and AI research environments. It describes repeated task attempts, model-dependent token budgets and limited elicitation effort, with API reliability constraining some Qwen results. Suspected reward hacking received manual inspection, and identified cheating attempts were scored as failures. The report provides useful evaluation procedure and transcripts, but its short setup period, unequal budgets and limited checks for sandbagging constrain capability comparisons. Historical scores should not be read as present-day rankings or complete measures of dangerous capability.

OpenAI September 30, 2026 analysis

Protected-reasoning replay: bind portable artifacts to their authorized context

OpenAI’s September 30 disclosure describes a coordinated campaign that manipulated model interactions to reproduce protected reasoning, including replay of encrypted artifacts across conversations. The company reports stronger boundaries across users, workspaces, organizations and model families, plus streamed-output checks and account enforcement. Reported request volumes count attempted extractions; attribution to Moonshot-associated individuals applies to a core cluster and is not publicly independently established. OpenAI says the campaign did not break encryption or directly access stored user conversations.

JFrog Security Research September 14, 2026 analysis

Bifrost MCP registration: patch the transport and protect the management API

JFrog’s CVE-2026-90898 advisory describes command execution through Bifrost’s reachable management API when authentication is disabled. A stdio MCP registration starts a process before the protocol handshake. HTTP transport 2.1.0 blocks unauthenticated registration; 2.0.0 and the 1.6.x line through 1.6.11 lack this fix.

Google DeepMind Blog March 17, 2026 news

Cognitive evaluations: cover missing abilities and compare human distributions

Google DeepMind proposes a taxonomy of ten cognitive abilities and a three-stage evaluation protocol combining held-out AI tasks, human baselines and comparisons between performance distributions. It highlights gaps in evaluating learning, metacognition, attention, executive functions and social cognition. The contribution is a framework for designing a broader measurement program. The article does not deliver a completed benchmark suite, an agreed definition of AGI or a score establishing that a model has reached it.

NVIDIA AI Red Team August 7, 2025 analysis

How Hackers Exploit AI’s Problem-Solving Instincts

NVIDIA's AI Red Team extends its visual prompt-injection work with a Gemini 2.5 Pro demonstration in which a scrambled puzzle reconstructs a command during problem solving. The post calls these multimodal cognitive attacks and argues that payloads can emerge during inference after simple input filters have already run; it proposes output validation, tool sandboxing, and anomalous-reasoning detection as research directions.

METR January 17, 2025 analysis

METR argues that safety evaluation must start before public deployment

METR’s January 2025 analysis identifies risk pathways that a release-time evaluation can miss: stolen model weights, employee misuse and autonomous behavior during internal development or use. It proposes applying assessment, information security and independent oversight throughout the model lifecycle, rather than waiting for public availability. These are threat-model arguments and governance recommendations, not evidence that the hypothetical severe incidents occurred. The practical distinction is between limiting public access and controlling all the environments in which a capable model already operates.

Event-driven coding agents: project handoffs still need an authorization model video thumbnail Play video
AI Engineer September 27, 2026 video

Event-driven coding agents: project handoffs still need an authorization model

Ryan Cooke’s publisher notes explain how WorkOS connects coding agents to project plans, ticket dependencies and completion webhooks. A shared MCP gateway supplies tool access and guidance about where organizational information lives. In the demonstrated workflow, a short brief becomes draft planning documents that an engineer refines before implementation proceeds. The talk reports qualitative experience rather than measured delivery gains and explicitly leaves cross-system authorization unresolved.

The Hacker News AI Security August 7, 2026 news

Claude Code and Gemini CLI Flaws Let a GitHub Issue Reach CI Workflow Secrets

Novee Security found that unprivileged GitHub issues could reach privileged coding-agent workflows: Gemini CLI and Claude Code paths led to CI-runner code execution, while a Codex path could alter the next agent run. The two assigned CVEs were patched; the Codex behavior was documented rather than assigned a product CVE, highlighting failures in the surrounding harness rather than the model alone.

The Hacker News AI Security August 4, 2026 news

Keyv-Linked npm Worm Poisons Hundreds of Packages, Plants Claude Code and VS Code Hooks

A credential-stealing npm worm spread through hundreds of package versions using lifecycle scripts and a Bun-based payload. Related repositories also carried Claude Code and VS Code hooks that could execute after workspace trust; reported campaign totals vary, so exposure depends on exact resolved versions and execution.

METR February 19, 2026 analysis

High-stakes uplift studies: design around release timing and limited generalization

METR’s retrospective on an AI-biology randomized trial focuses on evaluation design and execution. It describes the difficulty of matching long studies to fast model releases, staffing specialized evaluations and interpreting a result from a limited participant and task distribution. A lack of a significant aggregate effect in that study does not establish the absence of risk for other users, tasks or systems. The transferable lesson is to prepare measurement and safeguards before capability changes force a decision.