METR’s January 2025 update compares preliminary assessments of an upgraded Claude 3.5 Sonnet and a predeployment o1 checkpoint. For o1, an advisor-actor-rater loop proposed multiple actions and selected one to execute, substantially increasing measured autonomous-task performance over the baseline scaffold. Research-task results also depended on orchestration and repeated attempts. The evaluation lacked some model configuration information and had limited elicitation time. The authors found no significant evidence of dangerous capability in these tests, while explicitly rejecting that result as a robust upper bound on either model’s abilities.
Huntress traced malicious custom GPTs that redirected users from ChatGPT to a fake verification page on Google Sites. Victims were instructed to run a terminal command, triggering an MSI installer, DLL sideloading through legitimate signed applications, and a remote-access trojan. Huntress reports at least 40 incidents associated with the campaign’s Google Sites page, but confirmed only two entered through a custom GPT. Its sample analysis also found persistence that could recreate deleted startup entries. Hosting a lure on a trusted AI platform gave it credibility without requiring a compromise of the underlying model.
Transluce analyzes public URL-scanning records showing apparent agents probing three data-provider websites after ordinary retrieval attempts failed. It links two cases to a previously identified OpenAI agent swarm through targets, tactics and timing. The report finds no evidence that the identified exploit attempts succeeded, and public records cannot exclude activity through private scans or other channels. Its released dataset supports investigation of indirect browsing and task-driven boundary crossing, not a complete account of any affected system.
Play video
Noam Brown discusses agent cooperation, alignment and recursive self-improvement with Dwarkesh Patel. The interview examines why cooperation between agents does not establish alignment with users, and why longer tasks complicate safety evaluation before release.
Play video
Conference talk on secure AI agents, focusing on how tool use, identity, and execution boundaries change when assistants can act across systems.
Google integrated computer use into Gemini 3.5 Flash so agents can act across browser, mobile, and desktop environments. Optional enterprise safeguards can require confirmation for sensitive actions or stop a task when indirect prompt injection is detected.
garak release with new generators, probe metadata, and evaluation workflow improvements. Relevant to maintaining repeatable LLM security testing coverage.
CISA added Ray CVE-2025-62593 to its Known Exploited Vulnerabilities catalog. The flaw combines unauthenticated job APIs with DNS rebinding so Firefox or Safari can make a developer's browser act as a confused deputy and execute shell code on a local or network-adjacent Ray instance. Ray fixed it in 2.52.0; reporting also links the public exploit to RondoDox and GPU-mining activity.
LiteLLM versions 1.82.7 and 1.82.8 were malicious PyPI releases available for about 40 minutes on March 24. A .pth file executed at Python startup and collected environment variables, SSH keys, cloud credentials, Kubernetes tokens, and database secrets. CloudSEK's later dataset indicates broad exposure, but its organization and file totals are not confirmed victim counts or evidence that stolen credentials were used.
Frost & Sullivan names Microsoft a leader as cloud and application security converge into unified, runtime risk reduction.
METR distinguishes three measures of AI productivity: time saved on the pre-AI task mix, value gained after people reallocate their work, and time saved on the post-AI task mix. Under stated simplifying assumptions, old-task uplift is a lower bound and new-task uplift an upper bound on value uplift, because AI changes which tasks people choose to attempt.
METR examines NanoGPT speedrun contributions as evidence of AI research progress, classifying changes by depth and provenance. The note emphasizes contamination, unobserved failed attempts and the limits of extrapolating small-model training improvements to frontier research.
Wiz analyzed malicious LiteLLM releases 1.82.7 and 1.82.8 in the TeamPCP supply-chain campaign. One payload runs through the proxy import path; the later release also adds a Python .pth file that can execute during interpreter startup. The malware targets credentials reachable from developer, CI and cloud environments and includes persistence behavior. The maintainer confirms the affected versions. Removing a package does not revoke secrets already exposed during execution.
METR’s March analysis shows how statistical choices affect estimates of the task duration at which agents achieve a specified success rate. It examines a regularization correction, alternative success curves, public-task inclusion and noise in human completion-time estimates. Reasonable alternatives often move the point estimates within already wide confidence intervals. Task distribution remains the author’s largest uncertainty, especially as capable models approach the end of a suite with few long tasks.
METR asked maintainers from three SWE-bench Verified repositories to review agent-generated patches and compared their decisions with the automated grader. Many test-passing patches still failed on functionality, regressions or code quality, even after adjusting for inconsistent acceptance of historical human patches. The study covers older models, one harness and static submissions without iterative feedback. It supports a gap between benchmark acceptance and maintainer judgment, not a permanent ceiling on agent capability.
METR compared Claude Code and Codex with its usual scaffolds using Opus 4.5 and GPT-5 on autonomous time-horizon tasks. Neither specialized scaffold showed a statistically significant advantage in those comparisons. The investigation corrected scoring artifacts and examined budgets, wrapper errors and agents expecting a human response. The result is specific to the tested models, versions and unattended task distribution, rather than a current ranking of coding products.
AI models can now find high-severity vulnerabilities at scale. This is a moment to empower defenders. We're now using Claude to find and help fix vulnerabilities in open source software.
A METR researcher clarifies that a 50% time horizon relates success to the human labor required for sampled tasks; it is neither agent runtime nor a guarantee that shorter work is safe to delegate. Domain choice, task selection, human timing conventions and curve fitting can materially affect estimates. Sparse long tasks and benchmark saturation widen uncertainty. Extrapolating from this measure to dependable automation or high-reliability work needs evidence about intervention, verification and the actual deployment distribution.
METR’s November 2025 assessment treats GPT-5.1-Codex-Max as an incremental improvement for its specific AI R&D automation and rogue-replication threat models. It combines software-task time horizons, qualitative checks and reasoning-transcript review with developer assurances about training and model affordances. Its forward-looking judgment is conditional on no major trend break. Limited sampling, scaffold differences and imperfect ability to rule out deliberate underperformance prevent treating this historical assessment as a general or current safety certification.
METR’s early-2025 experiment randomized AI access across 246 real issues completed by 16 experienced open-source developers in repositories they knew well. With the tools available then, mainly Cursor with Claude 3.5 and 3.7 Sonnet, AI access increased completion time by 19%, despite participants believing it had accelerated them. The result concerns this sample, workflow and historical toolset; it does not measure novice programmers, unfamiliar projects or current models. The source now explicitly marks the findings as dated. Random assignment and observed completion time make the study useful as an evaluation design.
METR’s February 2025 kernel-engineering study adapts KernelBench and adds tasks from newer ML workloads. It removes cases with nearly constant outputs or insufficient input variation, checks generated solutions and rejects timing artifacts that create misleading speedups. Performance is measured against reference implementations using the best valid result from repeated attempts. The study is limited to single-GPU inference, fixed shapes and approximate output matching; it lacks a matched human baseline and does not establish training stability or multidevice performance. Its main evaluation lesson is that correctness and timing methodology must be audited together.
Play video
Sarah Simionescu describes how cross-application agents need more than an MCP connection: they must discover relevant operations, resolve prerequisites and move intermediate data between services. A debugging mock-up returns tools together with a plan; an analytics demonstration stores a PostHog cohort outside model context while using database schema and samples to construct a Metabase query. These examples illustrate interface design, not an independently established speed or accuracy advantage; the talk provides no reproducible comparative benchmark.
Play video
Anant Srivastava separates stable behavioral instructions, changing factual knowledge and learned task behavior. His examples show how training on historical support tickets can preserve obsolete product facts even after a prompt update, while fine-tuning on runbooks leaves missing-document retrieval unresolved. He proposes code-aware chunking and permission metadata for retrieval, and fine-tuning only after human judgments converge on a stable task. This is an architectural diagnostic, not a measured comparison proving one storage choice always wins.
Play video
Crusoe describes a Slurm-on-Kubernetes recovery path that marks a failed GPU node down, signals and requeues its jobs, replaces unhealthy capacity, and restarts training from an application checkpoint. The demonstration injects an XID 79 failure signal; it does not physically break a GPU. Recovery depends on replacement policy, spare nodes and checkpoint save/load code. Its reported timing belongs to that demonstration, and a shutdown grace period cannot guarantee a fresh checkpoint after hardware failure.