Contributing to U.K. financial sector resilience as a critical third party
The U.K. Treasury has designated Google Cloud EMEA as a critical third party (CTP) to the U.K. financial sector under the CTP regime. Here’s how that helps you.
Browse entries 385–408 of 1531. Return to the first page to search and filter the complete collection.
The U.K. Treasury has designated Google Cloud EMEA as a critical third party (CTP) to the U.K. financial sector under the CTP regime. Here’s how that helps you.
Our flagship Google for Startups program, Gemini Startup Forum: Cybersecurity, has selected its first 33 trailblazing startups.
Google DeepMind and partners announce a $10M funding call for multi-agent safety research.
Anthropic and Verizon mapping of AI-enabled cyber activity to MITRE ATT&CK. Relevant to threat modeling, red-team scenario design, and structured reporting of AI-enabled operations.
OpenAI's opt-in Advanced Account Security applies to ChatGPT and Codex. It replaces password login with passkeys or FIDO security keys, disables email and SMS recovery, shortens sessions, adds login alerts and session management, and automatically excludes conversations from model training. The stronger recovery model also means support cannot restore access for an enrolled user.
Play video
This AI Explained video reviews a major AI development through the lens of benchmarks and evaluation evidence. It is useful context for AI engineering, evaluation, governance, and operational risk.
Play video
This AI Explained video reviews a major AI development through the lens of governance and responsible deployment. It is useful context for AI engineering, evaluation, governance, and operational risk.
Three METR staff spent two hours simulating how their research organization might work with agents capable of roughly 200-hour tasks. The exercise advanced through two hypothetical workdays and exposed choices about prioritization, delegation, context and review. Participants used present-day research needs but assumed future agent capabilities. The account is a planning exercise for identifying organizational constraints, not an experiment measuring productivity or evidence that such agents are available.
OpenAI announced plans to acquire Promptfoo, highlighting automated AI security testing, red teaming, and evaluation as core enterprise requirements.
In a collaboration with researchers at Mozilla, Claude Opus 4.6 discovered 22 Firefox vulnerabilities over the course of two weeks.
Play video
This AI Explained video reviews a major AI development through the lens of governance and responsible deployment. It is useful context for AI engineering, evaluation, governance, and operational risk.
A METR researcher tested Opus 4.6 on simplified terminal reimplementations of two existing games. Initial playthroughs found recognizable gameplay alongside missing or broken mechanics, without a predefined scoring rubric. A July follow-up uncovered additional problems that the original inspection missed. The tasks benefited from documented rules and reduced scope, and more complete versions failed on initial attempts. This is a qualitative study of deliverable inspection, not a measure of general game-development productivity.
METR explains why its follow-up developer experiment did not provide a reliable estimate of current AI productivity gains. Developers increasingly avoided participation or withheld tasks they did not want to perform without AI; lower compensation added another selection concern. Concurrent agent use also complicated time accounting. The raw results suggested possible speedups but had wide uncertainty and omitted important users and tasks, prompting changes to the study design.
METR’s January 2026 update expands its time-horizon suite from 170 to 228 tasks, repairs or removes problematic tasks, and moves evaluation infrastructure from Vivaria to Inspect. Re-estimated model results generally remain within earlier confidence intervals, but task composition changes the fitted recent trend. More long tasks improve coverage, yet only five of the 31 tasks estimated at eight hours or more have measured human baselines. The metric measures success against human task duration, not uninterrupted agent runtime.
Play video
This AI Explained video reviews a major AI development through the lens of governance and responsible deployment. It is useful context for AI engineering, evaluation, governance, and operational risk.
METR’s August 2025 follow-up contrasts algorithmic scoring with human review of repository work. It evaluated 18 tasks drawn from two projects using Claude 3.7 Sonnet and a basic agent scaffold. Automated scoring credited some solutions, while none of the 15 manually assessed submissions met the study’s holistic mergeability standard. Missing documentation, inadequate tests and other maintenance requirements help explain the gap. The selected tasks, limited elicitation and small sample prevent a general estimate of coding-agent usefulness; the manual assessment also was not a direct measurement of maintainers’ repair time.
METR’s March 2025 research measures how agent success varies with the time human experts need to complete software and reasoning tasks. A fitted success curve yields a horizon at a chosen reliability level, while hierarchical resampling estimates uncertainty across task families, tasks and attempts. The original paper also extrapolated historical growth, but that forecast depends on the trend continuing and on transfer beyond the measured tasks. The source distinguishes its dated static results from a live chart. A 50% horizon describes a success threshold, not dependable execution of a project of that duration.
METR’s early-2025 DeepSeek-R1 assessment uses autonomous software tasks and research-engineering environments through a third-party model provider. Transcript inspection found hallucinated tool results, malformed calls and difficulty recovering from mistakes, showing how interface behavior can limit measured performance. The report also acknowledges modest elicitation, incomplete checks for strategic underperformance and potentially biased exclusion of server-error runs. Published example transcripts were selected to illustrate behavior, not estimate its prevalence. These results describe a historical model-and-scaffold combination rather than an upper bound on open-weight model capability.
METR’s early-2025 DeepSeek-V3 report combines autonomous-task testing, research-engineering tasks and a held-out knowledge check. The authors observed repeated actions, hallucinated tool outputs and misread numerical feedback, while scaffold changes substantially improved performance. They explicitly caution that elicitation was limited, curated transcripts were not representative and excluded server-error runs could inflate scores. The study did not robustly test deliberate underperformance. Its historical capability comparison is therefore conditional on the provider, scaffold and scoring choices, with useful lessons for designing more auditable evaluations.
Microsoft summarizes lessons from red teaming more than one hundred generative AI products, emphasizing system-level testing, human expertise, and automation.
Play video
This AI Explained video reviews a major AI development through the lens of agentic workflows and tool-use risk. It is useful context for AI engineering, evaluation, governance, and operational risk.
Play video
This AI Explained video reviews a major AI development through the lens of agentic workflows and tool-use risk. It is useful context for AI engineering, evaluation, governance, and operational risk.
Play video
This AI Explained video reviews a major AI development through the lens of agentic workflows and tool-use risk. It is useful context for AI engineering, evaluation, governance, and operational risk.
Play video
This AI Explained video reviews a major AI development through the lens of agentic workflows and tool-use risk. It is useful context for AI engineering, evaluation, governance, and operational risk.