Full Archive · Page 17

Research archive, page 17

Browse entries 385–408 of 1531. Return to the first page to search and filter the complete collection.

OpenAI News April 30, 2026 news

Introducing Advanced Account Security

OpenAI's opt-in Advanced Account Security applies to ChatGPT and Codex. It replaces password login with passkeys or FIDO security keys, disables email and SMS recovery, shortens sessions, adds login alerts and session management, and automatically excludes conversations from model training. The stronger recovery model also means support cannot restore access for an enrolled user.

METR March 19, 2026 analysis

Use tabletop simulations to find bottlenecks before expanding agent autonomy

Three METR staff spent two hours simulating how their research organization might work with agents capable of roughly 200-hour tasks. The exercise advanced through two hypothetical workdays and exposed choices about prioritization, delegation, context and review. Participants used present-day research needs but assumed future agent capabilities. The account is a planning exercise for identifying organizational constraints, not an experiment measuring productivity or evidence that such agents are available.

METR March 3, 2026 analysis

Playable agent demos need deeper inspection and explicit acceptance criteria

A METR researcher tested Opus 4.6 on simplified terminal reimplementations of two existing games. Initial playthroughs found recognizable gameplay alongside missing or broken mechanics, without a predefined scoring rubric. A July follow-up uncovered additional problems that the original inspection missed. The tasks benefited from documented rules and reduced scope, and more complete versions failed on initial attempts. This is a qualitative study of deliverable inspection, not a measure of general game-development productivity.

METR February 24, 2026 analysis

Developer-productivity studies: AI adoption can bias participation and task selection

METR explains why its follow-up developer experiment did not provide a reliable estimate of current AI productivity gains. Developers increasingly avoided participation or withheld tasks they did not want to perform without AI; lower compensation added another selection concern. Concurrent agent use also complicated time accounting. The raw results suggested possible speedups but had wide uncertainty and omitted important users and tasks, prompting changes to the study design.

METR January 29, 2026 analysis

Time Horizon 1.1: version task composition and scaffolding when comparing capability trends

METR’s January 2026 update expands its time-horizon suite from 170 to 228 tasks, repairs or removes problematic tasks, and moves evaluation infrastructure from Vivaria to Inspect. Re-estimated model results generally remain within earlier confidence intervals, but task composition changes the fitted recent trend. More long tasks improve coverage, yet only five of the 31 tasks estimated at eight hours or more have measured human baselines. The metric measures success against human task duration, not uninterrupted agent runtime.

METR August 13, 2025 analysis

Passing repository tests did not establish that AI patches were mergeable

METR’s August 2025 follow-up contrasts algorithmic scoring with human review of repository work. It evaluated 18 tasks drawn from two projects using Claude 3.7 Sonnet and a basic agent scaffold. Automated scoring credited some solutions, while none of the 15 manually assessed submissions met the study’s holistic mergeability standard. Missing documentation, inadequate tests and other maintenance requirements help explain the gap. The selected tasks, limited elicitation and small sample prevent a general estimate of coding-agent usefulness; the manual assessment also was not a direct measurement of maintainers’ repair time.

METR March 19, 2025 analysis

METR’s time-horizon method ties agent success to human task duration

METR’s March 2025 research measures how agent success varies with the time human experts need to complete software and reasoning tasks. A fitted success curve yields a horizon at a chosen reliability level, while hierarchical resampling estimates uncertainty across task families, tasks and attempts. The original paper also extrapolated historical growth, but that forecast depends on the trend continuing and on transfer beyond the measured tasks. The source distinguishes its dated static results from a live chart. A 50% horizon describes a success threshold, not dependable execution of a project of that duration.

METR March 5, 2025 analysis

DeepSeek-R1 evaluation highlights tool-interface failures and elicitation limits

METR’s early-2025 DeepSeek-R1 assessment uses autonomous software tasks and research-engineering environments through a third-party model provider. Transcript inspection found hallucinated tool results, malformed calls and difficulty recovering from mistakes, showing how interface behavior can limit measured performance. The report also acknowledges modest elicitation, incomplete checks for strategic underperformance and potentially biased exclusion of server-error runs. Published example transcripts were selected to illustrate behavior, not estimate its prevalence. These results describe a historical model-and-scaffold combination rather than an upper bound on open-weight model capability.

METR February 12, 2025 analysis

DeepSeek-V3 assessment shows how scaffold and error handling shape evaluation

METR’s early-2025 DeepSeek-V3 report combines autonomous-task testing, research-engineering tasks and a held-out knowledge check. The authors observed repeated actions, hallucinated tool outputs and misread numerical feedback, while scaffold changes substantially improved performance. They explicitly caution that elicitation was limited, curated transcripts were not representative and excluded server-error runs could inflate scores. The study did not robustly test deliberate underperformance. Its historical capability comparison is therefore conditional on the provider, scaffold and scoring choices, with useful lessons for designing more auditable evaluations.