Full Archive · Page 19

Research archive, page 19

Browse entries 433–456 of 1531. Return to the first page to search and filter the complete collection.

Orchestras, Not Factories: How the Fastest Builders Work — Charlie Holtz, Conductor video thumbnail Play video
AI Engineer September 27, 2026 video

Orchestras, Not Factories: How the Fastest Builders Work — Charlie Holtz, Conductor

Charlie Holtz describes Conductor’s approach to coordinating coding agents through shared cloud workspaces and organizational context. His examples include CI-enforced human review for migrations, carefully maintained agent instructions, and a database of working information queried by agents. The presentation combines workflow advice with product demonstrations; it does not measure a general productivity advantage.

You’re Not Thinking Big Enough: Rebuilding Food Systems with AI Agents — Cody Menefee, Firecrawl video thumbnail Play video
AI Engineer September 23, 2026 video

You’re Not Thinking Big Enough: Rebuilding Food Systems with AI Agents — Cody Menefee, Firecrawl

Cody Menefee proposes an AI-assisted pasture-rotation system combining grazing knowledge, field observations and GPS-collar controls. Cameras, drones or satellite imagery could supply changing pasture conditions, while an LLM recommends moves for a farmer to accept or reject. The talk identifies unresolved data and hardware-integration needs; improved farm capacity is a project goal rather than a demonstrated outcome.

METR February 17, 2026 analysis

Transcript-based time savings are an upper bound, not a productivity experiment

A METR exploratory analysis uses 5,305 coding-agent transcripts from seven staff to estimate time saved on AI-assisted tasks. A model estimates counterfactual human effort, while message activity approximates actual human time. The judge was checked against only 34 human estimates, and transcript selection excludes work done without AI. Task substitution, specialization and uncertain success judgments further limit interpretation. The resulting ratios are proposed soft upper bounds, not causal measurements of overall productivity.

METR February 10, 2026 analysis

A computational AI-timeline model makes automation assumptions inspectable

METR’s February 2026 modeling note builds a compact forecast of AI research automation from research labor, compute and software efficiency. Its equations connect capability progress to the share of tasks automated while representing remaining human work as a bottleneck. The accompanying model is useful for examining sensitivity to automation speed and research returns. The headline date follows selected parameter assumptions, many based on judgment, rather than an observed capability threshold. The model omits important effects, including parts of research judgment and the wider economy, and does not establish when real research organizations will become autonomous.

METR July 14, 2025 analysis

Human-time horizons need validation within each evaluation domain

METR’s July 2025 analysis asks whether task-completion horizons transfer beyond its original software evaluations. It reanalyzes nine existing benchmarks, fitting success against human completion time where suitable measurements are available. The relationship is useful in several settings but differs across domains and weakens when proxies such as video length or task fees stand in for human effort. Some fits require constrained assumptions, and agents, datasets and scaffolds are not matched across benchmarks. The study supports checking the metric’s validity locally rather than treating a single horizon as a universal measure of autonomy.