Full Archive · Page 28

Research archive, page 28

Browse entries 649–672 of 1531. Return to the first page to search and filter the complete collection.

METR November 22, 2024 analysis

RE-Bench compares research agents and human experts under explicit compute budgets

METR’s November 2024 RE-Bench release provides seven research-engineering environments, human-expert attempts and agent transcripts. Tasks involve objectives such as optimizing kernels, recovering model performance and fitting scaling laws under specified resources. Results depend strongly on time allocation: agents often benefit from selecting the best of many short attempts, while humans benefit from longer continuous work. The environments offer clearer goals and faster feedback than much real research, and cheating requires inspection. The benchmark measures bounded engineering performance, not autonomous completion of an open-ended research program.

AWS Security Blog July 7, 2026 analysis

Enforce zero data retention on Amazon Bedrock with Bedrock Projects and service control policies

With the introduction of models that require data sharing with third-party providers—such as Claude Fable 5—organizations need a way to centrally enforce data retention policies. Amazon Bedrock gives you control over whether your prompts and model outputs are retained after an inference request completes.