Play video
AI Engineer session on Real world MCPs in GitHub Copilot Agent Mode, presented by Jon Peck, Microsoft. It adds practical context for how teams are building and operating AI systems in production.
Play video
AI Engineer session on Break It 'Til You Make It: Building the Self-Improving Stack for AI Agents - Aparna Dhinakaran. It adds practical context for how teams are building and operating AI systems in production.
Play video
AI Engineer session on The Demo I Wish I'd Had: OpenAI's Agents SDK... serverless! - Brook Riggio. It adds practical context for how teams are building and operating AI systems in production.
Play video
AI Engineer session on Breaking the Chain: Agent Continuations for Resumable AI Workflows - Greg Benson. It adds practical context for how teams are building and operating AI systems in production.
Play video
AI Engineer session on The Future of Qwen: A Generalist Agent Model, presented by Junyang Lin, Alibaba Qwen. It adds practical context for how teams are building and operating AI systems in production.
Play video
AI Engineer session on MCP: Origins and Requests For Startups, presented by Theodora Chu, Model Context Protocol PM, Anthropic. It adds practical context for how teams are building and operating AI systems in production.
Play video
AI Engineer session on Stateful Agents, presented by Full Workshop with Charles Packer of Letta and MemGPT. It adds practical context for how teams are building and operating AI systems in production.
Play video
AI Engineer session on Multi model multimodal and multi agent innovations in Azure AI: Cedric Vidal. It adds practical context for how teams are building and operating AI systems in production.
Play video
AI Engineer session on Building Agents with Model Context Protocol - Full Workshop with Mahesh Murag of Anthropic. It adds practical context for how teams are building and operating AI systems in production.
Palo Alto Networks' Unit 42 says a Chinese-speaking threat actor used DeepSeek through the open-source Hermes Agent framework to launch attacks autonomously. After an initial Telegram instruction, the agent found internet-facing systems and selected public exploits.
Claude Mythos Preview is a new general-purpose language model that is strikingly capable at computer security tasks. This post provides technical details for researchers and practitioners who want to understand exactly how we have been testing this model, and what we have found over the past month.
In June, we revealed that we'd set up a small shop in our San Francisco office run by an AI shopkeeper. It did not do particularly well. We made some adjustments for phase two of Project Vend. The idea of an AI running a business doesn't seem as far-fetched as it once did.
METR’s November 2024 RE-Bench release provides seven research-engineering environments, human-expert attempts and agent transcripts. Tasks involve objectives such as optimizing kernels, recovering model performance and fitting scaling laws under specified resources. Results depend strongly on time allocation: agents often benefit from selecting the best of many short attempts, while humans benefit from longer continuous work. The environments offer clearer goals and faster feedback than much real research, and cheating requires inspection. The benchmark measures bounded engineering performance, not autonomous completion of an open-ended research program.
With the introduction of models that require data sharing with third-party providers—such as Claude Fable 5—organizations need a way to centrally enforce data retention policies. Amazon Bedrock gives you control over whether your prompts and model outputs are retained after an inference request completes.
Play video
AI Engineer session on AgentCraft: Putting the Orc in Orchestration, presented by Ido Salomon. It adds practical context for how teams are building and operating AI systems in production.
Play video
AI Engineer session on Agents need more than a chat - Jacob Lauritzen, CTO Legora. It adds practical context for how teams are building and operating AI systems in production.
Play video
AI Engineer session on Full Workshop: Build Your Own Deep Research Agents - Louis-François Bouchard, Paul Iusztin, Samridhi. It adds practical context for how teams are building and operating AI systems in production.
Play video
AI Engineer session on The Future of MCP, presented by David Soria Parra, Anthropic. It adds practical context for how teams are building and operating AI systems in production.
Play video
AI Engineer session on From Chaos to Choreography: Multi-Agent Orchestration Patterns That Actually Work, presented by Sandipan Bhaumik. It adds practical context for how teams are building and operating AI systems in production.
Play video
AI Engineer session on Agentic Engineering: Working With AI, Not Just Using It, presented by Brendan O'Leary. It adds practical context for how teams are building and operating AI systems in production.
Play video
AI Engineer session on Platforms for Humans and Machines: Engineering for the Age of Agents, presented by Juan Herreros Elorza. It adds practical context for how teams are building and operating AI systems in production.
Play video
AI Engineer session on Your MCP Server is Bad (and you should feel bad) - Jeremiah Lowin, Prefect. It adds practical context for how teams are building and operating AI systems in production.
Play video
AI Engineer session on Spec-Driven Development: Agentic Coding at FAANG Scale and Quality, presented by Al Harris, Amazon Kiro. It adds practical context for how teams are building and operating AI systems in production.
Play video
AI Engineer session on Why Agent Hype can fall short of reality, presented by Joel Becker, METR. It adds practical context for how teams are building and operating AI systems in production.