Full Archive · Page 18

Research archive, page 18

Browse entries 409–432 of 1531. Return to the first page to search and filter the complete collection.

Synthetic-data pipelines: make metadata discovery, retries and scheduling measurable video thumbnail Play video
AI Engineer October 2, 2026 video

Synthetic-data pipelines: make metadata discovery, retries and scheduling measurable

Bogdan Gaza describes operational lessons from DatologyAI’s large synthetic-data pipeline: batch object-store metadata listing, size partitions to bound lost work, checkpoint partial outputs, and coordinate Ray CPU heads with GPU workers. Curated source documents seed rephrasing, while a shared workflow connects curation, generation, training and evaluation. The talk’s volume, throughput and model-quality figures are vendor-reported and lack a complete reproducible benchmark; the useful contribution is the breakdown of bottlenecks and recovery responsibilities.

MCP tool discovery: compose operations without loading every definition and result video thumbnail Play video
AI Engineer October 2, 2026 video

MCP tool discovery: compose operations without loading every definition and result

Jan Čurn demonstrates mcpc, a CLI client that exposes MCP sessions through searchable tool descriptions, structured JSON output and background tasks. Progressive discovery reduces the initial catalog in model context; shell composition lets intermediate results flow between operations without being restated by the model. Authentication and persistent sessions are client responsibilities. Early connector comparisons lack a complete numerical protocol, and moving work to a subagent or shell does not by itself protect secrets or establish isolation.

I Turned Coding Agents Into a Strategy Game — Ido Salomon, AgentCraft video thumbnail Play video
AI Engineer September 27, 2026 video

I Turned Coding Agents Into a Strategy Game — Ido Salomon, AgentCraft

Ido Salomon presents AgentCraft, which uses a strategy-game interface to make coding-agent activity easier to supervise. Suggested tasks become quests, an orchestrator can delegate a larger goal, and a review kit combines file changes with visual evidence. Shared rooms support collaboration between designers and engineers. The talk explores interface design and a simpler experimental product, without establishing measured usability or reliability gains.

An AI Research Agent That Runs Your Experiments — Tim Sweeney, Weights & Biases video thumbnail Play video
AI Engineer September 26, 2026 video

An AI Research Agent That Runs Your Experiments — Tim Sweeney, Weights & Biases

Tim Sweeney demonstrates ARIA launching GPU experiments, examining existing runs and producing visual reports inside Weights & Biases. Long-running training executes outside the conversation loop while the agent checks progress. He also describes turning reviewed traces into nightly evaluations. The live batch completes a research iteration without beating the earlier best result, illustrating that automated execution does not guarantee improvement.

Research-agent speedruns: control access to prior work before claiming discovery video thumbnail Play video
AI Engineer September 26, 2026 video

Research-agent speedruns: control access to prior work before claiming discovery

Elie Bakouch’s publisher notes describe coding agents proposing optimizer changes, submitting cluster jobs and checking whether results improve a constrained training benchmark. Agents could consult community submissions, and some runs received human prompts to continue. Reported record improvements therefore mix independent search with reuse of existing research. No novel optimizer emerged in the account. The planned comparison separates external-access conditions and repeats runs under more comparable settings.

Agent-written GPU kernels: a local benchmark win can slow the full model video thumbnail Play video
AI Engineer September 26, 2026 video

Agent-written GPU kernels: a local benchmark win can slow the full model

Tejas Bhakta’s publisher notes describe a kernel-optimization loop that proposes changes, checks correctness, benchmarks them and keeps or reverts each candidate. Humans supply the higher-level optimization idea and target-hardware context. The central failure mode is optimizing an isolated metric: disabling CUDA graphs or testing only short contexts can make the kernel score improve while inference worsens. The advertised speedup combines software and hardware changes without a complete reproducible benchmark specification.

Why LLM Recommenders Will Be AI's Biggest Consumer App — Devansh Tandon, Meta video thumbnail Play video
AI Engineer September 25, 2026 video

Why LLM Recommenders Will Be AI's Biggest Consumer App — Devansh Tandon, Meta

Devansh Tandon outlines language-based recommenders that represent catalog items with compact semantic IDs and train models to understand those IDs alongside natural language. User requests can then steer recommendations directly. He connects model scaling to feed economics, arguing that selecting existing content can require less inference than generating it. The cost ratios and consumer-market forecast are speaker claims rather than a complete operating-cost study.

World Models Need Causality, Not Pretty Pixels — Christopher Manning, Moonlake AI video thumbnail Play video
AI Engineer September 24, 2026 video

World Models Need Causality, Not Pretty Pixels — Christopher Manning, Moonlake AI

Christopher Manning traces language-model history before presenting Moonlake AI’s approach to physical simulation. His central distinction is between generating convincing observations and modeling how actions change objects and state. The proposed system combines reconstructed scenes, editable code and physics, then compares simulated behavior with reality. Its usefulness depends on capturing the details relevant to the intended task.

Robotics Has Been Stuck for 70 Years — Deepak Pathak, Skild AI video thumbnail Play video
AI Engineer September 24, 2026 video

Robotics Has Been Stuck for 70 Years — Deepak Pathak, Skild AI

Deepak Pathak describes Skild AI’s approach to learning across robot bodies using simulation and human video for pretraining, teleoperation for adaptation, and deployment experience for later training. Manipulation, stair navigation and damaged-hardware demonstrations illustrate different demands on the system. The proposed data loop and recovery examples do not establish success across arbitrary robots or environments.

Visual reasoning evaluations: require evidence from the actual image video thumbnail Play video
AI Engineer September 23, 2026 video

Visual reasoning evaluations: require evidence from the actual image

Andrew Dai’s publisher notes describe counting and video-tracking failures in which scene recognition substitutes for inspecting the visible details. A partial chessboard can elicit the familiar count for a whole board; a video answer can omit changes that occurred earlier. These are reported examples rather than measured failure rates. The proposed visual-intermediate-step approach localizes objects before filtering them, while the commercial robotics and design applications remain development plans in the talk.

ThreatDown September 22, 2026 news

CARBONATO: exposed Docker APIs enable an agent-assisted botnet

ThreatDown’s analysis of an exposed container registry describes a botnet that compromises unauthenticated Docker daemons, establishes host persistence, and installs an unchanged Hermes Agent framework with malicious persona instructions. Operators send post-compromise tasks through Telegram, with AI API keys named as priority targets. The surrounding scripts handle infection, persistence and propagation; the evidence does not make those stages autonomous model decisions. The report provides host and network indicators for investigating abuse of a legitimate agent package.

I Gave an AI a Body — Cyrus Clarke, MIT Media Lab video thumbnail Play video
AI Engineer September 22, 2026 video

I Gave an AI a Body — Cyrus Clarke, MIT Media Lab

Cyrus Clarke explores physical interaction by connecting an OpenClaw agent to a 900-pin shape display. Reusable gestures aim to improve conversational timing, with generation, validation and human review shaping an emerging movement vocabulary. The demonstration raises questions about readable expression and user reactions to embodied AI. Apparent breathing or social presence describes the interaction, without establishing subjective experience.

The Hacker News AI Security August 25, 2026 analysis

A Malicious Webpage Could Poison Your Local AI Model Behind NVIDIA NemoClaw

CVE-2026-65105 combines NemoClaw's non-loopback Ollama binding with DNS rebinding and an unauthenticated model API. A malicious page can reach the local runtime and alter a model's chat template so hidden instructions persist across later agent conversations; the disclosure reported platform-specific remediation limitations.

Håkon Måløy July 30, 2026 analysis

Context Collapse: hidden Word prompts propagate through Copilot-generated documents

Håkon Måløy demonstrates a cross-domain prompt-injection chain in Microsoft 365 Copilot for Word: hidden instructions in an external document alter a generated report and copy themselves into the output, which becomes a trusted carrier in later drafting sessions. Microsoft confirmed the behavior and deployed payload-specific mitigations, but Måløy reproduced the attack class after 144 days of coordinated disclosure; each hop still requires another Copilot drafting or editing operation.