Play video
Dylan Couzon demonstrates an offline object-memory application using local detection, embeddings and Qdrant Edge storage. Retrieved sightings include images and timing metadata, showing how a persistent index supplies continuity beyond the current context window. Portability across reasoners or devices assumes the same embedding model; changing that representation needs a separate migration plan. The demo’s latency and footprint are workload-specific, and proposed whole-day assistants and shared memories are extensions rather than validated outcomes.
Play video
Jakub Hojsan uses an API migration to show why a plausible diff can be misread when a coding agent relies on stale model knowledge. A version-change rule triggers retrieval of upstream documentation or changelogs, while query-specific excerpts and retrieval traces make the evidence inspectable without loading entire pages. The method separates tool availability from actually invoking verification. The talk does not establish a fixed knowledge-age gap for every model or the vendor’s claimed search-cost advantage.
Play video
Justin Reock separates utilization, impact and cost when evaluating AI-assisted development. His pipeline example shows how rapid code generation can increase batching when builds and reviews remain slow, while the suggested measurements combine PR size and cycle time with failures, review corrections and developer experience. The talk’s organizational observations are associations and self-reported outcomes, not a causal estimate of AI’s effect. Faster releases or higher token consumption alone do not demonstrate greater business value.
Cleafy’s analysis distinguishes two AI uses in RATHat: the operator panel estimates victim value from stolen SMS messages, while device-side Gemini calls help locate controls when fixed UI automation fails. The malware first requires an Accessibility grant and a successful wireless-debugging pairing path; an operator can then deploy a separate shell-level service. That service can survive app removal until reboot. The analyzed samples do not show an LLM performing fraudulent transfers, and some native capture tools fail on Android 14 and later.
Play video
Zixuan Li introduces GLM-5.2 through its coding and agentic capabilities, adjustable thinking budget and open-weight deployment options. He separates the model from Z Code, a coding harness that also accepts other models. The talk explains the roles of local inference, domain fine-tuning and ecosystem tooling; its benchmark placements are Z.ai’s reported comparisons with incomplete evaluation conditions.
Play video
Suraj Gupta’s publisher notes distinguish agents doing recurring work from agents proposing improvements to that work. Warp’s triage example turns human feedback into a skill-change pull request, while persistent memory retains investigation findings with editing and provenance controls. A separate evaluation loop compares models on recurring task classes. The demonstration does not quantify memory savings, and customer-facing routing evaluations were planned rather than available in the account.
Play video
Kevin Hou describes Antigravity’s move toward model-led teams that create specialist subagents, listen for events and generate task-specific interfaces. Examples include an operating-system kernel demonstration and parallel investigations of evaluation differences. The reported costs and completion times belong to particular demonstrations; generated hypotheses and working applications still require independent review.
Play video
Roland Gavrilescu’s publisher notes propose preserving an agent’s prompts, skills, evaluations, tools and environment choices as a versioned configuration. Failures become tests, repeated procedures become skills, and human feedback defines what counts as useful work. Candidate changes should then face production experiments before promotion. The examples explain an improvement process, but do not provide controlled outcome measurements or establish that a higher-level agent can replace human judgment.
Play video
Erina Karati’s publisher notes use a simulated game village to expose memory failures: agents can retain a topic while losing its source, certainty or implications for later plans. The proposed evaluation records observations, memory writes, retrievals and belief changes across complete scenarios. It freezes the harness and evaluator while searching a limited policy space. A corrected rumor example is illustrative; the talk does not claim repeated evidence of general improvement.
Play video
Daksh Gupta compares likely agent-written and human-written pull requests using Greptile’s enterprise review data. Authorship is inferred from metadata, while reverts, flagged issues and review rounds serve as quality proxies. He reports broadly similar aggregate results with different failure patterns, then describes validating related code and running applications in sandboxes. The observational comparisons do not control all differences in task selection or difficulty.
Play video
Jason Ma’s publisher notes describe using a video-based progress model to find robot failures, then collecting human demonstrations of recovery and fine-tuning the policy. Napkin folding exposes both bad grasps and quality failures that a nominal task-completion measure can miss. The reported 24-hour success rate belongs to a particular folding evaluation, with no sample size supplied in the notes. Transfer to a new site and recovery from unfamiliar mistakes require separate evidence.
Play video
Adit Abraham’s publisher notes explain a document pipeline that combines layout detection, selective vision-model processing and targeted OCR corrections. The key failure mode is a model silently fixing the document itself, such as replacing a printed but incorrect total. Separate representations can support retrieval and structured reasoning, while iterative chart reconstruction can expose extraction mistakes. Reported benchmark improvements are incomplete comparisons; the notes provide no numerical tolerance establishing exact chart recovery.
Play video
Armen Aghajanyan describes Perceptron AI’s effort to combine perception, reasoning and robot control. Task-dependent token routing focuses computation, while zooming, tiling and revisiting video intervals help gather visual evidence. Joint video and control training aims to reduce dependence on teleoperation. The talk leaves parts of the training objective undisclosed and identifies temporal reliability and severe visual disruption as unresolved challenges.
Play video
Merve Noyan’s publisher notes describe using a vision-language model to label images, smaller judges to inspect overlaid boxes, and a task-specific detector for deployment. The workflow exposes two evaluation traps: agreement with generated labels differs from agreement with human ground truth, and filtering can discard too much training data. Standard image augmentations can also change the correct answer. The signature-detection example is qualitative, and reported run costs lack a complete workload specification.
Play video
Aditya Gautam’s publisher notes describe separating video perception, retrieval and review so short content changes remain tied to timestamps and source clips. Production traces help distinguish model errors from tool or retrieval failures before retraining. Small specialized models, caching and frame compression reduce work, but each shortcut can hide relevant content. The talk does not provide achieved accuracy or decision thresholds, and similarity alone does not establish a video’s original creator.
Microsoft’s CVE-2026-85889 record describes a missing-authentication flaw in Azure AI Foundry that allows network-based privilege escalation. The report concerns a managed-service security update; the public record provides limited details about the underlying attack path.
AI has made fundamental changes to the operating environment for cybersecurity. Explore exposure management guidance on recommended controls and take action and stay ahead of cyberthreats.
GreyNoise reports an AI-orchestrated PaperCut campaign using a DeepSeek model through the Codex harness, with at least 440 affected instances across 395 identified organizations. Its observed domain-admin outcomes were a smaller subset, and campaign development preceded the fastest compromises.
See how Microsoft Defender detects and disrupts AI-themed phishing, malware, and multi-stage attacks across the attack chain.
Microsoft examines an AI-assisted business email compromise campaign that used executive impersonation and fake invoices to target finance teams with ACH payment fraud.
New models, trained using NVIDIA Nemotron 3 Ultra, aim to catch rogue agent behavior before it executes, without the latency of large-model review.
Organizations need protection that operates in the gap between discovery and remediation.
Vendor guidance on operationalizing AI-enabled detection and response. Useful as an implementation signal for monitoring, containment, and response workflows around AI-influenced threats.
Play video
This AI Explained video reviews a major AI development through the lens of agentic workflows and tool-use risk. It is useful context for AI engineering, evaluation, governance, and operational risk.