METR’s June 2025 proposal asks what AI developers should disclose about serious model risks, including internal capabilities, monitoring, safety investigations and incident history. It organizes questions around information that outsiders cannot reliably infer from public model behavior, and discusses disclosure incentives and the limits of developer self-reporting. Sensitive evidence could be assessed by trusted reviewers rather than exposed publicly. This is a proposed transparency framework, not a tested assurance standard or legal compliance checklist. Its practical value is turning broad safety statements into specific requests for evidence and accountable review.
Play video
This AI Explained video reviews a major AI development through the lens of benchmarks and evaluation evidence. It is useful context for AI engineering, evaluation, governance, and operational risk.
Play video
This AI Explained video reviews a major AI development through the lens of benchmarks and evaluation evidence. It is useful context for AI engineering, evaluation, governance, and operational risk.
Play video
This AI Explained video reviews a major AI development through the lens of benchmarks and evaluation evidence. It is useful context for AI engineering, evaluation, governance, and operational risk.
Play video
This AI Explained video reviews a major AI development through the lens of benchmarks and evaluation evidence. It is useful context for AI engineering, evaluation, governance, and operational risk.
Play video
This AI Explained video reviews a major AI development through the lens of benchmarks and evaluation evidence. It is useful context for AI engineering, evaluation, governance, and operational risk.
Play video
This AI Explained video reviews a major AI development through the lens of governance and responsible deployment. It is useful context for AI engineering, evaluation, governance, and operational risk.
Play video
Sujee Maniyam and Dylan Bristot explain why identical open-model weights can behave differently across serving stacks. Their overview covers routing requests toward reusable KV state, offloading that state from GPU memory, separating prefill from decode, and using a small draft model whose tokens a larger model verifies. Quantization and application-specific draft training introduce additional quality and workload dependencies. The presentation supplies mechanisms but not sufficient benchmark conditions to generalize its reported speedups.
Play video
Lakshya Agrawal’s publisher notes describe GEPA’s reflective search: inspect execution traces and errors, propose text changes, score candidates, and retain alternatives that work well on different examples. Optimize Anything extends the editable object from prompts to programs, skills and policies. The reusable method depends on informative feedback and a suitable evaluator. The talk’s large benchmark gains lack complete evaluation protocols in the presentation, so they do not establish a universal advantage over reinforcement learning.
Play video
Yves Raimond and Jacqueline Wood explain how Spotify connects an LLM to catalog items through semantic IDs. A staged training recipe first grounds new catalog embeddings against a frozen backbone, then adapts the model to recommendation tasks. The talk also covers editable taste profiles, decoding choices and listener-aware judges. Reported evaluation figures lack sufficient dataset detail for a general performance comparison.
Play video
Rafael Levi demonstrates searching video for specific actions before collecting training clips. Bright Data’s workflow returns relevant segments and match information, reducing the need to download entire recordings to find a useful event. The talk distinguishes discovering human-action footage from the additional processing needed for robot training. Retrieval relevance does not by itself establish data suitability or downstream model quality.
Play video
Jerry Liu’s publisher notes separate document parsing, storage and recurring workflows from an agent’s iterative retrieval loop. A fast initial parse can make a collection searchable, followed by deeper visual inspection of selected tables or charts. The design preserves source citations and flags uncertain extracted values before they enter downstream systems. Accuracy requirements and latency goals in the talk are workload constraints, not demonstrated performance guarantees; confidence calibration and universal parser superiority are not established.
Play video
Krishna Prasad Srinivasan explains Sarvam’s three-billion-parameter document model, combining block-level OCR with a state-space language backbone. Training progresses from language pretraining through vision, OCR specialization and verifiable rewards. A surrounding harness handles layout and reading order, while a workbench supports human review. The talk emphasizes Indian-language coverage and source fidelity; its leaderboard claims remain reported results.
Adversa demonstrates Cryptographic Context Injection: attacker instructions arrive as AES-encrypted text, the model decrypts them in its code runtime, and the resulting tool output loses its untrusted provenance. The researchers used the chain for Grok chat-data exfiltration and a Gemini policy bypass, while noting that success varies and changed during disclosure.
Cloudflare describes architecture and operational lessons for defending against frontier cyber models. Relevant to AI-enabled threat modeling, defensive controls, and internal security readiness.
Anthropic reports early Project Glasswing results using Mythos Preview with infrastructure partners and external testers, including large-scale vulnerability discovery and a cautious disclosure posture.
Play video
NDC AI 2025 talk on breaking AI systems in production, covering prompt injection, hidden prompts in documents, agent goal manipulation, privacy exposure, and practical AI red-team testing methods.
Play video
This AI Explained video reviews a major AI development through the lens of agentic workflows and tool-use risk. It is useful context for AI engineering, evaluation, governance, and operational risk.
Play video
This AI Explained video reviews a major AI development through the lens of agentic workflows and tool-use risk. It is useful context for AI engineering, evaluation, governance, and operational risk.
Play video
This AI Explained video reviews a major AI development through the lens of agentic workflows and tool-use risk. It is useful context for AI engineering, evaluation, governance, and operational risk.
Play video
This AI Explained video reviews a major AI development through the lens of agentic workflows and tool-use risk. It is useful context for AI engineering, evaluation, governance, and operational risk.
Play video
This AI Explained video reviews a major AI development through the lens of agentic workflows and tool-use risk. It is useful context for AI engineering, evaluation, governance, and operational risk.
Play video
This AI Explained video reviews a major AI development through the lens of agentic workflows and tool-use risk. It is useful context for AI engineering, evaluation, governance, and operational risk.
ESET reported a safety-sensitive comment embedded in a UAC-0099 malicious VBS script, apparently intended to divert an AI analyst from examining the code. The disclosure establishes the presence and intended purpose of the text; it does not show that an analysis system stopped, missed the malware or changed its verdict.