Full Archive · Page 20

Research archive, page 20

Browse entries 457–480 of 1531. Return to the first page to search and filter the complete collection.

METR June 27, 2025 analysis

METR proposes evidence questions for frontier-model risk disclosure

METR’s June 2025 proposal asks what AI developers should disclose about serious model risks, including internal capabilities, monitoring, safety investigations and incident history. It organizes questions around information that outsiders cannot reliably infer from public model behavior, and discusses disclosure incentives and the limits of developer self-reporting. Sensitive evidence could be assessed by trusted reviewers rather than exposed publicly. This is a proposed transparency framework, not a tested assurance standard or legal compliance checklist. Its practical value is turning broad safety statements into specific requests for evidence and accountable review.

Open-model serving: evaluate cache locality, decoding and quantization together video thumbnail Play video
AI Engineer October 3, 2026 video

Open-model serving: evaluate cache locality, decoding and quantization together

Sujee Maniyam and Dylan Bristot explain why identical open-model weights can behave differently across serving stacks. Their overview covers routing requests toward reusable KV state, offloading that state from GPU memory, separating prefill from decode, and using a small draft model whose tokens a larger model verifies. Quantization and application-specific draft training introduce additional quality and workload dependencies. The presentation supplies mechanisms but not sufficient benchmark conditions to generalize its reported speedups.

GEPA: use execution feedback to improve prompts and agent programs video thumbnail Play video
AI Engineer September 26, 2026 video

GEPA: use execution feedback to improve prompts and agent programs

Lakshya Agrawal’s publisher notes describe GEPA’s reflective search: inspect execution traces and errors, propose text changes, score candidates, and retain alternatives that work well on different examples. Optimize Anything extends the editable object from prompts to programs, skills and policies. The reusable method depends on informative feedback and a suitable evaluator. The talk’s large benchmark gains lack complete evaluation protocols in the presentation, so they do not establish a universal advantage over reinforcement learning.

Teaching LLMs to Speak Spotify — Yves Raimond & Jacqueline Wood, Spotify video thumbnail Play video
AI Engineer September 25, 2026 video

Teaching LLMs to Speak Spotify — Yves Raimond & Jacqueline Wood, Spotify

Yves Raimond and Jacqueline Wood explain how Spotify connects an LLM to catalog items through semantic IDs. A staged training recipe first grounds new catalog embeddings against a frozen backbone, then adapts the model to recommendation tasks. The talk also covers editable taste profiles, decoding choices and listener-aware judges. Reported evaluation figures lack sufficient dataset detail for a general performance comparison.

Physical AI's Next Bottleneck Is Finding the Right Video — Rafael Levi, Bright Data video thumbnail Play video
AI Engineer September 24, 2026 video

Physical AI's Next Bottleneck Is Finding the Right Video — Rafael Levi, Bright Data

Rafael Levi demonstrates searching video for specific actions before collecting training clips. Bright Data’s workflow returns relevant segments and match information, reducing the need to download entire recordings to find a useful event. The talk distinguishes discovering human-action footage from the additional processing needed for robot training. Retrieval relevance does not by itself establish data suitability or downstream model quality.

Document context layers: combine fast parsing with selective source inspection video thumbnail Play video
AI Engineer September 23, 2026 video

Document context layers: combine fast parsing with selective source inspection

Jerry Liu’s publisher notes separate document parsing, storage and recurring workflows from an agent’s iterative retrieval loop. A fast initial parse can make a collection searchable, followed by deeper visual inspection of selected tables or charts. The design preserves source citations and flags uncertain extracted values before they enter downstream systems. Accuracy requirements and latency goals in the talk are workload constraints, not demonstrated performance guarantees; confidence calibration and universal parser superiority are not established.

From Scratch to SOTA: Training a 3B State-Space Vision Model — Krishna Prasad Srinivasan, Sarvam video thumbnail Play video
AI Engineer September 23, 2026 video

From Scratch to SOTA: Training a 3B State-Space Vision Model — Krishna Prasad Srinivasan, Sarvam

Krishna Prasad Srinivasan explains Sarvam’s three-billion-parameter document model, combining block-level OCR with a state-space language backbone. Training progresses from language pretraining through vision, OCR specialization and verifiable rewards. A surrounding harness handles layout and reading order, while a workbench supports human review. The talk emphasizes Indian-language coverage and source fidelity; its leaderboard claims remain reported results.

Adversa AI Trusted AI Blog August 20, 2026 analysis

Zero-click Grok data theft: Cryptographic Context Injection attack leaks chat histories

Adversa demonstrates Cryptographic Context Injection: attacker instructions arrive as AES-encrypted text, the model decrypts them in its code runtime, and the resulting tool output loses its untrusted provenance. The researchers used the chain for Grok chat-data exfiltration and a Gemini policy bypass, while noting that success varies and changed during disclosure.