Full Archive · Page 15

Research archive, page 15

Browse entries 337–360 of 1531. Return to the first page to search and filter the complete collection.

Local agent memory: separate persistence, retrieval and permission to share video thumbnail Play video
AI Engineer October 2, 2026 video

Local agent memory: separate persistence, retrieval and permission to share

Dylan Couzon demonstrates an offline object-memory application using local detection, embeddings and Qdrant Edge storage. Retrieved sightings include images and timing metadata, showing how a persistent index supplies continuity beyond the current context window. Portability across reasoners or devices assumes the same embedding model; changing that representation needs a separate migration plan. The demo’s latency and footprint are workload-specific, and proposed whole-day assistants and shared memories are extensions rather than validated outcomes.

Ground coding-agent dependency reviews in current upstream evidence video thumbnail Play video
AI Engineer October 2, 2026 video

Ground coding-agent dependency reviews in current upstream evidence

Jakub Hojsan uses an API migration to show why a plausible diff can be misread when a coding agent relies on stale model knowledge. A version-change rule triggers retrieval of upstream documentation or changelogs, while query-specific excerpts and retrieval traces make the evidence inspectable without loading entire pages. The method separates tool availability from actually invoking verification. The talk does not establish a fixed knowledge-age gap for every model or the vendor’s claimed search-cost advantage.

Measure AI development impact beyond usage and pull-request throughput video thumbnail Play video
AI Engineer September 30, 2026 video

Measure AI development impact beyond usage and pull-request throughput

Justin Reock separates utilization, impact and cost when evaluating AI-assisted development. His pipeline example shows how rapid code generation can increase batching when builds and reviews remain slow, while the suggested measurements combine PR size and cycle time with failures, review corrections and developer experience. The talk’s organizational observations are associations and self-reported outcomes, not a causal estimate of AI’s effect. Faster releases or higher token consumption alone do not demonstrate greater business value.

Cleafy September 28, 2026 news

RATHat: model-assisted targeting and UI recovery rely on an existing Android foothold

Cleafy’s analysis distinguishes two AI uses in RATHat: the operator panel estimates victim value from stolen SMS messages, while device-side Gemini calls help locate controls when fixed UI automation fails. The malware first requires an Accessibility grant and a successful wireless-debugging pairing path; an operator can then deploy a separate shell-level service. That service can survive app removal until reboot. The analyzed samples do not show an LLM performing fraudulent transfers, and some native capture tools fail on Android 14 and later.

GLM-5.2: Open Weights, Near-Frontier Intelligence — Zixuan Li, Z.ai video thumbnail Play video
AI Engineer September 27, 2026 video

GLM-5.2: Open Weights, Near-Frontier Intelligence — Zixuan Li, Z.ai

Zixuan Li introduces GLM-5.2 through its coding and agentic capabilities, adjustable thinking budget and open-weight deployment options. He separates the model from Z Code, a coding harness that also accepts other models. The talk explains the roles of local inference, domain fine-tuning and ecosystem tooling; its benchmark placements are Z.ai’s reported comparisons with incomplete evaluation conditions.

Improving agent skills and memory through reviewed changes and task evaluations video thumbnail Play video
AI Engineer September 27, 2026 video

Improving agent skills and memory through reviewed changes and task evaluations

Suraj Gupta’s publisher notes distinguish agents doing recurring work from agents proposing improvements to that work. Warp’s triage example turns human feedback into a skill-change pull request, while persistent memory retains investigation findings with editing and provenance controls. A separate evaluation loop compares models on recurring task classes. The demonstration does not quantify memory savings, and customer-facing routing evaluations were planned rather than available in the account.

Get Out of the Model's Way — Kevin Hou, Google Antigravity video thumbnail Play video
AI Engineer September 27, 2026 video

Get Out of the Model's Way — Kevin Hou, Google Antigravity

Kevin Hou describes Antigravity’s move toward model-led teams that create specialist subagents, listen for events and generate task-specific interfaces. Examples include an operating-system kernel demonstration and parallel investigations of evaluation differences. The reported costs and completion times belong to particular demonstrations; generated hypotheses and working applications still require independent review.

Agent improvement loops: version the whole configuration and test user outcomes video thumbnail Play video
AI Engineer September 26, 2026 video

Agent improvement loops: version the whole configuration and test user outcomes

Roland Gavrilescu’s publisher notes propose preserving an agent’s prompts, skills, evaluations, tools and environment choices as a versioned configuration. Failures become tests, repeated procedures become skills, and human feedback defines what counts as useful work. Candidate changes should then face production experiments before promotion. The examples explain an improvement process, but do not provide controlled outcome measurements or establish that a higher-level agent can replace human judgment.

Long-running agent memory: test provenance, uncertainty and privacy together video thumbnail Play video
AI Engineer September 26, 2026 video

Long-running agent memory: test provenance, uncertainty and privacy together

Erina Karati’s publisher notes use a simulated game village to expose memory failures: agents can retain a topic while losing its source, certainty or implications for later plans. The proposed evaluation records observations, memory writes, retrievals and belief changes across complete scenarios. It freezes the harness and evaluator while searching a limited policy space. A corrected rumor example is illustrative; the talk does not claim repeated evidence of general improvement.

AI-Generated Code Is Already Competing With Human Code — Daksh Gupta, Greptile video thumbnail Play video
AI Engineer September 25, 2026 video

AI-Generated Code Is Already Competing With Human Code — Daksh Gupta, Greptile

Daksh Gupta compares likely agent-written and human-written pull requests using Greptile’s enterprise review data. Authorship is inferred from metadata, while reverts, flagged issues and review rounds serve as quality proxies. He reports broadly similar aggregate results with different failure patterns, then describes validating related code and running applications in sandboxes. The observational comparisons do not control all differences in task selection or difficulty.

Robot reliability: evaluate recovery and new-site transfer separately video thumbnail Play video
AI Engineer September 24, 2026 video

Robot reliability: evaluate recovery and new-site transfer separately

Jason Ma’s publisher notes describe using a video-based progress model to find robot failures, then collecting human demonstrations of recovery and fine-tuning the policy. Napkin folding exposes both bad grasps and quality failures that a nominal task-completion measure can miss. The reported 24-hour success rate belongs to a particular folding evaluation, with no sample size supplied in the notes. Transfer to a new site and recovery from unfamiliar mistakes require separate evidence.

Document ingestion: correct OCR without rewriting the source video thumbnail Play video
AI Engineer September 23, 2026 video

Document ingestion: correct OCR without rewriting the source

Adit Abraham’s publisher notes explain a document pipeline that combines layout detection, selective vision-model processing and targeted OCR corrections. The key failure mode is a model silently fixing the document itself, such as replacing a printed but incorrect total. Separate representations can support retrieval and structured reasoning, while iterative chart reconstruction can expose extraction mistakes. Reported benchmark improvements are incomplete comparisons; the notes provide no numerical tolerance establishing exact chart recovery.

From VLM/VLA's to Embodied Agents — Armen Aghajanyan, Perceptron AI video thumbnail Play video
AI Engineer September 23, 2026 video

From VLM/VLA's to Embodied Agents — Armen Aghajanyan, Perceptron AI

Armen Aghajanyan describes Perceptron AI’s effort to combine perception, reasoning and robot control. Task-dependent token routing focuses computation, while zooming, tiling and revisiting video intervals help gather visual evidence. Joint video and control training aims to reduce dependence on teleoperation. The talk leaves parts of the training objective undisclosed and identifies temporal reliability and severe visual disruption as unresolved challenges.

VLM-generated labels: judge annotations and preserve task meaning during training video thumbnail Play video
AI Engineer September 23, 2026 video

VLM-generated labels: judge annotations and preserve task meaning during training

Merve Noyan’s publisher notes describe using a vision-language model to label images, smaller judges to inspect overlaid boxes, and a task-specific detector for deployment. The workflow exposes two evaluation traps: agreement with generated labels differs from agreement with human ground truth, and filtering can discard too much training data. Standard image augmentations can also change the correct answer. The signature-detection example is qualitative, and reported run costs lack a complete workload specification.

Video moderation pipelines: preserve brief events and audit the evaluator video thumbnail Play video
AI Engineer September 23, 2026 video

Video moderation pipelines: preserve brief events and audit the evaluator

Aditya Gautam’s publisher notes describe separating video perception, retrieval and review so short content changes remain tied to timestamps and source clips. Production traces help distinguish model errors from tool or retrieval failures before retraining. Small specialized models, caching and frame compression reduce work, but each shortcut can hide relevant content. The talk does not provide achieved accuracy or decision thresholds, and similarity alone does not establish a video’s original creator.