Application architecture, developer workflow, tooling, and production patterns for building AI systems.
AI Engineering
Application architecture, developer workflow, tooling, and production patterns for building AI systems.
- Core concepts for ai engineering
- Useful references, notes, and curated examples
- Practical links back to AI systems and operational risk
- It creates better language for technical and governance discussions
- It helps teams connect theory to deployed workflows
- It supports more repeatable review and decision-making
- Researchers and builders working with AI systems
- Security and governance teams
- Leaders looking for current reference material
Current notes, events, and source material
These items are included because they add useful evidence, framing, implementation detail, or upcoming context for teams working in this area.
The Defender’s Window
OpenAI describes a staged program for AI-assisted defense: use agents to review code and infrastructure, triage alerts, enumerate attack paths, and validate security invariants while retaining strong isolation and least privilege. Its recommended rollout starts with internet-facing services and vulnerability backlogs, moves security review into CI, requires focused fixes and regression tests, and expands from read-only triage to narrowly bounded automation only after teams build evidence and confidence.
Now in preview: Find and fix software vulnerabilities with CodeMender
Google opened a preview of CodeMender, an AI code-security agent delivered through Gemini Enterprise Agent Platform and AI Threat Defense. It is designed to inspect code, identify and validate potentially exploitable defects, and produce targeted fixes, with Google’s specialized Gemini 3.5 Flash Cyber model initially restricted to governments and trusted partners.
Transforming Bedrock Guardrails events into OCSF with CloudWatch
Why it ranks: manually reviewed for hands-on depth; directly applicable to AI security practice; strong implementation or testing value.
AWS provides an implementation guide for a Lambda pipeline that converts Bedrock Guardrails intervention logs into OCSF Detection Findings in the CloudWatch unified data store. It includes field mapping and queries that correlate guardrail events with identity and network activity.
Piloting the world's first double-blind AI evaluations
Google DeepMind, Singapore's AI Safety Institute, OpenMined, AVERI, and MLCommons are piloting an external evaluation in a confidential-computing environment. The evaluator's hidden tests and Google's Gemini Flash Lite weights remain private from one another, reducing benchmark contamination without transferring either sensitive asset.
OpenShell: inspect the runtime controls behind NVIDIA’s agent safety launch
Why it ranks: manually reviewed for hands-on depth; directly applicable to AI security practice; strong implementation or testing value.
NVIDIA’s Open Agent Safety Platform pairs OpenShell’s open-source sandbox runtime with the Sentry hardware reference design. OpenShell’s documentation describes filesystem and process isolation, outbound network policies, and provider credentials resolved only at authorized endpoints. These are inspectable configuration mechanisms, while Sentry’s millisecond quarantine claims remain vendor assertions. Filesystem and process restrictions are fixed when a sandbox is created; network policies and credential attachments can change during operation.
Auditing in the age of (good enough) AI
Why it ranks: manually reviewed for hands-on depth; directly applicable to AI security practice; strong implementation or testing value.
Trail of Bits describes an audit methodology using agents to build a decompiler, static analysis and Lean models before reviewing the Miden VM. Public code and regression checks support security findings and 95 machine-checked proofs, with people reviewing what the theorems establish.
AWS Deception Benchmark tests false positives in AI security review
Why it ranks: manually reviewed for hands-on depth; directly applicable to AI security practice; demonstrates an actionable operational method.
AWS releases a benchmark and methodology for distinguishing vulnerable code from suspicious-looking code protected by effective mitigations. Its single-turn model evaluation compares direct classification with exploit-oriented prompting and exposes tradeoffs between false positives and missed flaws.
OpenLeash Adds a Human Check to Risky AI Agent Actions
SecurityWeek profiles OpenLeash, an authorization layer that evaluates proposed agent actions and can block them or request human approval. The project’s public repository provides a personal runtime using agent hooks and provider traffic, with a decision engine, local history and desktop integration. Its hosted business control plane is outside that repository. Public implementation materials make it inspectable, while the profile offers no independent efficacy benchmark.
Extend Amazon Bedrock Guardrails to Tool Interactions Using the Strands Agents SDK
AWS extends Bedrock Guardrails beyond model input and output with three Strands lifecycle checkpoints: inspect inbound user or retrieved content, validate tool arguments before execution, and inspect tool results before they re-enter the model or leave the system. The implementation mixes service guardrails with lower-latency schema, regex, and allowlist checks.
Securing the future of AI agents
Google DeepMind frames increasingly capable agents as potential insider threats and proposes an AI Control Roadmap that combines access controls with supervisors that inspect plans, reasoning, and actions. Its internal prototype analyzed one million coding-agent tasks, but most flags reflected mistakes or overreach rather than adversarial behavior, making this a control design and measurement guide rather than proof of solved monitoring.
NVIDIA AI Red Team: An Introduction
NVIDIA’s 2023 AI red-team introduction organizes assessments across the ML lifecycle, infrastructure and organizational risk. It combines conventional security testing, model attacks and harm scenarios, then illustrates lifecycle boundaries, privilege separation and tabletop exercises. The framework helps teams identify affected components and assign responsibility across data collection, training, deployment and monitoring.
Authenticate legitimate AI agent traffic with AWS WAF Bot Control
AWS provides a four-step technical guide to authenticating automated agents with Web Bot Authentication: deploy WAF Bot Control, sign requests with Ed25519 HTTP Message Signatures, write rules against verification labels, and monitor attempts through WAF logs and CloudWatch.
Diagnose prompt-cache misses without widening an agent’s tool access
OpenAI’s prompt-caching guidance explains how to compare requests for changes that invalidate shared prefixes, place explicit breakpoints, and preserve tool definitions while changing which tools are callable. GPT-6 can also receive appended reasoning-effort updates without rewriting the earlier prefix. Workload savings still need measurement.
Agents API separates managed orchestration from execution environments
OpenAI’s Agents API beta combines a managed Codex harness with hosted, partner or customer-controlled execution environments. The launch explains context compaction, tool discovery, programmatic calls and subagent coordination, with an implementation example.
Agentic security: Detection and response at machine speed
AWS outlines four areas for securing autonomous workloads: distinct agent identities with temporary scoped credentials, continuous behavioral monitoring, tiered automated containment and traceable delegation across agent teams. It recommends separating sensitive-data access, untrusted inputs and external communication. The article introduces an AWS/SANS framework and links to the longer implementation guidance.
Securing Claude Code: The New Compliance API, Local Visibility, and Identity Governance
This guide maps three complementary control layers for local coding agents: enforced Claude Code settings, Anthropic's Compliance API transcripts for local sessions, and endpoint telemetry such as OpenTelemetry, hooks, configuration inventory, and EDR. It also identifies important gaps: cloud transcripts do not capture unused local plugins or off-platform model sessions, endpoint logs lack business intent, and retained transcripts can become a sensitive data store.
NVIDIA Forms 37-Member Open Secure AI Alliance and Open-Sources NOOA Framework
NVIDIA launched the Open Secure AI Alliance and contributed NOOA, an Apache-2.0 Python framework that represents agent state, capabilities, prompts, and typed contracts in classes with built-in testing and tracing. NVIDIA reports 86.8% on CyberGym L1 with GPT-5.5, blocked network access, and trajectory checks; the repository warns that generated Python can exfiltrate or delete data and that its AST and module filters are not a containment boundary.
AI patch benchmarks need comparable tasks and complete repair checks
Trail of Bits critiques the FLAWED patching benchmark’s aggregation of deliberately misleading prompts, restricted testing and differing model settings. Its reanalysis distinguishes patches that block a supplied exploit from repairs that preserve behavior across the application.
OWASP AIBOM Generator
The OWASP AIBOM Generator creates CycloneDX-aligned inventories for Hugging Face models, visualizes model metadata and dependencies, and scores field completeness. It is a practical starting point for recording model provenance and supply-chain inputs, but an inventory does not establish that a component is safe or that its metadata is accurate.
Run open weight models on Amazon Bedrock in AWS European Sovereign Cloud
AWS explains running Gemma 4 through Bedrock in its European Sovereign Cloud, including regional inference, IAM permissions, audit logging and data-handling controls. Stateful response storage and model-specific retention require separate attention.
Implement custom authentication for tools integration using request Lambda interceptor in AgentCore Gateway
AWS demonstrates an interim AgentCore Gateway pattern for legacy tool APIs: validate the caller's JWT again in a deterministic request Lambda, retrieve a service credential from Secrets Manager, and construct the downstream Basic Auth header without exposing the secret to the model or changing the tool schema. The post explicitly treats this as a bridge to modern authentication, not a target architecture.
OWASP ASI02: tool misuse and exploitation — the definitive security guide
This OWASP ASI02 guide separates accidental and adversarial tool misuse across misinterpreted requests, ignored constraints, poisoned tool descriptions, supply-chain injection, and unsafe multi-tool chains. It connects documented coding-agent incidents to attack surfaces, detection patterns, preventive architecture, and agent-specific containment and forensic questions.
Best Practices for Securing LLM-Enabled Applications
NVIDIA's AI Red Team organizes LLM application risk around prompt injection, information leakage, and probabilistic failure. It recommends treating model output as untrusted, narrowing and parameterizing tool actions, keeping authorization outside the prompt, protecting retrieved-document permissions through the response and logging path, and designing multi-tool workflows to fail closed when an intermediate result is invalid.
Play video
Black Hat Asia 2026 | Model Files → Memory Corruption → RCE: The Triple-Stage AI Attack Chain
Ji'an Zhou and Lei Lu show how a malicious model artifact can move beyond familiar pickle or Lambda-layer deserialization bugs into native memory corruption. Their Black Hat briefing builds an end-to-end three-stage chain from a crafted model file through controlled heap layout and control-flow hijacking to reliable code execution, then evaluates the attack against real inference systems.