AI Engineer · August 24, 2025

Evals Are Not Unit Tests — Ido Pesok, Vercel v0

Evals Are Not Unit Tests — Ido Pesok, Vercel v0 video thumbnail
Why it matters

AI Engineer session on Evals Are Not Unit Tests, presented by Ido Pesok, Vercel v0. It adds practical context for how teams are building and operating AI systems in production.

My takeaway: Evals Are Not Unit Tests — Ido Pesok, Vercel v0 is a tooling signal. The practical read is to watch how AI red-team and evaluation tooling is adding probes, integrations, and workflows that can become repeatable controls.
Keep exploring

More curated notes connected through AI Engineering and Prompt Engineering.

OpenAI News · guide

The Defender’s Window

OpenAI describes a staged program for AI-assisted defense: use agents to review code and infrastructure, triage alerts, enumerate attack paths, and validate security invariants while retaining strong isolation and least privilege. Its recommended rollout starts with internet-facing services and vulnerability backlogs, moves security review into CI, requires focused fixes and regression tests, and expands from read-only triage to narrowly bounded automation only after teams build evidence and confidence.

Google Cloud Security Blog · tool

Now in preview: Find and fix software vulnerabilities with CodeMender

Google opened a preview of CodeMender, an AI code-security agent delivered through Gemini Enterprise Agent Platform and AI Threat Defense. It is designed to inspect code, identify and validate potentially exploitable defects, and produce targeted fixes, with Google’s specialized Gemini 3.5 Flash Cyber model initially restricted to governments and trusted partners.

Google DeepMind Blog · framework

Piloting the world's first double-blind AI evaluations

Google DeepMind, Singapore's AI Safety Institute, OpenMined, AVERI, and MLCommons are piloting an external evaluation in a confidential-computing environment. The evaluator's hidden tests and Google's Gemini Flash Lite weights remain private from one another, reducing benchmark contamination without transferring either sensitive asset.