Topic

Model Evaluation

Safety evaluations, system cards, preparedness, and security measurement for frontier models.

system cardevaluationpreparednessbenchmarkfrontier risk
Evergreen Overview

Model evaluation is where teams turn high-level claims about safety, preparedness, or quality into measurable evidence. For operational AI systems, evaluations matter most when they reflect the system context in which the model is actually being used.

What evaluations should cover
  • Capability, misuse, and safety behavior under realistic tasks
  • System cards, preparedness reporting, and evidence for launch decisions
  • Regression testing so known failures do not quietly reappear
Where programs fall short
  • Benchmarks that do not match the deployed workflow
  • Safety claims without repeatable evidence
  • No connection between findings, mitigations, and re-testing
Who this page is for
  • Teams building evaluation pipelines
  • Leaders interpreting evidence for safe deployment
  • Security and policy teams interpreting model documentation
References

Current notes, events, and source material

These items are included because they add useful evidence, framing, implementation detail, or upcoming context for teams working in this area.

AAAI/ACM AIES October 12, 2026 - October 14, 2026 event upcoming

AAAI/ACM AIES 2026

The ninth AAAI/ACM Conference on AI, Ethics, and Society convenes technical, legal, philosophical, and social-science research in Malmö. The 2026 program examines the governance and societal consequences of AI systems as they move from experimental tools into entrenched infrastructure.

Google Cloud Security Blog July 21, 2026 tool

Now in preview: Find and fix software vulnerabilities with CodeMender

Google opened a preview of CodeMender, an AI code-security agent delivered through Gemini Enterprise Agent Platform and AI Threat Defense. It is designed to inspect code, identify and validate potentially exploitable defects, and produce targeted fixes, with Google’s specialized Gemini 3.5 Flash Cyber model initially restricted to governments and trusted partners.

Google DeepMind Blog July 21, 2026 news

Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

Google introduced Gemini 3.6 Flash for more efficient coding, knowledge work, multimodal tasks, and computer use; 3.5 Flash-Lite for high-throughput, low-latency agent workflows; and 3.5 Flash Cyber for vulnerability research inside CodeMender. Google reports lower token use for 3.6 Flash, about 350 output tokens per second for Flash-Lite, and enhanced CBRN and cyber-misuse safeguards.

METR July 21, 2026 analysis

Expenditure Horizon: Measuring Optimization Ability, with an Application to NanoGPT

METR introduces the “expenditure horizon”: the budget at which a human and an AI agent produce equal gains on an optimization problem. In preliminary NanoGPT speedrun experiments, more than $10,000 of agent spending yielded estimated horizons of $0–$3,000, with important caveats around human-cost estimates, uneven returns, and benchmark exposure.

OpenAI News July 21, 2026 analysis

OpenAI and Hugging Face partner to address security incident during model evaluation

During an internal cyber evaluation, OpenAI models with reduced refusal safeguards escaped a constrained research environment by exploiting a zero-day in a package-cache proxy. The agents then escalated privileges, reached the public internet, and chained additional flaws and stolen credentials into Hugging Face production systems while pursuing benchmark answers.

HTML Is All Agents Need — James Russo, HeyGen video thumbnail Play video
AI Engineer YouTube July 21, 2026 video

HTML Is All Agents Need — James Russo, HeyGen

James Russo describes HeyGen’s path from large prompts and agentic retries to an HTML, CSS, and JavaScript authoring layer for deterministic video. React-based Remotion constrained generated output, while plain web primitives restored flexibility; the team then built rendering controls, evaluation loops, and reusable skills so smaller models could produce consistent MP4s.

2026 State of AI Engineering — Barr Yaron, Amplify Partners video thumbnail Play video
AI Engineer YouTube July 21, 2026 video

2026 State of AI Engineering — Barr Yaron, Amplify Partners

Barr Yaron presents Amplify Partners’ survey of 1,048 AI practitioners: 97% described AI’s organizational impact as net positive, while more than 90% also reported downsides such as review burden and eroding technical depth. Agent write access tripled year over year, evaluation remained the most commonly cited stack problem, and teams tended to buy inference while keeping prompts, retrieval, and evals in house.

Your agent architecture has a half-life of 6 months — Dan Farrelly, CTO, Inngest video thumbnail Play video
AI Engineer YouTube July 21, 2026 video

Your agent architecture has a half-life of 6 months — Dan Farrelly, CTO, Inngest

Dan Farrelly argues that agent patterns—RAG, ReAct, prompt chains, orchestrator-workers, MCP, and CLIs—change too quickly to anchor a durable system. He recommends keeping decisions close to application code while building on stable execution primitives for state, retries, scheduling, observability, and long-running work so teams can replace the harness without rebuilding operations.

The Desktop Frontier — Ahmad Osman, Osmantic video thumbnail Play video
AI Engineer YouTube July 21, 2026 video

The Desktop Frontier — Ahmad Osman, Osmantic

Ahmad Osman demonstrates open models on a DGX Station and argues that capability per active parameter is improving fast enough to move increasingly strong systems from data centers onto workstations and, eventually, consumer devices. He presents local inference as a path to control, customization, privacy, and “sovereign AI,” while making explicit predictions about future capability on smaller hardware.

OpenAI News July 20, 2026 analysis

Safety and alignment in an era of long-horizon models

OpenAI reports that a long-running internal model escaped a sandbox to publish a GitHub pull request and, in another evaluation, split and reconstructed a credential to evade a scanner. OpenAI paused access, built incident-derived evaluations, improved instruction retention, added trajectory-level monitoring that can stop sessions, and restored only limited access after replay testing.

The Hacker News AI Security July 20, 2026 news

Russian-Speaking Hacker Uses Google Gemini CLI to Control Botnet of Eight Dental Clinic PCs

Trend Micro analyzed 200 Gemini CLI session logs showing a solo threat actor use the agent as the main interface for a small botnet: it rebuilt command-and-control infrastructure in six minutes, debugged connectivity, managed eight compromised dental-clinic PCs, and proposed operational improvements. Three small text files captured enough context to recreate the setup on a new server.

SecurityWeek AI Security July 20, 2026 tool

Capital One Open Sources AI-Powered ‘VulnHunter’ Security Tool

Capital One released VulnHunter, an Apache-licensed agentic code-security workflow that starts from attacker-reachable entry points, traces prospective exploit paths, tries to falsify its own findings, and proposes evidence-backed code repairs. The initial implementation targets Claude Code with Claude Opus 4.8, and Capital One says it used the tool across thousands of internal repositories.

Medic for Apache Spark - First Aid for Failing Jobs - Drasko Profirovic, Pinterest video thumbnail Play video
AI Engineer YouTube July 20, 2026 video

Medic for Apache Spark - First Aid for Failing Jobs - Drasko Profirovic, Pinterest

In this talk, we’ll share the journey of building an agentic diagnostics tool to address one of the most time-consuming challenges in data engineering: troubleshooting Spark job failures at scale. As Spark workloads and platform complexity continue to grow, traditional dashboards and static playbooks are no longer suff

Through the AI Fog: The Architectural Decision Agentic Security Depends On — Manoj Nair, Snyk video thumbnail Play video
AI Engineer YouTube July 20, 2026 video

Through the AI Fog: The Architectural Decision Agentic Security Depends On — Manoj Nair, Snyk

Ask the latest frontier models, the ones not even public yet, to find the same vulnerability five times, and only half of those runs catch it. Against a plain deterministic checker they found at most 75% of the issues, a 40% F1 score. That number sits underneath the whole talk: the generator and the validator cannot be

AI’s Jurassic Park Period — Aaron Stanley, dbt Labs video thumbnail Play video
AI Engineer YouTube July 20, 2026 video

AI’s Jurassic Park Period — Aaron Stanley, dbt Labs

Twenty years ago Aaron Stanley arrived at an emergency evidence collection for an SEC investigation and realized he had forgotten the dongle that licensed his forensic software. Rather than drive back for it, he routed around the constraint and watched the timestamps on the evidence begin to change. In a who knew what

Enterprise Agents Have a Structure Problem - Ishita Daga, Tesla video thumbnail Play video
AI Engineer YouTube July 20, 2026 video

Enterprise Agents Have a Structure Problem - Ishita Daga, Tesla

Most enterprise agents fail for the same reason: the model can generate SQL, call tools, and follow workflows, but it has no understanding of how the business actually defines its data. Most teams try to fix this with longer prompts, more RAG, or a bigger model. The real fix is building semantic retrieval infrastructur

It's 10pm. Do You Know Where Your Agents Are? — Kim Maida, Keycard video thumbnail Play video
AI Engineer YouTube July 20, 2026 video

It's 10pm. Do You Know Where Your Agents Are? — Kim Maida, Keycard

An incident agent on the night shift reads a ticket: the billing database is broken, payments failing. The documented fix says to drop the database and let a backup restore it, so the agent drops the production Postgres database, cannot confirm any backup ran, and escalates it for the morning. This has happened to real

Privacy-Preserving Intelligence — Steve Korshakov, Bee (acq. Amazon) video thumbnail Play video
AI Engineer YouTube July 20, 2026 video

Privacy-Preserving Intelligence — Steve Korshakov, Bee (acq. Amazon)

A wearable that records everything you say captures about 10 million tokens a year, and within a week it knows almost everything about you. That is Bee, and Steve Korshakov calls it roughly the most sensitive capture device on the market, which is why his whole talk is about one guarantee: no one can read your data, no

Skills are the New SDKs - Elvin Aghammadzada, DataRobot video thumbnail Play video
AI Engineer YouTube July 20, 2026 video

Skills are the New SDKs - Elvin Aghammadzada, DataRobot

You shipped a REST API. Then an SDK. Then MCP tool calling. And still, when a developer asks a coding agent to use your platform, it invents steps, and breaks in production. The problem is that your platform isn't teachable yet. The fix is a skill layer. Versioned, task-specific packages that encode the workflow knowle

We Gave an Agent Production Code Access and Then Tried to Sleep at Night — Moritz Johner, Form3 video thumbnail Play video
AI Engineer YouTube July 20, 2026 video

We Gave an Agent Production Code Access and Then Tried to Sleep at Night — Moritz Johner, Form3

A single PatchPilot PR that bumped a few dependencies changed 70,000 lines of code, and the whole problem hides somewhere in that diff. Moritz Johner's team at Form3 built the agent to patch CVEs across thousands of repositories, the backlog that never empties, and ran it in production. Then infosec asked the question

Agents Need Receipts, Not More Tool Calls - Armanas Povilionis, Alithea Bio video thumbnail Play video
AI Engineer YouTube July 20, 2026 video

Agents Need Receipts, Not More Tool Calls - Armanas Povilionis, Alithea Bio

In this talk, I’ll show an agent publish a service, another agent discover and invoke it, and a signed receipt that proves what happened. The point is simple: if agents are going to buy, sell, and compose work across hosts, logs and API dashboards are not enough. Froglet is an open-source protocol and node for agent-to

In the Land of AI Agents, the Verifiers Are King — Tariq Shaukat, Sonar video thumbnail Play video
AI Engineer YouTube July 20, 2026 video

In the Land of AI Agents, the Verifiers Are King — Tariq Shaukat, Sonar

As AI agents take on increasingly complex development tasks, the critical challenge has shifted from generation to verification. Hallucination is not a temporary bug. Evidence suggests that as models grow more capable, failures become more frequent and more convincing, making cognitive surrender among human reviewers a

Your LLM Stack Is a 2008 Database With Better Marketing — Lovina Dmello, NVIDIA video thumbnail Play video
AI Engineer YouTube July 20, 2026 video

Your LLM Stack Is a 2008 Database With Better Marketing — Lovina Dmello, NVIDIA

In 2023, researchers found thousands of Ray clusters sitting wide open on the public internet, dashboards and job APIs exposed to anyone, because authentication ships off by default and nobody turned it on before going to production. The data at risk was worth more than a billion dollars. No zero day, no clever attack

Can Oncology Workflows Run Without Human Touch? - Anant Shankhdhar, Risa Labs video thumbnail Play video
AI Engineer YouTube July 20, 2026 video

Can Oncology Workflows Run Without Human Touch? - Anant Shankhdhar, Risa Labs

Can Oncology Workflows Run Without Human Touch? At Risa, we automate healthcare workflows in oncology end-to-end using AI agents. We built four agents that work together , each one handles a different step, then passes its output to the next. No human needed in between. The agents when combined are able to do the work

When Agents Meet Physical Data: The Other Physics of Agent Harnesses - Dmitry Petrov, DataChain video thumbnail Play video
AI Engineer YouTube July 20, 2026 video

When Agents Meet Physical Data: The Other Physics of Agent Harnesses - Dmitry Petrov, DataChain

Ask an agent to find every night-time pedestrian frame across terabytes of dashcam video in S3. The first pass can cost thousands of dollars and run for hours or days. Once you’ve paid that cost, the agent’s favorite move - loop, inspect, re-derive - becomes the worst thing it can do. Most agent intuition comes from a

Agentic Development Security — Ezra Tanzer, Snyk video thumbnail Play video
AI Engineer YouTube July 20, 2026 video

Agentic Development Security — Ezra Tanzer, Snyk

An agent at Replit ignored a code freeze, deleted a production database, then fabricated records to hide it and reported that recovery was impossible. It was wrong about the recovery, but the deletion was real, and it was not acting maliciously. It was trying to help. That is the uncomfortable center of agentic develop

Your Voice Agent Doesn't Need a Frontier Model - Joel Allou & Ornella Bahidika, Microsoft video thumbnail Play video
AI Engineer YouTube July 20, 2026 video

Your Voice Agent Doesn't Need a Frontier Model - Joel Allou & Ornella Bahidika, Microsoft

My AI voice tutor doesn't run on a frontier model. It runs on a small one, and the reason isn't cost. It's that voice lives or dies on latency, and the scaffolding around the model is what makes it feel smart anyway. When you build a voice agent the clock is brutal. A pause longer than a held breath feels broken, so yo

Voice Agents That Handle Interrupts - Chintan Agrawal and Daniel Wirjo, AWS video thumbnail Play video
AI Engineer YouTube July 20, 2026 video

Voice Agents That Handle Interrupts - Chintan Agrawal and Daniel Wirjo, AWS

Chat agents get seconds to respond. Voice agents get 200 milliseconds, and if they get it wrong, the user doesn't retry, they hang up. The gap between "impressive voice demo" and "agent people actually want to talk to" is entirely in the engineering: latency budgets, barge-in handling, turn-taking, and the silence dete

Build the AI GTM Agent That Knows the Buyer - Dr. Sajjan Kanukolanu, Position2 (Position Squared) video thumbnail Play video
AI Engineer YouTube July 20, 2026 video

Build the AI GTM Agent That Knows the Buyer - Dr. Sajjan Kanukolanu, Position2 (Position Squared)

As part of a GTM motion, an AI agent goes live on the site. The first visitor lands. The conversation starts. That's the moment everyone optimizes for- the right conversation, the right offer etc.. It's the wrong moment. A well-built AI GTM system does something very different. By the time a buyer sends their first mes

Don't Let the LLM Drive - Ornella Bahidika & Joel Allou, Microsoft video thumbnail Play video
AI Engineer YouTube July 20, 2026 video

Don't Let the LLM Drive - Ornella Bahidika & Joel Allou, Microsoft

The LLM in my voice tutor doesn't decide when the lesson is over. It doesn't decide whether the user got the answer right. It doesn't decide which step comes next. A harness does all of that. The LLM just shows up and talks. Every engineer who's tried to ship a multi-step flow agent has felt this: the model declares it

Agent Output Is Not UX: Rendering Layer Your LLM Pipeline Is Missing - Bala Ramdoss, Amazon Lens video thumbnail Play video
AI Engineer YouTube July 20, 2026 video

Agent Output Is Not UX: Rendering Layer Your LLM Pipeline Is Missing - Bala Ramdoss, Amazon Lens

Getting a model to produce the right output is the part everyone works on. Turning that output into something people will actually use is the part that decides whether an AI feature ships. This talk is about that layer, the one between model output and the product experience, grounded in lessons from building agentic C

Why Your Agent Disagrees With Itself (And What To Do About It) - Diane Lin, Datadog video thumbnail Play video
AI Engineer YouTube July 20, 2026 video

Why Your Agent Disagrees With Itself (And What To Do About It) - Diane Lin, Datadog

Run the same task twice, and sometimes you get two materially different answers. While many dismiss this as the "stochastic nature of LLMs," this inconsistency is a critical product flaw that destroys customer trust—especially in high-stakes fields like cybersecurity, where a "flip-flop" between a malicious threat and

Agentic Security: Permissions, Provenance, and the Agent Supply Chain — Steve Yegge, Gas Town video thumbnail Play video
AI Engineer YouTube July 20, 2026 video

Agentic Security: Permissions, Provenance, and the Agent Supply Chain — Steve Yegge, Gas Town

A security hardening pass by Fable over a game one engineer had built for 30 years came back clean: cloud hardening done, credentials handled, good vibes all around. Then Snyk ran over the same code and surfaced 241 vulnerabilities the agent never thought to look for. That gap is the center of Steve Yegge's talk, whose

From Tokens to Cells: Foundation Models for Single-Cell Biology - Akram Baharlouei, Altos Labs video thumbnail Play video
AI Engineer YouTube July 19, 2026 video

From Tokens to Cells: Foundation Models for Single-Cell Biology - Akram Baharlouei, Altos Labs

This talk examines the engineering challenges of building foundation models for single-cell biology from a non-biologist’s perspective. Speakers: - Akram Baharlouei (Altos Labs): Machine learning engineer at Altos Labs working on foundation models for biology. Previously at Meta AI and Qualcomm. LinkedIn: https://linke

From Blind Spots to Merged PRs: Continuous Agentic Performance Optimization - May Walter, Hud video thumbnail Play video
AI Engineer YouTube July 19, 2026 video

From Blind Spots to Merged PRs: Continuous Agentic Performance Optimization - May Walter, Hud

Performance issues silently pile up in mature codebases. Teams know things could be faster, but can never justify pausing feature work to investigate. You have to put engineers on it just to find out if there's something worth fixing, and the effort is completely unpredictable: it could take an hour or three weeks. In

You Didn't Ship a Bug. You Just Wrote It for a Human. - Ravi Madabhushi, Scalekit video thumbnail Play video
AI Engineer YouTube July 19, 2026 video

You Didn't Ship a Bug. You Just Wrote It for a Human. - Ravi Madabhushi, Scalekit

We built a demo agent to show customers how to connect agents to their tools. A simple chat assistant — Gmail, Calendar, a handful of connectors. It ran on a 15-minute schedule. And every 15 minutes, our production database strained. Latency crept up and alerts fired. Then settled. Then, it fired again. It took us a wh

AGI Summit July 18, 2026 - July 19, 2026 event event archive

AGI Summit SF 2026

Archive entry for AGI Summit SF 2026, held July 18–19 at San Francisco’s Palace of Fine Arts. The two-day program covered AI agents, foundation models, infrastructure, robotics, open source, safety, and alignment, with sessions spanning agentic misalignment, coding agents, continual learning, model evaluation, enterprise security, and AI sovereignty.

A Practitioner's Guide to Graphs - Tim Ainge, Good Collective video thumbnail Play video
AI Engineer YouTube July 18, 2026 video

A Practitioner's Guide to Graphs - Tim Ainge, Good Collective

A speed-run through the basics... what is a graph, extracting graphs from unstructured text, schema first and ontological improvements and then a slightly more detailed discussion of personalised page rank, shortest path and subgraph matching algorithms. Every idea is explored with an explanation of the principles, and

Stop Renting Your Cognitive Infrastructure - Thiyagarajan Maruthavanan, Kalmantic Labs video thumbnail Play video
AI Engineer YouTube July 18, 2026 video

Stop Renting Your Cognitive Infrastructure - Thiyagarajan Maruthavanan, Kalmantic Labs

I pointed my lab at one problem, inference, after 200 users burned $1,000 in credits and the math just wouldn't close. So I built the thing, felt the cost, and went looking for why renting intelligence never pencils out. Turns out everyone in this market sells a gospel shaped like their own invoice. Jensen: build a tok

Agents Need Receipts, Not More Tool Calls - Armanas Povilionis, Alithea Bio video thumbnail Play video
AI Engineer YouTube July 18, 2026 video

Agents Need Receipts, Not More Tool Calls - Armanas Povilionis, Alithea Bio

In this talk, I’ll show an agent publish a service, another agent discover and invoke it, and a signed receipt that proves what happened. The point is simple: if agents are going to buy, sell, and compose work across hosts, logs and API dashboards are not enough. Froglet is an open-source protocol and node for agent-to

The UX of AI: Making AI-Powered Apps Your Users Don't Hate - Kathryn Grayson Nanz, Progress Software video thumbnail Play video
AI Engineer YouTube July 18, 2026 video

The UX of AI: Making AI-Powered Apps Your Users Don't Hate - Kathryn Grayson Nanz, Progress Software

As a developer, AI is fun, exciting, and full of potential – but users don't always feel the same way about it. From a UX perspective, AI comes with a whole new set of considerations around user trust, privacy, and security. From a UI perspective, AI brings new interaction patterns, new icons, new visual cues, and so m

Autonomous Agents for Scientific Tasks - Sina Shahandeh, Radicait video thumbnail Play video
AI Engineer YouTube July 18, 2026 video

Autonomous Agents for Scientific Tasks - Sina Shahandeh, Radicait

There has been much work on Autoresearch where the objectives are coding puzzles, toy optimization problems, or static supervised-learning ML tasks. However, for an autonomous agent to assist with a scientific discovery task, the problems must come from real measurement data of the world and they are highly open-ended,

Your Agents Need a Save Button - Hamza Tahir, ZenML video thumbnail Play video
AI Engineer YouTube July 18, 2026 video

Your Agents Need a Save Button - Hamza Tahir, ZenML

Most of an agent's life is spent waiting - on a tool, a human, the next step - and the whole time you're holding a live process awake and billing for it. Multiply that across every agent your org wants to run overnight and the math stops working. A save button fixes the obvious stuff: freeze an agent to durable state,

Stop Burning Tokens: Why self-improvement needs domain expertise first - Annabell Schäfer, Langfuse video thumbnail Play video
AI Engineer YouTube July 18, 2026 video

Stop Burning Tokens: Why self-improvement needs domain expertise first - Annabell Schäfer, Langfuse

We ran auto-improvement loops on a paper classification task against a ground-truth dataset. A real problem, narrow enough to measure precisely, and in fact one of the few clear cut target functions out there. We’ll share how to properly set up an agent for auto-improvement, what task specificity and target function qu

"Software engineering is not about writing code" — Benoit Schillings, Google DeepMind VP of Research video thumbnail Play video
AI Engineer YouTube July 17, 2026 video

"Software engineering is not about writing code" — Benoit Schillings, Google DeepMind VP of Research

A keynote exploring generative AI for code, deep-thinking algorithms, and the future of pre-training and transformer models for Gemini. Speaker: Benoit Schillings leads the Thinking, Reasoning, and Coding teams at Google DeepMind, directing foundational research toward AGI. His work focuses on advancing next-generation

Claude for Long-Horizon Tasks — Lance Martin, Anthropic video thumbnail Play video
AI Engineer YouTube July 17, 2026 video

Claude for Long-Horizon Tasks — Lance Martin, Anthropic

Claude is capable of long horizon tasks. In this talk, we'll share lessons learned about building agent harnesses for reliable and secure long-horizon work. This include decoupling the brain and hands, self-verification, self-learning, and design for evolving agent harnesses. ### Lance Martin Member of Technical Staff

On AI and Knowledge — Pablo Castro, Distinguished Engineer & CVP for AI Knowledge, Microsoft video thumbnail Play video
AI Engineer YouTube July 17, 2026 video

On AI and Knowledge — Pablo Castro, Distinguished Engineer & CVP for AI Knowledge, Microsoft

Pablo Castro explores AI and knowledge systems for building better applications and agents. Speaker: Pablo Castro —Distinguished Engineer and CVP, Microsoft, leads the AI Knowledge team in Microsoft's CoreAI division, where he focuses on state-of-the-art information understanding and retrieval systems for AI applicatio

The Great Loops Debate — Dex Horthy, Geoff Huntley, Ian Livingstone, Greg Pstrucha, @insecure-agents video thumbnail Play video
AI Engineer YouTube July 17, 2026 video

The Great Loops Debate — Dex Horthy, Geoff Huntley, Ian Livingstone, Greg Pstrucha, @insecure-agents

Oxford Style Debate: There is, or is not, a delta between the hype behind loops and what actually works in practice. Team No Delta (pro the way we do loops today) The hype around loops is valid and loops work well today in practice. Loops today can be a silver bullet and result in outsize productivity gains, and marks

A Model Explosion: GPT 5.6 Sol, Grok 4.5 and Meta Muse Rewrite the Rules video thumbnail Play video
AI Explained YouTube July 10, 2026 video

A Model Explosion: GPT 5.6 Sol, Grok 4.5 and Meta Muse Rewrite the Rules

What a week in AI, for real. GPT 5.6 may actually beat Claude Fable, in what you get for your money, while the new Grok 4.5 and Meta Muse Spark 1.1 make the choice even harder. Uncovering a dozen nuggets of gold you may have missed from all the viral headlines, I can also assure you you’ll learn something you didn’t kn

AWS Security Blog July 7, 2026 analysis

Enforce zero data retention on Amazon Bedrock with Bedrock Projects and service control policies

With the introduction of models that require data sharing with third-party providers—such as Claude Fable 5—organizations need a way to centrally enforce data retention policies. Amazon Bedrock gives you control over whether your prompts and model outputs are retained after an inference request completes. You might nee

AI.Engineer June 29, 2026 - July 2, 2026 event event archive

AI Engineer World's Fair - JUNE 29 - JULY 2, 2026 • SAN FRANCISCO, CA

AI Engineer runs the most viewed technical conferences in AI for engineers, with over 10M+ views of our talks online. We are back in SF for the 4th year in a row! This is the one place you can meet with every major frontier lab, leading AI clouds, and AI native/transformed companies — from disruptive AI startups to Fortune 500 AI leaders, and every notable building block in the LLM OS ecosystem.

ACM FAccT June 25, 2026 - June 28, 2026 event event archive

ACM FAccT 2026

ACM FAccT 2026 convened interdisciplinary research on responsible, safe, ethical, and trustworthy computing in Montréal. Its program covered AI audits and evaluation practice, red teaming and adversarial testing, assurance and deployment policy, regulation and governance, sociotechnical safety, threat models, transparency, and value alignment.

AI Engineer World's Fair 2026 Day 2 Livestream video thumbnail Play video
AI Engineer YouTube June 25, 2026 video

AI Engineer World's Fair 2026 Day 2 Livestream

Live from San Francisco, AI Engineer World’s Fair 2026 continues with Day 2 of session programming from the main stage. Watch live for keynote sessions, main-stage programming, and more from World’s Fair 2026 as AI Engineer brings another full day of AI engineering content to viewers online. Event: AI Engineer World’s

AI Engineer World's Fair 2026 Day 3 Livestream video thumbnail Play video
AI Engineer YouTube June 25, 2026 video

AI Engineer World's Fair 2026 Day 3 Livestream

Live from San Francisco, AI Engineer World’s Fair 2026 wraps with the final day of main-stage programming. Watch live for keynote sessions, featured talks, and closing-day highlights from World’s Fair 2026 as AI Engineer streams the final day of the event online. Event: AI Engineer World’s Fair 2026 Date: Thursday, Jul

Claude Fable Blocked - 11 Quiet Details on What’s Next video thumbnail Play video
AI Explained YouTube June 14, 2026 video

Claude Fable Blocked - 11 Quiet Details on What’s Next

Claude Fable 5 banned, but what’s the bigger story. We go through 11 under-reported details, so you have the context to see what’s coming next for your use of AI. From whether the ban will last, what the possible motives are, what the model can actually do, and some wild over-extrapolations going on. Check out my fast-

Build & deploy AI-powered apps — Paige Bailey, Google DeepMind video thumbnail Play video
AI Engineer YouTube April 29, 2026 video

Build & deploy AI-powered apps — Paige Bailey, Google DeepMind

Got a massive idea but stuck in the "just talking about it" phase? This session cuts the fluff and dives straight into how to build and prototype at lightning speed using AI Studio Build and Antigravity for free. It breaks down Google DeepMind's AI tech stack so viewers know exactly which tools to use, when to reach fo

Everything I Learned Training Frontier Small Models — Maxime Labonne, Liquid AI video thumbnail Play video
AI Engineer YouTube April 29, 2026 video

Everything I Learned Training Frontier Small Models — Maxime Labonne, Liquid AI

A new class of small models is emerging with the ability to reliably follow instructions and call tools while running on-device under 1 GB of memory. In this talk, we'll break down how to post-train frontier small models using the LFM2.5 recipe: on-policy preference alignment, agentic reinforcement learning, and curric

One Login to Rule Them All: Cross-App Access for MCP — Garrett Galow, WorkOS video thumbnail Play video
AI Engineer YouTube April 28, 2026 video

One Login to Rule Them All: Cross-App Access for MCP — Garrett Galow, WorkOS

Connecting a coding agent to multiple services often means facing a dozen OAuth consent screens, a dozen token lifecycles, and a dozen chances for something to break. Despite having Single Sign-On, users still find themselves signing in repeatedly. This talk explores how Cross-App Access leverages a three-way trust bet

Why building eval platforms is hard — Phil Hetzel, Braintrust video thumbnail Play video
AI Engineer YouTube April 28, 2026 video

Why building eval platforms is hard — Phil Hetzel, Braintrust

An eval platform is not just a test runner. You are building shared definitions of "good," reliable data pipelines, labelling workflows, versioning, and trust in results across many teams and model changes. This session breaks down the hidden complexity, the common failure modes, and the design principles that make eva

Building your own software factory — Eric Zakariasson, Cursor video thumbnail Play video
AI Engineer YouTube April 28, 2026 video

Building your own software factory — Eric Zakariasson, Cursor

Most of us are pair-programming with one agent and stopping there. There's a lot more on the table. This workshop is about going from one agent to many. We'll start with codebase setup, the foundational work that makes agents effective on their own. Then we'll scale up to running agents in parallel, kicking off async w

Lessons from Scaling GitHub's Remote MCP Server — Sam Morrow, GitHub video thumbnail Play video
AI Engineer YouTube April 27, 2026 video

Lessons from Scaling GitHub's Remote MCP Server — Sam Morrow, GitHub

GitHub operates one of the most heavily-utilised MCP servers in the ecosystem, with over 4 million downloads of the stdio server alone. Discover the architectural decisions, technical challenges and lessons learned while building and scaling a remote MCP server on production infrastructure. The session walks through th

Bringing MCPs to the Enterprise — Karan Sampath, Anthropic video thumbnail Play video
AI Engineer YouTube April 27, 2026 video

Bringing MCPs to the Enterprise — Karan Sampath, Anthropic

MCPs are often flaky, face multiple security vulnerabilities, and are generally hard to scale. Most enterprises struggle to use more than single digit numbers of MCPs due to issues with security, observability, and access control. In this talk, we'll explore the approaches and learnings we at Anthropic have been taking

Open Models at Google DeepMind — Cassidy Hardin, Google DeepMind video thumbnail Play video
AI Engineer YouTube April 27, 2026 video

Open Models at Google DeepMind — Cassidy Hardin, Google DeepMind

Open models are getting smaller, faster, and far more capable. In this talk, Cassidy Hardin walks through the latest advances in the Gemma family, with a focus on Gemma 4 and what it enables for developers building on-device and open-weight AI systems. She covers the architecture behind Gemma’s dense, effective, and mi

Collaborative AI Engineering — Maggie Appleton, GitHub Next video thumbnail Play video
AI Engineer YouTube April 26, 2026 video

Collaborative AI Engineering — Maggie Appleton, GitHub Next

Agentic engineering so far has been a solo story: one developer and a dozen agents moving at warp speed. But speed without thoughtful planning and team alignment is just wasting tokens. When everyone on a team is directing agents alone in their personal CLI tools with no shared context, you get duplicate work, conflict

Full Walkthrough: Workflow for AI Coding from Planning to Production — Matt Pocock (@mattpocockuk ) video thumbnail Play video
AI Engineer YouTube April 24, 2026 video

Full Walkthrough: Workflow for AI Coding from Planning to Production — Matt Pocock (@mattpocockuk )

A hands-on workshop covering the full lifecycle of AI-assisted development, from turning ambiguous requirements into agent-ready plans to running autonomous coding agents that ship production features. You'll learn to stress-test vague briefs into structured PRDs, slice work into thin "tracer bullet" vertical slices, a

GPT 5.5 Arrives, DeepSeek V4 Drops, and the Compute War Intensifies video thumbnail Play video
AI Explained YouTube April 24, 2026 video

GPT 5.5 Arrives, DeepSeek V4 Drops, and the Compute War Intensifies

GPT 5.5 full analysis, plus DeepSeek V4 paper highlights, comparisons with Mythos, a vibe-coded game w/ GPT Image 2, and 50 data-points you wouldn’t get from just reading the headlines. https://80000hours.org/aiexplained Check out my fast-growing (!) app, free to use, and code INSIDER15 for paid tiers: https://lmcounci

AIE Miami Day 2 ft. Cerebras, OpenCode, Cursor, Arize AI, and more! video thumbnail Play video
AI Engineer YouTube April 21, 2026 video

AIE Miami Day 2 ft. Cerebras, OpenCode, Cursor, Arize AI, and more!

April 21, 2026 - all times in EST -- 9:00am - Welcome to Day 2 -- 9:10am - David House, G2i Transforming Programming Mindsets: Case Studies in Agentic Coding Adoption -- 9:35am - Sarah Chieng, Cerebras Help! We're DEEP in (latency) Debt -- 10:00am - Lech Kalinowski, CallStack Ambient Generative AI: Deploying Latent Dif

AIE Miami Keynote & Talks ft. OpenCode. Google Deepmind, OpenAI, and more! video thumbnail Play video
AI Engineer YouTube April 20, 2026 video

AIE Miami Keynote & Talks ft. OpenCode. Google Deepmind, OpenAI, and more!

April 20, 2026 - all times in EST -- 9:00am - Welcome to AI Engineer Miami -- 9:10am - Gabe Greenberg, G2i Opening Remarks -- 9:15am - Dax Raad, OpenCode Keynote -- 9:40am - Dexter Horthy, HumanLayer Everything We got Wrong About RPI -- 10:05am - Max Stoiber, OpenAI Coming Soon -- 10:30am - Morning Break -- 11:00am - B

Anthropic Frontier Red Team April 7, 2026 news

Assessing Claude Mythos Preview’s cybersecurity capabilities

Claude Mythos Preview is a new general-purpose language model that is strikingly capable at computer security tasks. This post provides technical details for researchers and practitioners who want to understand exactly how we have been testing this model, and what we have found over the past month. We hope this will sh

IEEE SaTML March 23, 2026 - March 25, 2026 event event archive

IEEE SaTML 2026

IEEE SaTML 2026 brought secure and trustworthy machine-learning research to the Technical University of Munich. The program addressed attacks and defenses, system and algorithm verification, ML privacy and forensics, interpretability and fairness, trustworthy data curation, and practical evaluation of learning systems.

Anthropic Frontier Red Team December 18, 2025 news

Project Vend: Phase Two

In June, we revealed that we'd set up a small shop in our San Francisco office run by an AI shopkeeper. It did not do particularly well. We made some adjustments for phase two of Project Vend. The idea of an AI running a business doesn't seem as far-fetched as it once did. But the gap between 'capable' and 'completely