1Password's AI patching benchmark is misleading

Executive Summary

1Password’s August 2026 report claimed AI models produced clean fixes only 26% of the time, but the figure is misleading. The benchmark included trials where agents were told to apply wrong fixes, experiments that prohibited testing, and a sample of six difficult bugs. Reanalysis of trials where agents could run code and were not misdirected shows 86% of patches blocked the exploit. The headline may discourage defenders from using AI patching, while Trail of Bits released two agent skills—post‑patch‑validation and review‑walkthrough—to improve patch quality.


Intelligence Metadata - Source Publisher: Trail of Bits - Published Date: 2026-09-15T11:00:00+00:00 - Category: research

Original Description: 1Password’s FLAWED report, published on August 6, 2026, gives defenders a misleading picture of AI patching. Its headline says models produced clean fixes only 26% of the time. That figure includes experiments that deliberately instructed agents to apply the wrong fix, along with experiments in which agents could not compile or test their patches. The report risks making defenders less effective by discouraging them from using technology that could help them fix more vulnerabilities. Teams th...

"I am a man of fixed and unbending principles, the first of which is to be flexible at all times."

— Everett Dirksen
Source: Trail of Bits