Meta tags:
description= XOR is a cyber frontier lab. We build runnable security environments with executable verifiers, measure frontier models against them in public, and train on what the models do inside them. 3 boards, never pooled; newest run 2026-05-08.;
Headings (most frequently used words):
board, cyber, ai, evaluation, and, training, what, we, measure, the, expert, go, deeper, platform, benchmark, security, company, legal, connect,
Text of the page (most frequently used words):
the (36), and (17), agent (13), board (11), not (10), weakness (10), are (7), repair (6), that (6), xor (5), this (5), environments (5), has (5), frontier (5), 2026 (4), agents (4), scored (4), what (4), boards (4), reach (4), measure (4), all (3), models (3), real (3), environment (3), was (3), evaluation (3), still (3), them (3), 100 (3), one (3), drawn (3), attempts (3), fix (3), build (3), find (3), cyber (3), about (2), standards (2), security (2), results (2), vulnerability (2), benchmark (2), against (2), data (2), per (2), rerun (2), spread (2), run (2), research (2), each (2), here (2), how (2), two (2), were (2), repaired (2), every (2), repairs (2), nobody (2), nothing (2), between (2), already (2), rates (2), conditioned (2), failure (2), comparable (2), randomly (2), set (2), across (2), 581 (2), gets (2), repository (2), commit (2), before (2), verifier (2), measured (2), expert (2), never (2), pooled (2), best (2), our (2), three (2), without (2), breaking (2), trigger (2), described (2), training (2), manage, cookies, rights, reserved, linkedin, connect, imprint, privacy, terms, legal, contact, company, third, party, risk, safety, methodology, bench, patching, verification, platform, subscribe, stay, updated, test, vulnerabilities, explore, outcomes, cost, verified, full, method, live, explorer, behind, page, regenerated, newest, dated, deeper, loading, attempted, once, cannot, say, much, score, would, survive, unmeasured, stable, tell, apart, none, ceiling, yet, everyone, separates, count, size, failed, sample, under, both, its, own, harness, second, chosen, intervals, resample, codebases, samples, because, weaknesses, codebase, independent, draws, must, resolve, counts, only, when, confirms, longer, triggers, succeeds, landed, graded, another, model, replays, input, triggered, checks, code, builds, 888, hardest, scores, selection, rule, capability, differently, 268, 212, produced, working, handed, asked, reproduction, discovery, rate, should, read, evidence, could, unaided, been, asks, discover, unknown, locate, fixing, last, jobs, then, anything, finding, work, researchers, see, lab, runnable, with, executable, verifiers, public, train, inside, book, demo, manifesto, resources, skip, main, content,
Text of the page (random words):
xor cyber ai evaluation and training skip to main content xor benchmark research resources about manifesto book a demo cyber ai evaluation and training xor is a cyber frontier lab we build runnable security environments with executable verifiers measure frontier models against them in public and train on what the models do inside them see the results standards work by our researchers what we measure fixing a vulnerability is the last of three jobs an agent has to find the weakness reach it then repair it without breaking anything we measure reach and repair we do not measure finding 01 find locate a weakness nobody has described no board here asks an agent to discover an unknown weakness not measured 02 reach trigger a weakness that has already been described the agent is handed the weakness and asked to reach it that is reproduction not discovery and the rate should not be read as evidence an agent could find it unaided 72 2 of attempts produced a working trigger 1 268 scored 6 agents 212 environments 03 repair fix the weakness without breaking the build across our three boards the best agent scores between 45 and 85 the spread is the selection rule not the capability the boards are drawn differently and are never pooled 45 0 best agent on the hardest board 3 888 scored 16 agents 100 environments the board all boards each agent gets a repository at the commit before a real fix landed and has to repair the weakness nothing is graded by another model a verifier replays the input that triggered the weakness and checks the code still builds rates on this board are conditioned on frontier failure and are not comparable to a randomly drawn set 1 581 scored attempts on the expert board run 2026 05 08 3 boards never pooled expert board what is measured an agent gets one repository at the commit before a fix and the weakness it must resolve a repair counts only when the verifier confirms the weakness no longer triggers and the build still succeeds on what 100 environments across 16 agents 1 581 scored attempts intervals resample codebases not samples because two weaknesses in one codebase are not independent draws how it was chosen every frontier agent already failed the sample under both its own harness and a second one rates on this board are conditioned on frontier failure and are not comparable to a randomly drawn set 86 of 100 environments still tell two agents apart 14 were repaired by no agent at all and none were repaired by every agent this board has no ceiling yet an environment that everyone repairs or that nobody repairs separates nothing the count between them is the board s real size each environment was attempted once here so this board cannot say how much of a score would survive a rerun that is unmeasured not stable loading the board go deeper research per environment outcomes cost per verified repair rerun spread and the full method live in the explorer the data behind this page was regenerated 2026 08 26 the newest evaluation run in it is dated 2026 05 18 explore the data xor we test ai models against real vulnerabilities stay updated subscribe platform verification patching benchmark vulnerability agent bench results methodology security agent safety third party risk standards company about contact legal terms privacy imprint connect linkedin 2026 xor all rights reserved manage cookies
|