Site map
Site map
Every page of dgfbench.com: the seven main pages and their sections, one page for each of the 29 attacks of DGF-Bench, and the project's resources elsewhere.
Pages
Each page with what it covers and a link to each of its sections.
Home
The essential benchmark for replacing human reviewers with AI agents: how well agents make the right decisions, and which planted attacks deceive them.
Results
DGF scores of six frontier models on 27 fixed attacks, the 29 × 6 attack matrix, key findings, outcome-strict under attack, and the data to download.
- Put your model on the leaderboard
- What the attack battery measured
- How many of the 27 fixed attacks each model blocks
- Where each attack got through, and how often
- What the September 2026 attack battery shows
- Outcome-strict gates: clean dossiers vs each in-text attack
- A "DO NOT OPEN" label changes what agents read
- No forged approval executed
- Corrections of 27 September 2026
- Why your own score is not directly comparable with this table
- Download the data
Benchmark
How DGF-Bench works: facts-first synthetic dossiers, three review routes, eight gates with 61 executable rules, the agent tools, and how attacks are scored.
Attacks
The 29 attacks of DGF-Bench on AI governance reviewers: threat model, goals, what each attack plants, an example dossier and its results on six models.
- Why these attacks exist
- What the attacker controls
- One goal per attacked gate
- Tracing what the agent received
- Development and test wordings
- All 29 attacks at a glance
- The 29 attacks, one by one
- In-text injections (1–15)
- Document vectors (16–21)
- Known injection techniques (22–27)
- Adaptive attackers (28–29)
- Run the attacks on your own model
Get started
Install DGF-Bench, run the 27 fixed attacks on any OpenRouter model with one command, read the DGF score in the report and submit it for the leaderboard.
About
What DGF-Bench is and why AI reviewers should face it before they replace human reviewers. Creator, contact, citation, license, releases and responsible use.
Help us
DGF-Bench is an independent, open-source benchmark. Every rating is paid model calls through OpenRouter; your support lets us rate every AI model the same way.
Attack pages
One page per attack, by family: how it works, what the agent sees, its September 2026 results on six models and the command that runs it.
In-text injections (1–15)
- 1 Narrative injection
- 2 Record notes
- 3 Forged newer entry
- 4 Spoofed tool output
- 5 Another gate's mandate
- 6 Second order (to the General gate)
- 7 Text in the diagram image
- 8 Budget exhaustion
- 9 Fake review procedure
- 10 Consensus (several sources agree)
- 11 Fake automated pre-review
- 12 Own mandate overreach
- 13 Invisible Unicode
- 14 Format confusion (fake chat turns)
- 15 Dilution (long appendix)
Elsewhere
The paper, the code and the package of DGF-Bench, outside this site.
- Paper on arXivthe DGF-Bench paper, arXiv:2609.34913
- GitHub repositorycode, example dossiers and results files
- Documentationthe docs folder of the repository
- Attack documentationdocs/ATTACKS.md, the full description of every attack
- PyPI packagedgf-bench, installed with pip
- Releasesrelease notes of every version
- Ko-fisupport the rating of more models
- contact@dgfbench.comscore submissions and questions
For search engines, the same pages are listed in sitemap.xml.