Our methods are peer-reviewed and presented at the venues that set the agenda for AI evaluation and safety.
EvoFlint formalizes the evolutionary, multi-turn red-teaming approach that powers Flint, treating model failure discovery as a diversity-driven search across conversation space rather than a fixed test suite. The result is a living atlas of how production language models break down over extended interactions, not just whether they pass.
Formalizes the evolutionary multi-turn red-teaming approach behind Flint.
A learned, feedback-driven masking policy for discrete diffusion language-model pretraining.
We don't just publish, we set the agenda, moderating and presenting alongside the leading labs.
Presenting red-teaming as a diverse search problem over model behavior.
A panel with heads of Trust & Safety and safety engineering from Anthropic, OpenAI, Google, and Mercor.
Exploring the tension between organizational automation and personalization in agentic systems.