Our work shows up at top-tier venues, and we help shape the AI safety community.
Engineering and research notes from the team. Each card opens the full post on our blog.
Multi-turn attacks are social maneuvers — trust-laundering, strategic ambiguity, gradual escalation. Why fifty years of trust research belongs in your evals and red team.
Read the post →300 clean probes, zero leaks, a green dashboard — and a hidden breach. Why pooled pass rates mislead, and how stratified sampling turns a number into evidence.
Read the post →We found exactly where a retail agent failed, built verifiable data against it, and lifted first-try success from 42.5% to 75%.
Read the post →Not all multimodal attacks fail for the same reason. We map where today's safety systems hold and where they break down.
Read the post →A four-stage pipeline for synthetic chart-VQA data with verified answers and calibrated difficulty.
Read the post →Two reports, same pass rate, failing completely differently. Flint maps how models fail, not just how often.
Read the post →Generating image-prompt-response triples for multimodal safety fine-tuning: what worked, what didn't.
Read the post →How graded human annotations (our Centific partnership) feed back into Flint's learning loop.
Read the post →Introducing Flint: automated multi-turn stress-testing that surfaces concrete policy violations before real users do.
Read the post →Formalizes the evolutionary multi-turn red-teaming approach behind Flint.
Prof. Tony Geng, Rice University: a learned, feedback-driven masking policy.
CEO Anish Das Sarma presenting on red-teaming as a diverse search problem.
Our panel with heads of Trust & Safety and safety engineering from Anthropic, OpenAI, Google, and Mercor. Anish moderating.
Co-hosted with Centific.