Resources

Our work, in the open.

Our work shows up at top-tier venues, and we help shape the AI safety community.

From the blog

Deep dives for practitioners.

Engineering and research notes from the team. Each card opens the full post on our blog.

Red Teaming

Social science knowledge as a special tool for your team

Multi-turn attacks are social maneuvers — trust-laundering, strategic ambiguity, gradual escalation. Why fifty years of trust research belongs in your evals and red team.

Read the post →
Evaluation

Your eval passed. That doesn't mean your system is safe.

300 clean probes, zero leaks, a green dashboard — and a hidden breach. Why pooled pass rates mislead, and how stratified sampling turns a number into evidence.

Read the post →
Agents

Closing the Loop: Finding and Fixing a Customer-Service Agent's Failure Modes

We found exactly where a retail agent failed, built verifiable data against it, and lifted first-try success from 42.5% to 75%.

Read the post →
Multimodal

Beyond Attack Detection: Why Multimodal Safety Is Fighting the Wrong Battle

Not all multimodal attacks fail for the same reason. We map where today's safety systems hold and where they break down.

Read the post →
Dataset

Stop Waiting on Labeled Data. Generate Your Evals Instead.

A four-stage pipeline for synthetic chart-VQA data with verified answers and calibrated difficulty.

Read the post →
Evaluation

Beyond Pass Rates: What AI Safety Tests Don't Tell You

Two reports, same pass rate, failing completely differently. Flint maps how models fail, not just how often.

Read the post →
Engineering

Building a Production-Grade Multimodal SFT Pipeline

Generating image-prompt-response triples for multimodal safety fine-tuning: what worked, what didn't.

Read the post →
Partnership

Human Judgment at Scale

How graded human annotations (our Centific partnership) feed back into Flint's learning loop.

Read the post →
Launch

Don't Ship That Chatbot (Until You Read This)

Introducing Flint: automated multi-turn stress-testing that surfaces concrete policy violations before real users do.

Read the post →
Publications and conferences

Published, not just claimed.

KDD 2026

EvoFlint: An Evolutionary Atlas of Multi-Turn LLM Vulnerabilities

Formalizes the evolutionary multi-turn red-teaming approach behind Flint.

ICML 2026

InfoDLM: Information-Adaptive Discrete Diffusion LM Pretraining

Prof. Tony Geng, Rice University: a learned, feedback-driven masking policy.

Community leadership

Convening the AI field.

KDD 2026

AI Integrity as a Search Problem: Diversity-Driven Behavioral Evaluation

CEO Anish Das Sarma presenting on red-teaming as a diverse search problem.

TrustCon 2026

"Who Owns What? Responsible AI in Dynamic Production Systems"

Our panel with heads of Trust & Safety and safety engineering from Anthropic, OpenAI, Google, and Mercor. Anish moderating.

NeurIPS 2025

"Agentic AI: Organizational Automation vs. Personalization at Scale"

Co-hosted with Centific.