Blog

See the work, up close.

Engineering and research notes straight from the team building it. Swipe through how we find, fix, and guard the ways frontier AI fails.

Part 3Red Teaming

Attacks have shapes

A pass/fail verdict can't tell an agent that never wobbled from one that nearly folded. A trust curve can — and the Cliff, the Ramp, the Sawtooth and the Banked Ramp are recognizable on sight.

6 min readRead article →
Part 2Agents

Closing the Loop, Part 2: Evaluating Agent Failures and Improving the Harness

An evaluator that says where, why, and with what proof an agent failed — then two harness fixes built from those diagnoses that lift pass3 without retraining the model.

6 min readRead article →
Evaluation

Beyond Average Accuracy: Choosing an LLM Judge You Can Actually Trust

We benchmarked 21 LLMs as red-teaming judges, four runs each. The top is a statistical tie — and the judges that look best on average are the ones you can trust least on a single call.

15 min readRead article →
Part 2Red Teaming

Social science knowledge as a special tool for your team

Trust-laundering, strategic ambiguity, gradual escalation — multi-turn attacks are social maneuvers. Why fifty years of trust research belongs in your evals and your red team.

5 min readRead article →
Part 1Evaluation

Your eval passed. That doesn't mean your system is safe.

300 clean probes, zero leaks, a green dashboard — and a hidden breach. Why pooled pass rates mislead, and how stratified sampling from social science turns a number into real evidence.

7 min readRead article →
Agents

Closing the Loop: Finding and Fixing a Customer-Service Agent's Failure Modes

We found exactly where a retail agent failed, built verifiable data against it, and lifted first-try success from 42.5% to 75% on an eval it can't game.

8 min readRead article →
Multimodal

Beyond Attack Detection: Why Multimodal Safety Is Fighting the Wrong Battle

Not all multimodal attacks fail for the same reason. We map where today's safety systems hold and where they break down.

7 min readRead article →
Dataset

Stop Waiting on Labeled Data. Generate Your Evals Instead.

A four-stage pipeline for synthetic chart-VQA data with verified answers and calibrated difficulty.

8 min readRead article →
Evaluation

Beyond Pass Rates: What AI Safety Tests Don't Tell You

Two reports, same pass rate, failing completely differently. Flint maps how models fail, not just how often.

6 min readRead article →
Engineering

Building a Production-Grade Multimodal SFT Pipeline

Generating image-prompt-response triples for multimodal safety fine-tuning: what worked, what didn't.

9 min readRead article →
Partnership

Human Judgment at Scale

How graded human annotations (our Centific partnership) feed back into Flint's learning loop.

5 min readRead article →
Launch

Don't Ship That Chatbot (Until You Read This)

Introducing Flint: automated multi-turn stress-testing that surfaces concrete policy violations before real users do.

5 min readRead article →