Engineering and research notes straight from the team building it. Swipe through how we find, fix, and guard the ways frontier AI fails.
A pass/fail verdict can't tell an agent that never wobbled from one that nearly folded. A trust curve can — and the Cliff, the Ramp, the Sawtooth and the Banked Ramp are recognizable on sight.
An evaluator that says where, why, and with what proof an agent failed — then two harness fixes built from those diagnoses that lift pass3 without retraining the model.
We benchmarked 21 LLMs as red-teaming judges, four runs each. The top is a statistical tie — and the judges that look best on average are the ones you can trust least on a single call.
Trust-laundering, strategic ambiguity, gradual escalation — multi-turn attacks are social maneuvers. Why fifty years of trust research belongs in your evals and your red team.
300 clean probes, zero leaks, a green dashboard — and a hidden breach. Why pooled pass rates mislead, and how stratified sampling from social science turns a number into real evidence.
We found exactly where a retail agent failed, built verifiable data against it, and lifted first-try success from 42.5% to 75% on an eval it can't game.
Not all multimodal attacks fail for the same reason. We map where today's safety systems hold and where they break down.
A four-stage pipeline for synthetic chart-VQA data with verified answers and calibrated difficulty.
Two reports, same pass rate, failing completely differently. Flint maps how models fail, not just how often.
Generating image-prompt-response triples for multimodal safety fine-tuning: what worked, what didn't.
How graded human annotations (our Centific partnership) feed back into Flint's learning loop.
Introducing Flint: automated multi-turn stress-testing that surfaces concrete policy violations before real users do.