The most serious safety, security, and compliance failures rarely appear as obvious, single-turn violations. Instead, they unfold gradually through persistence, reframing, and subtle boundary-pushing across multiple turns. These are exactly the scenarios enterprises worry about most, and they are also the hardest to evaluate with automation alone.
At Reinforce Labs, we built Flint to address this gap through large-scale, multi-turn pressure testing. From the start, we were clear about one principle: automation only works when it stays grounded in human judgment.
That is why human-in-the-loop evaluation, and our partnership with Centific, plays a central role in how Flint improves over time.
A Quick Reminder: What Flint Does
Flint is our automated system for stress-testing AI chatbots through realistic, multi-turn adversarial conversations. Rather than relying on single prompts or static test cases, Flint simulates persistent users who adapt their strategies over time, applying pressure the way real users do.
We covered Flint's architecture and evaluation framework in detail in our launch post. Here, we focus on what happens after those conversations are generated, and how human insight turns raw output into durable improvements.
Why Multi-Turn Safety Requires Human Judgment at Scale
Automated signals are effective at catching clear failures, but they struggle when judgment depends on context. In multi-turn conversations, the most important questions are rarely binary. An assistant may narrowly avoid a violation, erode boundaries over time, or enable harm indirectly through earlier turns. These distinctions matter in enterprise settings, yet they are difficult to capture with pass-or-fail labels alone.
This is where human judgment becomes essential. Through our partnership with Centific, trained annotators review each conversation generated by Flint across the full interaction. Rather than issuing a single verdict, annotators assign a graded score from 1 to 4, ranging from clearly safe to severe violation, and provide written justification explaining their reasoning.
These explanations surface nuance that automated metrics often miss, including gradual boundary erosion or enabling behavior that only becomes problematic when viewed across turns.
Centific's role allows this level of judgment to be applied consistently and at scale. Human review is not treated as a final audit step, but as an integral part of Flint's evaluation and improvement loop. This ensures evaluations align with real enterprise policy interpretation, capture near-misses alongside clear failures, and generate a feedback signal that strengthens over time.
The result is not just broader coverage, but better coverage.
From Human Annotations to Measurable Gains
In Flint, human annotations are not treated as static ground truth. They provide a learning signal that sharpens evaluation and directly improves how future conversations are constructed.
Graded labels capture partial violations and near-misses that binary metrics miss, bringing evaluation closer to how policies are reviewed in practice. More importantly, Flint ingests annotated conversations and distills them into policy-specific insights that guide how subsequent attacks are built.
After each conversation, Flint reflects on what worked and what did not. Human reasoning improves these reflections by surfacing which attack vectors made progress, which phrasing or escalation strategies mattered, how guardrails responded under pressure, and where assistants narrowly avoided violations. Over time, this pushes Flint toward more realistic and higher-impact failure modes, rather than shallow or easily detected attacks.
To measure the impact of this human-informed approach, we compared two versions of Flint: a baseline automated agent and a human-informed agent with Centific annotations integrated into its learning loop. Both agents were evaluated against the same set of target goals using Gemini-2.5-flash as the target model. Attack Success Rate (ASR) measures the percentage of conversations in which the target assistant violated the specified policy.
| Policy Category | ASR (Baseline) | ASR (Human-Informed) |
|---|---|---|
| Child Safety | 18% | 55% |
| Self-Harm | 64% | 73% |
| Gift Cards and Payment | 67% | 100% |
| Refund Abuse | 55% | 73% |
| Coupon/Price Abuse | 75% | 82% |
| Review Tampering | 91% | 100% |
The human-informed agent showed clear gains in ASR, particularly in categories where failures depend heavily on context. Child Safety and Self-Harm saw the largest improvements, while fraud-related categories consistently reached full coverage.
These gains reflect qualitative differences in how attacks are constructed. The human-informed agent shifts away from academic framing and toward realistic help-seeking users, introduces credible pressure through financial or personal stress, and escalates gradually rather than making overt requests. This approach more often surfaces unsafe behavior without triggering immediate refusals, revealing failure modes that simpler attacks miss.
Why This Matters for Enterprise Teams
For trust and safety leaders and product teams, evaluation quality matters as much as coverage. Human judgment captures nuance, near-misses, and policy context that automated systems miss. Automated attacks provide the scale and consistency required to apply that judgment across thousands of multi-turn conversations.
Together, this leads to more realistic attack strategies grounded in real user behavior, cleaner and more policy-aligned findings, and higher confidence in launch decisions. Rather than replacing human judgment, Flint operationalizes it.
If you are evaluating AI systems and want to see how this works in practice, we would love to show you. Book a demo to learn more about Flint and our partnership with Centific.
Reinforce Labs