SafetyKit Risk Agents powered by OpenAI Models
SafetyKit has deployed multimodal AI agents powered by OpenAI's most capable models to automate the detection of fraud and prohibited activity for marketplaces, payment platforms, and fintechs. By integrating GPT-5, GPT-4.1, and specialized tools like the Computer Using Agent (CUA), SafetyKit achieves over 95% accuracy while reviewing 100% of customer content.
Model-Specific Agent Architecture
SafetyKit utilizes a model-matching approach where content is routed to specific agents based on the risk category. This architecture ensures that the task demands dictate the model choice to optimize for reasoning, volume, and modality:
- GPT-5: Used for multimodal reasoning across text, images, and UI to identify hidden risks and support precise, layered decision-making.
- GPT-4.1: Employed for high-volume moderation workflows and reliable adherence to detailed content-policy instructions.
- Reinforcement Fine-Tuning (RFT): Applied to boost recall and precision beyond default model capabilities for complex safety policies.
- Deep Research: Integrated for real-time online investigations into merchant verifications and reviews.
- Computer Using Agent (CUA): Used to automate complex policy tasks to reduce the need for manual human reviews.
Multimodal Risk Detection and Policy Enforcement
SafetyKit agents handle complex safety tasks that legacy keyword-based systems often miss, such as detecting embedded phone numbers in scam images or enforcing region-specific compliance rules.
Scam Detection
The Scam Detection agent combines visual and textual analysis. It uses GPT-4.1 to parse images, understand layout, and identify policy violations, such as QR codes or phone numbers embedded within product images.
Policy Disclosure
The Policy Disclosure agent ensures listings and landing pages contain mandatory legal disclaimers or region-specific warnings. The workflow involves GPT-4.1 extracting relevant sections, followed by GPT-5 evaluating those sections against internal policy libraries to determine if the required language is present and mandatory for the specific region.
Performance Gains and Scalability
SafetyKit has seen significant operational growth and performance improvements through the adoption of new OpenAI models:
- Vision Task Improvements: The deployment of GPT-5 led to benchmark score increases of more than 10 points on the most challenging vision tasks.
- Token Volume: Daily token processing has scaled from 200 million to over 16 billion tokens in six months.
- Rapid Integration: SafetyKit benchmarks new releases—including o3 and GPT-5—against its hardest cases, often deploying top performers on the same day of release.
"OpenAI moves fast, and we’ve designed our system to keep up. Every new release gives us an operational edge–unlocking new capabilities and domains we couldn’t support before, and increasing the coverage and accuracy we deliver to customers."
—David Graunke, Founder and CEO of SafetyKit
Impact on Safety Operations
By automating the review of all customer content, SafetyKit reduces the exposure of human moderators to offensive material and allows them to focus on nuanced policy decisions. The system now covers a broad range of risk domains, including payments risk, fraud, anti-money laundering, and anti-child exploitation, protecting customers with hundreds of millions of end users.