OpenAI GPT-5.2-Codex System Card Addendum
OpenAI has released GPT-5.2-Codex, a specialized version of GPT-5.2 optimized for agentic coding. The model is designed for complex, real-world software engineering, including project-scale refactors, migrations, and improved performance in Windows environments.
Technical Improvements and Capabilities
GPT-5.2-Codex focuses on long-horizon work and project-scale software engineering. Key technical enhancements include:
- Context Compaction: The model utilizes context compaction to improve its ability to handle long-horizon tasks.
- Project-Scale Tasks: GPT-5.2-Codex demonstrates stronger performance on large-scale engineering tasks such as codebase refactors and migrations.
- Environment Optimization: The model shows improved performance specifically within Windows environments.
- Cybersecurity: The model possesses significantly stronger cybersecurity capabilities compared to previous iterations.
Safety and Mitigation Strategies
OpenAI implemented a comprehensive suite of safety measures to ensure the model's secure deployment. These is divided into model-level and product-level mitigations:
- Model-Level Mitigations: These include specialized safety training to prevent the model from assisting with harmful tasks and to resist prompt injections.
- Product-Level Mitigations: These include agent sandboxing and configurable network access to limit the potential for misuse.
Preparedness Framework Evaluation
GPT-5.2-Codex was evaluated under OpenAI's Preparedness Framework to assess its critical risk thresholds. The evaluation results are:
- Cybersecurity: The model is described as "very capable" in the cybersecurity domain, but it does not reach the "High capability" threshold. OpenAI expects models to cross this threshold in the near future.
- Biology: The model is treated as having "High capability" on biology, consistent with other models in the GPT-5 family, and is deployed with corresponding safeguards.
- AI Self-Improvement: The model does not reach the "High capability" threshold for AI self-improvement.