OpenAI gpt-image-1 API Release
OpenAI has released gpt-image-1, a natively multimodal image generation model now available via API. This release allows developers and businesses to integrate high-quality, professional-grade image generation, text rendering, and visual editing capabilities directly into their own applications.
Model Capabilities and Versatility
gpt-image-1 is the same model that powers image generation within ChatGPT. It is designed to be versatile across diverse aesthetic styles and can faithfully follow custom guidelines while leveraging world knowledge. Key technical capabilities include:
Accurate Text Rendering: The model can accurately render text within generated images.
Visual Editing: The model supports high-fidelity visual edits and the ability to transform rough sketches into polished graphic elements.
Diverse Styling: It provides the flexibility to experiment with different aesthetic styles to suit various professional and consumer needs.
Enterprise Integration and Use Cases
Several leading companies are already integrating gpt-image-1 into their creative and business tools:
- Adobe: Integrating image generation capabilities into Firefly and Express apps to provide creators with more aesthetic style options.
- Canva: Exploring the use of
gpt-image-1for design generation and editing within Canva AI and Magic Studio, specifically for transforming sketches into graphic elements. - GoDaddy: Experimenting with the creation of editable logos, background removal, and the generation of professional typography for brand content and social media assets via GoDaddy Airo®.
- HubSpot: Exploring the creation of marketing and sales collateral for social media, email marketing, and landing pages.
- Instacart: Testing the generation of images for recipes and shopping lists.
- invideo: Utilizing the model for improved text generation, fine-grain editing controls, and advanced style guidance in AI video production.
Safety and Moderation
gpt-image-1 employs the same safety guardrails as the image generation features in GPT-4o, including restrictions on generating harmful images. All generated images include C2PA metadata for provenance.
Developers using the API have control over moderation sensitivity via the moderation parameter:
- auto (default): Standard filtering.
- low: Less restrictive filtering.
OpenAI states that customer API data is not used for training by default, and all inputs and outputs remain subject to API usage policies.
Pricing Structure
gpt-image-1 is priced per token, with distinct rates for text and image tokens:
| Token Type | Price per 1M Tokens |
|---|---|
| Text input tokens (prompt text) | $5 |
| Image input tokens (input images) | $10 |
| laage output tokens (generated images) | $40 |
In practical terms, this results in approximately $0.02, $0.07, and $0.19 per generated image for low, medium, and high-quality square images, respectively.