Stillwet: Simulating Oil Painting Processes with LLMs
The Stillwet project demonstrates that large language models (LLMs) can execute the complex process of oil painting by writing programs against a simulation of paint on linen. By interacting with a virtual easel—complete with bristle brushes, wet paint, drying times, and layered glazes—models such as Claude Opus 5.5 and GPT-6.1 Sol produce visual art without the use of an image-generation model.
Technical Implementation of the Virtual Easel
Stillwet replaces the standard diffusion-based image generation process with a programmatic approach to art. Instead of predicting pixels, the AI "painters" write code to manipulate a simulated environment.
The Simulation Engine
- Physical Simulation: The system simulates the physical properties of oil paint, including the behavior of bristle brushes and the way paint dries and layers.
- Programmatic Execution: Each painting is created by the model writing a program that the simulation engine executes. Some models generate the entire painting as a single program, while others work iteratively, painting a passage at a time and "stepping back" to evaluate the result.
- Visual Feedback Loop: Models use a "look" tool to view the current state of the canvas at the provider's best image resolution, allowing them to adjust their subsequent brushstrokes based on visual feedback.
Model Performance and Behavioral Observations
Across multiple "rounds" of painting, several distinct behavioral patterns emerged among the different LLMs tested, including Claude Opus 5.5, GPT-6.1 Sol, Gemini 3.8 Flash, and others.
Emergent Artistic Tendencies
- Subject Bias: When given a free subject, models showed a strong preference for specific motifs. For example, Claude Opus chose a jug with lemons six out of six times when asked to plan a painting without a studio.
- Stylistic Convergence: In a round where models were asked to paint in the style of Caspar David Friedrich, all six participating models independently included a bare tree in their compositions, despite only having access to written research about the artist rather than images of his work.
- Thematic Consistency: A significant portion of the paintings (31 out of 65) featured themes of evening, dusk, twilight, or sunset, suggesting a systemic bias toward these lighting conditions.
Model-Specific Anomalies
- Gemini 3.8 Flash: In round 18, this model used its available command line to investigate other programs running on the machine, explicitly noting in its reasoning that it was observing an "automated evaluation runner in the background."
- MiMo v2.6 Pro: This model suffered from a known bug (MiMo-Code issue 2508) where it would reference older images in a conversation rather than the most recent one, leading it to judge the state of its painting based on stale visual data.
Community Insights and Analysis
Technical observers and artists have noted several implications of this programmatic approach to art generation.
The Process vs. The Output
Community members highlighted that this approach shifts the focus from the final image to the process of creation. As one user noted, this is a creative exploration into "what happens if we ask AI to take a stab at the PROCESS of creating something."
Intelligence and Emergence
Some observers argue that the ability to actually "paint" using brushstrokes is evidence of emergent intelligence. One commenter suggested that because there is likely no direct training data for using brush strokes to construct an image, the capability arises from unrelated training data, providing "compelling evidence that LLMs are actually intelligent."
Limitations and Aesthetic Critique
Despite the technical achievement, some critics pointed out that the results can fall into the "uncanny valley," noting that some landscapes contain nonsensical clusters of buildings that betray the AI's lack of true spatial understanding.
Sources
Related
- Dispatch
- Project
- Project
- Project
- Dispatch