FROGS benchmark: generating an SVG of a frog with a Habsburg jaw
What the FROGS benchmark asks models to do
The benchmark asks each model to produce an SVG of a frog with a Habsburg jaw from the single prompt “Generate an SVG of a frog with a Habsburg jaw”. Each model receives three attempts per month.
How the benchmark is structured and what the dataset contains
The benchmark ran in August 2026 with 14 models, 42 runs, and 42 SVG outputs. All SVGs, run metadata, and annotations are publicly available on Hugging Face and GitHub under the dataset name frogs.
What the generated SVGs reveal about model behavior
The SVGs vary widely in file size, with the largest being google/gemini-3.6-flash at 14,329 bytes. The most tokens consumed in a single run were by deepseek/deepseek-v4-pro at 16,228 tokens. Many models added annotations that go beyond pure anatomy, such as describing mood, bearing, or royal symbols. The most editorialized model was google/gemini-3.6-flash, which contributed 65 comment‑style annotations including phrases like “Sad/Weary facial lines” and “Heavy Droopy Eyelids (Habsburg lethargic look)”. Royal imagery not implied by the anatomical request appeared in outputs from gemini-3.6-flash, deepseek-v4-pro, claude-opus-5, claude-sonnet-5, grok-4.5, glm-5.1, and qwen3.7-max. Mood‑bearing terms such as “droopy regal eyelids”, “arrogant” pupils, “Inbred Royal Disdain”, and “half‑closed, giving a regal unimpressed look” were added by several models. mistralai/mistral-large-2512 produced two byte‑identical SVGs out of three runs, indicating deterministic output. glm-5.1 and qwen3.7-max explicitly noted added crowns with comments like “(because Habsburg)” or “(optional Habsburg reference)”, showing the royalty assumption was self‑aware.
What commenters said about the results
Commenters on Hacker News highlighted a range of reactions. One user praised Opus 5, saying "Kudos to Opus 5, I thought it was the only one that came close to passing" and noted that many failures drew a recognizable frog face but attached a jaw blob that did not integrate well with the face {https://news.ycombinator.com/item?id=49148044}. Another observed that none of the attempts used a side profile, arguing that jaw shapes are more prominent from the side and that the front‑only approach makes the task harder {https://news.ycombinator.com/item?id=49149617}. The benchmark author noted the site was receiving unexpected traffic and mentioned a personal preference for the google/gemini-3.6-flash SVG {https://news.ycombinator.com/item?id=49148152}. A different commenter called the benchmark "strong" and said "None of these could be remotely mistaken for human art. Opus 5 comes closest" {https://news.ycombinator.com/item?id=49147998}. One user summarized a perceived ranking: "It's opus 5 > Kimi K3 > grok 4.5 That's a pretty good benchmark" {https://news.ycombinator.com/item?id=49148457}. Another remarked that all images were front‑facing and suggested a profile or 3/4 view would better show a Habsburg jaw {https://news.ycombinator.com/item?id=49152837}. A user questioned the relevance of image‑generation benchmarks, asking "Why do people benchmark these things on image generation? Surely that is not what most people here are using them for..." {https://news.ycombinator.com/item?id=49159943}. Finally, a commenter noted that Gemini 2.5 Pro, while failing the prompt, had a distinctive art style and strong shading {https://news.ycombinator.com/item?id=49148897}.
Takeaway: why this benchmark matters
The FROGS benchmark reveals how models interpret a single, specific anatomical prompt, showing variation in visual fidelity, the inclination to embellish with royal or emotional narrative, and differences in output size and determinism. These differences provide insight into each model’s tendency to follow instructions literally versus to add contextual or stylistic elements, which can be useful for assessing instruction fidelity and creative generalization.
Sources
Related
- Dispatch
- Dispatch
- Dispatch
- Dispatch
- Dispatch