FROGS benchmark: August 2026 results for generating an SVG of a frog with a Habsburg jaw

Overview of the FROGS benchmark

The benchmark asked models to draw a frog with a Habsburg jaw via SVG, collecting 42 runs in August 2026.

Scale and output statistics

14 models each had three attempts, yielding 42 SVGs; the largest file was google/gemini-3.6-flash at 14,329 bytes and the highest token consumption was deepseek/deepseek-v4-pro with 16,228 tokens in a single run.

Annotation and editorializing patterns

Most models added anatomical commentary; google/gemini-3.6-flash contained the most editorial notes (65), describing mood, royal framing, and anatomical exaggeration.

Royalty and mood additions beyond the prompt

Several models inserted royal imagery or mood descriptors not requested, such as crowns, regal eyelids, and references to inbreeding.

Deterministic and self‑aware outputs

mistralai/mistral-large-2512 produced two byte‑identical SVGs across its three runs; glm-5.1 and qwen3.7-max explicitly noted added crowns as "(because Habsburg)" or "(optional Habsburg reference)" in their SVG comments.

Community reaction on Hacker News

Commenters praised Opus 5, noted the lack of side‑view attempts, and discussed the benchmark’s relevance.

@jnwatson: Fable 5 on Max knocks it out of the park: https://imgur.com/a/usR8K7G Definitely has some creative flourishes. (I made no extra prompting. Just the above text. Single shot.)

@hn_throwaway_99: I thought this was great, and hilarious. Kudos to Opus 5, I thought it was the only one that came close to passing. Interestingly, I thought many of the failures drew the frog face OK, and they had some type of big blob for the jaw, so they knew "Hapsburg jaw" meant a protruding jaw, but it wasn't really connected to the frog face in any way that made sense. Small side note, the first gemini-2.5-pro one totally reminded me of some sad faced meme or Pepe the frog from somewhere. Anyone know what I'm referring to, tried to find it.

@krisoft: Curious that none of the attempts draw the frog from side profile. If i have to draw this i would immediately know that drawing a recognisable frog is the easy part of job. Expressing a particular jaw shape and melding it on the frog is the hard part. And jaw shapes are more prominent from the side. Even absence of thinking this through you would think that some frogs will be from the front, some from the side. Just by chance. And yet all appears to go for the harder pose.

@thebigship: Hi all, the site is getting hugged to death, thank you, was not expecting this kind of warm response. I will be working to make this more reliable, in the meantime, sign up for my newsletter: https://www.jaymollica.com/blog/ also my favorite SVG was def the google/gemini-3.6-flash edit: ok better now I think

@wren6991: Opus 5 clearly frogmaxxed. gemini-3.6-flash runs 2 and 3 responded best to the royal portrait context.

@getnormality: This is a strong benchmark! None of these could be remotely mistaken for human art. Opus 5 comes closest.

@fennecfoxy: Isn't that therefore a benchmark specifically on "it can generate an SVG of this exact thing" and naught else?

@rush86999: It's opus 5 > Kimi K3 > grok 4.5 That's a pretty good benchmark

@vb-8448: It's curious that all images are front facing ... although, for a habsburg jaw, a profile or 3/4 profile picture would be better.

@dehrmann: How do models approach SVG generation? In one version, I imagine them actually trying to reason about them as an LLM. In another, I imagine something closer to a GAN.

@ricardobeat: Gemini 2.5 Pro fails, but has a distinctive art style that is quite nice. It seems to understand shading to a much higher level than all other models.

@linksnapzz: A friend’s favorite prompt is “Batman & Julia Child; in the kitchen laughing at a ham”. Sounds simple, but has been surprisingly tough.

@leumon: Can you also try the new deepseek v4 flash?

@bigglywiggler: https://jumpshare.com/s/rcdD4i6MRY3AijzQG7W9 GPT 5.6 absolutely smashed it

@abound: My personal benchmark is a directory containing a bunch of research papers on the physics of popping popcorn kernels, and a prompt about creating high fidelity, photo realistic 3D models of all the different kinds of popped kernels. Fable (surprisingly? unsurprisingly?) refused to do it last time I tried, and the results from other models are, well, fine , but there's still plenty of headroom on this particular one.

@cmoski: Why do people benchmark these things on image generation? Surely that is not what most people here are using them for...

@evan_: Hopsburg Jaw

@riazrizvi: The secret to great interview questions and challenge tests is keeping them secret. Posting them on HN and getting them onto the front page puts them in jeopardy.

@ianberdin: Check out my MacBook SVG benchmark. From my experience, it demonstrates the Real model’s behavior. However, I notice the errors it makes, which are similar to the mistakes made by the mistake model in code. https://playcode.io/blog/macbook-svg-benchmark

@zirkuswurstikus: https://imgur.com/a/2DFUpGZ ChatGPT MMD

@getnormality: This is a strong benchmark! None of these could be remotely mistaken for human art. Opus 5 comes closest.

@fennecfoxy: Isn't that therefore a benchmark specifically on "it can generate an SVG of this exact thing" and naught else?

@rush86999: It's opus 5 > Kimi K3 > grok 4.5 That's a pretty good benchmark

@vb-8448: It's curious that all images are front facing ... although, for a habsburg jaw, a profile or 3/4 profile picture would be better.

@dehrmann: How do models approach SVG generation? In one version, I imagine them actually trying to reason about them as an LLM. In another, I imagine something closer to a GAN.

@ricardobeat: Gemini 2.5 Pro fails, but has a distinctive art style that is quite nice. It seems to understand shading to a much higher level than all other models.

@linksnapzz: A friend’s favorite prompt is “Batman & Julia Child; in the kitchen laughing at a ham”. Sounds simple, but has been surprisingly tough.

@leumon: Can you also try the new deepseek v4 flash?

@bigglywiggler: https://jumpshare.com/s/rcdD4i6MRY3AijzQG7W9 GPT 5.6 absolutely smashed it

@abound: My personal benchmark is a directory containing a bunch of research papers on the physics of popping popcorn kernels, and a prompt about creating high fidelity, photo realistic 3D models of all the different kinds of popped kernels. Fable (surprisingly? unsurprisingly?) refused to do it last time I tried, and the results from other models are, well, fine , but there's still plenty of headroom on this particular one.

@cmoski: Why do people benchmark these things on image generation? Surely that is not what most people here are using them for...

@evan_: Hopsburg Jaw

@riazrizvi: The secret to great interview questions and challenge tests is keeping them secret. Posting them on HN and getting them onto the front page puts them in jeopardy.

@ianberdin: Check out my MacBook SVG benchmark. From my experience, it demonstrates the Real model’s behavior. However, I notice the errors it makes, which are similar to the mistakes made by the mistake model in code. https://playcode.io/blog/macbook-svg-benchmark

@zirkuswurstikus: https://imgur.com/a/2DFUpGZ ChatGPT MMD

Sources