Google has been steadily dropping newer Flash models over the past few months, but hardcore users have only been waiting for one thing: a new frontier-class model. The recent Gemini 3.8 Flash release brought some interesting coding improvements, but it would be a disservice to call it a breakthrough model.
Gemini models have long had a reputation for producing awkward, unoriginal web designs when asked to create frontend code. That might finally be on its way out, if early test results of what appears to be Gemini 4 Pro are anything to go by.
The leak from Lentils was among the first to mention it a few days ago, saying Google had begun testing early checkpoints of Gemini 4 Pro within the company under the code name "argon."
Gemini 4 Pro checkpoints have finally started appearing internally a few days ago.
— Lentils (@Lentils80) September 14, 2026
This is the first ever output from the model, internally codenamed "argon". It took 2.4 minutes on High thinking effort.
It has a 256k token "output limit", compared to 64k in previous Gemini… pic.twitter.com/x76ihjstTG
Since then, others have also shared their early Gemini 4 Pro test results on X. A demonstration posted by developer Bee on X featuring a fully monographic website that looked extremely clean was pretty impressive.
Although the model took fourteen minutes to produce the page, the final result had nothing whatever in common with the robotic templates which we typically receive from AI tools. Bee said that he was surprised at how clean and well-polished the monographic website had become "it looks so good", and mentioned that Google had done a great job with the design, the earlier problems regarding design taste now appearing to have been resolved.
Gemini 4 Pro is unreal...
— Bee (@thtbee_) September 17, 2026
I'm surprised by how clean and polished it made this monographic website... this looks so good!
It took 14 minutes to build this.. seems like the design taste issues earlier Gemini models had is finally solved!
Google really cooked! https://t.co/M0g1UCHeRv pic.twitter.com/RSIWPLJsNO
The audio effects used to create the sound of sketching on canvas paper are also a neat touch that adds to the interface's tactile feel.
Another AI benchmarking account also shared Gemini 4 Pro's output in an attempt to create an Xbox controller SVG. We've seen similar tests from Gemini models in the past, but none have had the same level of polish as this one. So it's almost certain this is Gemini 4 Pro's work of art.
🚨 Gemini 4 Pro (Xbox controller SVG)
— Lumina (@LuminaBench) September 19, 2026
This was from "gemini 3.7 flash" in arena btw which was routing to 4 pro, it took around 20 minutes and looks crazy
This example was from last night, today I can’t seem to get anything close to this https://t.co/pohi5FsOPo pic.twitter.com/I2ZG789L1N
Way back in July, Logan Kilpatrick hinted that the team is cooking with "Gemini 3.5 Pro," but it looks like Google may not have been too excited about the results at the time. With all the progress made since then, it only makes sense that the team would skip the 3.5 moniker and go straight to 4 Pro. Plus, with these early results, it might be worth the big jump instead of a point release.
That said, while Gemini 4.0 Pro is still in the oven, Google is also working on enhancements for the Gemini desktop app, which we spotted in testing. In short, Projects will work a lot like the workspaces in ChatGPT or Claude. You’ll have dedicated hubs for reference files and custom instructions, so each new conversation uses that shared context instead of making you start over every time.
These tests also come at an interesting time, as we're seeing some of the first signs of what could be Anthropic's upcoming Fable 5.2 model. So it'll be interesting to see what Google has been working on for the past few months and how it compares to Anthropic's next frontier model.
The results from these current tests suggest that Gemini 4 Pro could possibly have a real chance of becoming a worthy frontier-class model from Google. But we'll have to wait and see how it fares in head-to-head comparisons with Astra, Grok 4.7, and, of course, Fable 5.2 to see how it stacks up against the competition.