I gave Claude, Codex, and Antigravity nothing however screenshots of an app, and one rebuilt it virtually pixel for pixel


Vibe-coding has modified quite a bit about what it takes to build software. Not too way back, somebody with little to no technical background would have needed to sit by way of hours of coding tutorials, be taught the fundamentals of a programming language, and spend even longer determining why their code refused to work.

Now, you possibly can describe what you need in plain English, level an AI coding agent in the correct route, and find yourself with one thing practical with out understanding each line beneath it. But we have moved effectively past LLMs merely spitting out traces of code behind the scenes. Coding fashions have gotten significantly better at the visual side of development too.

Give them a screenshot of an interface, and in principle, they need to be capable of work backwards from the completed product like determining the format, spacing, colours, typography, and parts wanted to recreate it. That made me marvel how far this has really gone. So I gave Claude Code, Codex, and Google Antigravity the very same screenshots of an app, gave them no supply code or design recordsdata, and requested each to rebuild it from scratch.

I intentionally selected an app the fashions had been unlikely to already know

Familiarity would have ruined the take a look at

I initially thought I’d go together with an app that is a bit extra complicated but additionally well-known. For occasion, a social media app like Instagram or Snapchat, or one thing visually distinctive like Google Maps, the place recreating the interface convincingly would really be a problem.

The difficulty is that these are all apps the fashions are already more likely to be very acquainted with. Even if I solely gave them screenshots, there would at all times be questions of whether or not they had been reverse-engineering the UI from what they might see or just falling again on what they already knew the app was speculated to seem like.

I assumed that defeated the complete function of the experiment, so I deliberately went with a distinct segment app that most individuals in all probability have not even heard of: Foqos. It’s an open-source focus app built around helping you block distracting apps and keep off your cellphone, however extra importantly, for this take a look at, it has a clear, distinctive interface with out being so complicated that the comparability turns right into a take a look at of backend engineering as a substitute.

It’s additionally not well-known within the conventional sense, and whereas an LLM might technically observe down its supply code as a result of the challenge is open supply, that wasn’t one thing I allowed throughout the take a look at. This is the immediate I used for every software:

I used the CLI model of all three instruments and set each up in its personal separate listing with the very same set of reference photographs. I used Opus 5 for Claude Code, GPT-6 Sol for Codex, and Gemini 3.1 Pro for Antigravity, with larger reasoning enabled on all.

Codex was the perfect all-rounder

It made one hilarious mistake, although

Some time in the past, I wrote an article on XDA about how Codex nails virtually every little thing besides frontend work. Since then, although, OpenAI has moved on to its newer GPT-6 household, together with Astra, Sol, and Luna, and the distinction in frontend work is fairly noticeable. For this take a look at, I used GPT-6 Sol, which is constructed particularly for extra complicated coding and agentic workflows.

Out of all three, Codex’s consequence was simply the closest to the unique. Before I get into every little thing it did proper, although, I’ve to level out the funniest mistake it made. Instead of calling the app Foqos, it one way or the other determined to rename it Fogos.

Beyond that, although, the general proportions, spacing, card shapes, and placement of components had been impressively correct. At a look, it genuinely regarded like the identical app.

In the screenshots I offered, the Support button used a easy coronary heart icon, and Codex reproduced it in a method that really regarded like a part of the UI. Claude, alternatively, changed it with a literal purple coronary heart emoji, which instantly stood out in opposition to the in any other case polished interface. Antigravity was someplace within the center on this instance.

The similar utilized to the opposite UI components too. Codex persistently did a greater job of treating the small particulars as precise interface parts moderately than approximations. Its settings icon, buttons, card outlines, and profile part all felt a lot nearer to the reference, whereas Claude and Antigravity had been extra more likely to swap in generic-looking icons or barely reinterpret the styling. Claude, for instance, merely added emojis wherever it might!

In phrases of the general fonts and typography, Codex was additionally the closest match. The measurement, weight, and hierarchy of the textual content felt rather more trustworthy to the unique, whereas Claude tended to make some components barely bigger and Antigravity regarded somewhat extra restrained total. None of them matched each single element completely, however Codex was the one one the place I might look between the reference and the recreation with out instantly noticing that one thing felt off.

Claude did an honest job

It obtained artistic the place I actually did not need it to

Beyond the factors I discussed above, I feel Claude did a reasonably good job total. It was really the one one of many three that made among the recreated controls genuinely interactive. The toggles may very well be switched on and off as a substitute of simply being static UI components. Its button placement was additionally typically fairly near the reference, and the general construction of the screens made sense.

Where it fell behind Codex was largely within the smaller visible particulars. The emoji-heavy icon decisions had been the obvious instance, however there have been additionally a couple of locations the place the spacing, typography, and sizing felt barely off in comparison with the unique. Nothing was dramatically unsuitable, however facet by facet with the reference, it regarded extra like an excellent recreation than an virtually actual copy.

This was the one consequence the place I noticed clear hallucination, although! The New Profiles part I offered solely included the next sections: Name, Blocking Strategy, and Blocked Apps. I didn;t present a scereenshot of the total web page. Codex and Antigravity caught to the choices confirmed within the screenshot I had offered, whereas Claude added an Options part of its personal with two toggles: Enable Live Activity and Strict Mode.

Given that these choices really exist in the actual app, I obtained interested by whether or not Claude had one way or the other looked for Foqos regardless of my directions. So, I requested it instantly. Claude stated it hadn’t searched the net in any respect and that each rows had been invented based mostly on frequent conventions in focus apps.

According to it, “Live Activity” seemed like a believable characteristic for this type of app, whereas “Strict Mode” is a standard setting in app blockers. That makes the hallucination much more attention-grabbing. Claude technically broke the temporary by including UI that wasn’t seen within the screenshots, but it surely one way or the other hallucinated options that had been really there.

Antigravity had no main flaws, but it surely nonetheless got here final

It regarded good, simply not shut sufficient

Antigravity was in all probability the toughest one to criticize as a result of there wasn’t actually something unsuitable with its consequence. The app regarded polished, the screens had been structured accurately, and nothing instantly jumped out as damaged or weird.

The drawback was that, in contrast with Codex and even Claude in some areas, it merely took extra liberties with the design. Its playing cards had been somewhat extra crammed out, some spacing was completely different, and several other UI components felt extra like Antigravity’s interpretation of the screenshots than a direct recreation.

For occasion, Antigravity missed among the subtler shade cues within the reference. Both Codex and Claude observed that elements of the exercise grid used barely lighter purple shades to differentiate sure cells, whereas Antigravity rendered the part rather more uniformly. It’s a small element, however facet by facet, it made its model really feel noticeably flatter and fewer trustworthy to the unique.

Again, there was nothing unsuitable with Antigravity’s consequence, however there additionally wasn’t something about it that made me cease and suppose it had really nailed the reference. When the entire level of the experiment was to see which software might get closest to the screenshots, these small variations added up, and that was in the end why it got here final.

Ultimately, this was an excellent attention-grabbing experiment as a result of all three instruments obtained a lot nearer than I anticipated. None of them produced an ideal duplicate, however they had been all capable of infer a shocking quantity from screenshots alone, together with layouts, spacing, colours, part construction, and even some interactions!



Source link