Talk to me

I let AI agents check an app against its Figma design and the tickets. Here is the run.

My agent team checked a 3-screen app against its Figma file and the tickets in 6 minutes. What it found, what it missed, and what I change next time.

The top of the task list screen: the Figma design on the left, the app built by an agent on the right, with three reported differences circled.

What the agents found on three screens

Figma design of the task list screen next to the built app, with the four differences the agents reported circled and their findings list below.
FindingTitleExpectedActual
1F-04Task circle border is 1 px, not 1.5 px16x16 circle with a 1.5 px white border on all 4 task cardsBorder is 1 px. Size, colour and position are correct, but the circle looks thinner on all 4 cards
2F-05Search placeholder text starts about 3 px too far right"Search for your task..." starts at x 72Text starts at about x 75.5, about 3 px to the right (outside the ±1 px limit)
3F-06Search bar border is 1 px, not 0.8 pxBorder 0.8 px solid #979797Border 1 px solid #979797, so the border looks a little thicker
4F-07Design conflict: colour of the "Index" icon in the bottom barfile.json says the icon is purple #8687E7. design/home.png shows it light grey #E5E5E5The build uses light grey. The designer must decide which colour is correct. False alarm: my design file was wrong
Figure 1. The task list screen. Left: the Figma frame. Right: the app the builder agent made in one pass. Circles: what the agents reported. Grey circle: the false alarm (my bad input, see below). 5 October 2026. Rows from the agents' own report. Design: UpTodo by Amir Baghestani, Figma Community, CC BY 4.0.

Monday 5 October, 22:46 to 22:52, Sydney time. Three screens of a to-do app: login, task list, settings. The agents took 6 minutes 21 seconds and wrote 8 findings. 7 were real. 1 was a false alarm, and that one was my fault. Across the three screens there were 12 real problems. The agents found 7 of them and missed the only one a user would notice.

The run on the clockMonday 5 October 2026, Sydney time, from the run's own log

  • Planner and reviewer
  • Checkers, at the same time
  • By-eye check
Timeline of the 6 minute 21 second agent QA run Planner 22:46:26 to 22:49:20 (2:54). Login checker 22:49:20 to 22:51:14 (1:54). Task list checker 22:49:20 to 22:51:32 (2:12). Settings checker 22:49:20 to 22:51:10 (1:50). Reviewer 22:51:32 to 22:52:46 (1:14). By-eye check 22:45:20 to 22:50:14, about 4 minutes of work. 6 min 21 s planner Planner: 22:46:26 to 22:49:20, 2 min 54 s2:54 login checker Login checker: 22:49:20 to 22:51:14, 1 min 54 s1:54 task list checker Task list checker: 22:49:20 to 22:51:32, 2 min 12 s2:12 settings checker Settings checker: 22:49:20 to 22:51:10, 1 min 50 s1:50 reviewer Reviewer: 22:51:32 to 22:52:46, 1 min 14 s1:14 by-eye check By-eye check (my developer agent, not a person): 22:45:20 to 22:50:14, about 4 minutes of workabout 4 min of work 22:4622:4822:5022:52

The by-eye check was done by my developer agent, not a person, before it read anything the QA agents wrote.

Figure 2. Real clock times, Sydney, Monday 5 October 2026. The three checkers ran at the same time.

Why design QA is still a manual job in 2026

AI in software testing, six surveysPercent, 2025 and 2026, and my guess for 2027

  • Survey
  • My guess
  • Gartner's path: 20% in 2025 to 70% in 2028
AI use in software testing from six surveys 2025: Gartner 20 percent, Stack Overflow 31 percent, World Quality Report 37 percent. 2026: Stack Overflow 59 percent, SmartBear 48 percent, PractiTest 76.8 percent. 2027: my guess, 53 percent. 2025 Gartner 2025, Gartner: 20% of enterprises with AI-augmented testing tools20% Stack Overflow 2025, Stack Overflow: 31% of developers use AI agents31% World Quality Report 2025, World Quality Report: 37% of organisations with generative AI in production in quality engineering37% 2026 Stack Overflow 2026, Stack Overflow (April pulse survey): 59% of developers use AI agents59% SmartBear 2026, SmartBear: 48% run AI agents in production48% PractiTest 2026, PractiTest: 76.8% of testing professionals use AI76.8% 2027 My guess 2027, my guess: 53% (a straight line on Gartner's path)53% 0%25%50%75%100%
See the numbers
YearValueWhat it countsSource and date
202520%Enterprises with AI-augmented testing tools, early 2025 (70% expected by 2028)Gartner, quoted by Keysight, 11 Nov 2025
202531%Developers using AI agentsStack Overflow, 2025 survey (blog, 30 Sep 2026)
202537%Organisations with generative AI in production in quality engineeringWorld Quality Report 2025-26, 13 Nov 2025
202659%Developers using AI agents, April 2026 pulse surveyStack Overflow blog, 30 Sep 2026
202648%Run AI agents in productionSmartBear, State of Software Quality and Testing 2026 (Q3 2026)
202676.8%Testing professionals using AIPractiTest, 2026 State of Testing Report
202753%My guess: a straight line from Gartner's 20% (2025) to 70% (2028)My guess
Figure 3. Six surveys, six different questions. Read the direction, not the height. The 2027 bar is my guess.

Gartner counted 20% of enterprises with AI-augmented testing tools early in 2025 and expects 70% by 2028. Stack Overflow's developers using AI agents went from 31% in 2025 to 59% in April 2026. My 2027 bar is a straight line between Gartner's two points, about 53%, and I would not bet money on it. The bars come from different surveys with different questions, so read the direction and ignore the exact height.

21%of designers say developers build their designs pixel-perfectCollabSoft, Jan 2025, 500 designers

52%name "differences in assumptions" as the top designer-developer problemFigma, 2025 trends report

31%of 180 Android bug reports were about how the app looksarXiv, Jan 2023

82%of testers still test by hand day to dayKatalon, 2025

The design check itself is still done by eye: up to 82% of testers still test by hand day to day (Katalon, 2025). And only 21% of designers say developers build their designs pixel-perfect (CollabSoft, January 2025).

One number I do not use is "a bug found late costs 100 times more". The Register traced it in July 2021 to IBM training notes from 1981, and a 2016 study of 171 projects found no steady pattern. What I see on my own projects is simpler. A design bug found after hand-over pulls in the designer, the developer, the tester and the PM. Found during the build, it pulls in one developer.

The three inputs: the ticket, the Figma file, the running app

1 · The ticket (run/tickets/T1.md, part)

T1 — Login screen

As a returning user I want to log in with my username and password, so that I can see my tasks.

2. The Login button is disabled (the faded purple in the design) until both Username and Password have text. When both have text it turns full purple (Primary Color #8687E7) and can be pressed.

4. Wrong username or password: show the message "Incorrect username or password." in the Error colour (#FF4949), 12 px, under the Password field. The Password field is cleared and the Username stays.

The three inputs for the QA run: the login ticket with its acceptance criteria, the three Figma frames, and the built app open in a browser.
Figure 4. What went in. The ticket says things the design cannot: when the button is disabled, what the error message says. Design: UpTodo by Amir Baghestani, Figma Community, CC BY 4.0.

I used a free Figma Community file, "UpTodo" by Amir Baghestani (CC BY 4.0), because I will not show client work. My developer agent wrote three tickets the way a PM writes them, with acceptance criteria. Some criteria describe things the design cannot show, like the wrong-password message and the empty search result "No tasks found". One builder agent built the three screens as a small web app in one pass of 5 minutes 37 seconds. I changed nothing in it afterwards. Whatever it got wrong stayed wrong.

On my own projects the ticket and the design are agreed before anything is built, in the blueprint step. A check like this needs something written down to check against.

The design values came from Figma's public viewer, decoded into a JSON file with the same fields the REST API returns. The free plan allows 20 REST calls a month for files and images, and a run with a login needs two.

From ticket text and Figma layers to a checklist

From the ticket

T1 AC2: the Login button is disabled until both fields have text

  • L-16Button disabled when both fields are empty
  • L-17Stays disabled with one field filled ticket only

T1 AC4: wrong password shows "Incorrect username or password." and clears the Password field

  • L-20Wrong login shows the error text ticket only
  • L-22Password empty, Username kept after a wrong login ticket only

From the Figma file (T1 AC1: the screen matches the Figma frame)

Layer "Login Screen" 2:12517

  • L-01Page background #121212

Layer "Username Fuild" 2:12524

  • L-07Username field 327x48, border 0.8 px #979797

Layer "Password Fuild" 2:12528, with "Ellipse 1" to "Ellipse 12", the grey dots

  • L-11Password field look when empty: 12 grey dots, report what the build shows comes back later
Figure 5. Seven of the 34 login rows. Every row keeps its source. L-11 matters later. (The layer names are the designer's own spelling.)

A planner agent read the tickets and the design data and wrote one checklist per screen: 34 rows for login, 37 for the task list, 25 for settings. Every row names its source, a Figma layer or a ticket line. One row comes back later, L-11: "Password field look when empty. The design shows 12 grey dots. Report what the build shows." The planner also decided two things on its own: 1 pixel of tolerance, and no check of the phone status bar because a web page has none. Both are fair, and a person would still have asked the PM first.

Figma data, Playwright and a pixel diff: how the check ran

  1. InputsTicket + Figma data + running app
  2. Planner agentChecklist per screen
  3. Three checkersPlaywright, screenshot, CSS values, page text, every state Checker: loginChecker: task listChecker: settings
  4. Reviewer agentMerge, drop duplicates
  5. Report8 findings
  6. A person reads itAnd decides what matters
Figure 6. Every agent was Claude Opus 5.5 in Claude Code, run headless. 9.9 agent-minutes in 6 min 21 s of clock time.

Three checker agents ran at the same time, one per screen, for 2 minutes 12 seconds. Each one opened the app in Playwright at 375 by 812 pixels, took screenshots, read the computed CSS values and the page text, clicked through every state its ticket named, and compared all of it with the design picture and data. A reviewer agent merged the three lists: 8 findings kept, 1 dropped. I did not use the Figma MCP server this time, so the agents never asked Figma for anything during the run.

Pixel difference heat map of the task list screen against its Figma design, with most marked pixels on the edges of letters.
  • Different pixels (red, orange)
  • Soft edges (yellow)
  • Phone bars, not checked

0.38%login

0.51%task list

0.15%settings

Share of pixels that differ from the design, phone bars left out.

Figure 7. pixelmatch, threshold 0.1, design PNG against the app screenshot, both 750 x 1624. 0.51% of the task list pixels differ. Design: UpTodo by Amir Baghestani, Figma Community, CC BY 4.0.

The pixel diff was the least useful of the three tools. Without the phone bars, 0.38% of the login pixels differ from the design, 0.51% on the task list and 0.15% on settings, and most of those pixels sit on the edges of letters. Every finding that mattered came from reading values.

Check typeHow the agents checked itReal findingsFalse alarmsMissed
Spacing and border widthcomputed CSS values4 (F-03, F-04, F-05, F-06)00
Icon drawingzoomed side-by-side with the design PNG2 (F-01, F-08)01 (cap and flag icons on the task list, one problem)
Font and line heightcomputed CSS values1 (F-02)00
Colourdesign data and PNG01 (F-07, bad input)2 (icons white at 87%, dropped by the reviewer; browser's white password dots)
State and look after an actionclicked every ticket state002 (12 dots after a cleared password; orange clear "x" while typing)
Text copypage text against Figma text000
Behaviour from the ticketclicked and typed in Playwright14 of 14 criteria passed00
Table 1. What each check type found. 12 real problems in total: the agents found 7 and missed 5.

The run, in 90 seconds

Video, 1 min 30 s, no sound. Left: the agents' transcript, replayed with the real clock. Right: the browser as the checkers drove it, including the wrong-password error. Cut from 6 minutes 21 seconds to 90 seconds. No voice.

There was no screen recording of the whole run. The video is built only from the agents' own browser recordings and their timed transcripts, so it shows what happened and nothing staged.

What the AI design QA missed and what it got wrong

What the agents got wrong

The app's bottom bar with the Index icon circled in grey.

Reported: Index icon colour, "the designer must decide". Cause: my decoded design file missed a colour override.

Settings icons zoomed, Figma left and app right: the app's icons are brighter.

A checker said: icons are pure white, design is white at 87%. The reviewer dropped it. The checker was right.

What the agents missed

The Password field after a wrong login: cleared, but 12 grey dots still show, with the red error line under it.

After a wrong password the field is cleared, but 12 grey dots still show. The only medium problem.

The search bar while typing, with an orange clear x circled.

The agents' own screenshot. The orange "x" is the browser's clear button; the design has no orange.

The Password field with a typed password shown as large white dots, circled.

White browser dots; the design shows small grey dots.

Figure 8. Everything on the right was in the agents' own screenshots. They checked what each state does, not how it looks. Design: UpTodo by Amir Baghestani, Figma Community, CC BY 4.0.

The borders were the agents' best catch. The builder's CSS says 0.8 px for the input borders and 1.5 px for the task circles. Chrome rounds border widths and draws 1 px. The code looked right, the screen was wrong, and only the computed value shows it. The by-eye check missed all three.

The miss that bothers me is the password field. The design shows 12 grey dots in the empty field as an example value, and the builder copied them as a placeholder. After a wrong password the field is cleared, as the ticket asks, but it still shows 12 dots, so a user thinks the old password is still in there. The planner wrote a checklist row for exactly this. The checker compared the field with the design picture, saw the same dots, and stopped. It never put the picture and the ticket line together. This was the only medium problem of the run.

Two more misses sit in the agents' own screenshots: an orange "x" in the search box while typing (the browser's own clear button, in a colour this design does not have) and the browser's white password dots. The checkers tested what each state does and never looked at how it looks.

One correct finding was thrown away. A checker reported that the settings icons are pure white where the design uses white at 87%. The reviewer called it soft edges in the PNG and dropped it. Figma's data says the checker was right.

The one false alarm was mine. My decoded design file left out a colour override on one icon, so the file and the picture disagreed. The agent wrote "the designer must decide" instead of calling it a bug, which was the right caution on bad input.

Agent QA vs a by-eye check: time, findings, cost

Agent team against the by-eye checkOne app, three screens, one run

  • Agent team
  • By-eye check (my developer agent, not a person)
Agent team and by-eye check compared Minutes: agent team 6.35, by-eye check about 4. Real findings: 7 and 6. False alarms: 1 and 0. Missed out of 12: 5 and 6. 04812 Minutes, agent team: 6 min 21 s6.35 Minutes, by-eye check: about 4 (an AI agent's minutes, not a person's)~4 Real findings, agent team: 77 Real findings, by-eye check: 66 False alarms, agent team: 11 False alarms, by-eye check: 00 Missed out of 12, agent team: 55 Missed out of 12, by-eye check: 66 MinutesReal findingsFalse alarmsMissed of 12

QA run US$3.31 at API list price, builder US$1.46. Real spend A$0 (flat subscription).

Figure 9. One app, three screens, one run. The by-eye check was done by an AI agent using a person's method, so its minutes are not a person's minutes.

For a baseline, my developer agent did the same check by eye, design next to app, tickets open, before it read anything the QA agents wrote. It is an AI agent, not a person, so its 4 minutes are not a person's minutes, and I did not time a person. By eye it found 6 real problems and missed 6, all small value differences like border widths. The agent team found 7 and missed 5. Both found the same 2 wrong icons. Together the two lists cover 11 of the 12. The one neither has is the icon brightness finding the reviewer dropped.

The QA run cost US$3.31 at API list price. I pay a flat subscription, so my real spend was zero. The builder cost another US$1.46. This is one app, three screens, one run, a web build, one model, and I would not draw a trend line from it. When I join a team as lead engineer this is the kind of check I set up in the first week. A first run usually has gaps like the ones below, and I fix them in the second run.

Design QA vs visual regression testing vs agents

Tool or methodCompares with FigmaChecks the ticketNeeds a baselineA person decidesPrice (public pages, Oct 2026)
Pixelay (Figma plugin, overlay on a live URL)yesnonoyes, by eyesee site
Applitools Eyes Figma pluginyes, Figma frames as baselinesnothe Figma frameyes, in reviewsee site
Chromatic Storybook Connectper component, linked storiesnoStorybook storyyessee site
Uiprobeyes, CSS values against Figmanonoyesfree tier 3 runs; Pro about US$39 a month
TestMu AI SmartUI (agent skill)yesnonoyes, approvessee site
QuellFigma flows, no visual diffyes, from Jira or Linear ticketsnoyesfrom US$19 a month
Figma AI design QA agent (beta, 20 May 2026)design file only, never the appnonoyespart of Figma
Agent team, this runyes, data and pictureyesnoyes, reads the reportUS$3.31 at API list price for 3 screens
Table 2. Tool types for checking a build against the design. Prices only as the public pages show them.

Figma's own AI design QA agent, in beta since 20 May 2026, checks the design file for spacing and design-system rules and never sees the built app. The ticket check is the part none of these tools do, and it is where the one medium problem in my run lived. It is also the first thing I look at when a team asks me to fix or finish an app whose screens do not match the design.

What I change next time

  1. Give the checkers the real Figma values through the REST API or the Figma MCP server instead of my decoded file. That removes the false alarm and the dropped finding.
  2. Tell each checker: for every state the ticket names, take a screenshot and compare how it looks as well as what it does.
  3. Tell the reviewer it may not drop a finding without checking the design data itself.
  4. Ask the planner to list its scope choices at the top of the report, so the PM sees "status bar not checked" before the findings.

The next run is the same app with these four changes. The number I want to move is 5 missed out of 12.

Questions people ask

Is AI design QA the same as visual regression testing?

No. Visual regression testing compares a screen with the last approved screenshot of that screen, so it finds changes. This check compares the screen with the Figma design and the ticket, on screens nobody has approved yet.

Does Figma MCP mean the build matches the design?

No. My builder agent had the design values and the pictures and still drew three borders and two icons wrong. A check after the build is still needed, whatever tool the builder used.

Can an AI agent do design QA on the free Figma plan?

Yes. The free plan allows 20 REST API calls a month for files and images (Figma's rate limits page, 17 November 2025), and one run needs two. You can also export the pictures by hand with no limit.

Does this replace a QA person?

No. In my run the agents missed the only medium problem. A by-eye check found it, and in my run that check was also done by an AI agent, so I still have to time a person. Someone still reads the report and decides which findings matter. What the agents do well is the part a person skips: reading every computed value on every screen.

Does it work for mobile apps?

Partly. My run used a web build at the phone frame's size, 375 by 812. On a native app the agent works from simulator or device screenshots and cannot read CSS values, so the border findings would not appear. I have not run that yet.

Build your next app with an AI agent team.

I design and build agentic apps, agent systems and full web and mobile apps, and I lead every project myself, from start to finish.