Post 4 of 5 in our series on AI-assisted code and accessibility. Find our previous posts here, here, and here.
So far in our series on agentic coding and accessibility, we’ve focused on the how of using LLMs to revolutionize the accessibility of coding practices.
But we haven’t focused on the where and the when.
After all, baseline models, prompts, skills, Evinced Harness, and Evinced Resolve, all do their work in the developer’s environment (IDE). That is not at all a bad place to do this work. But it’s not the only place.
And at the developer’s desk, prior to a code commit, is not the only time.
Say hello to the pipeline
To “get to accessible” by relying solely on accessibility guidance in the IDE would be a tall order. Code arrives from multiple teams, from contractors, from a legacy repo nobody has touched since 2021, from dependency bumps, and from machine-generated pull requests. It’s very possible that a meaningful share of what reaches production never met an accessibility-aware developer or tool along the way. And developer habits, while they do change, are hard to change uniformly.
So where do organizations already catch what individuals miss?
In the pipeline, the company’s Continuous Integration and Deployment system (“CI/CD”).
Linting rules, unit tests, security scans: CI/CD is where quality stops being a personal habit and becomes company policy. Indeed, CI/CD is where an engineering team sets its minimum expectations for the code that the team delivers.
If we could just get accessibility to be one of these enforced expectations, the accessibility industry would be transformed.
How we got here
In the past, there were good reasons why accessibility was difficult to justify putting into CI/CD systems.
- It slowed down the cycle. The ability to fix issues automatically as part of the process didn’t exist. All the system could do was scan, and fail the build based on some minimum conditions. That has the potential to slow the development cycle down. Because an accessibility check that does no more than report a problem creates quite a lot of work: the developer must find the problem in the codebase, reproduce it, assign it, fix it, verify it, and merge it.
- Developers struggled to get the assistance they needed. Writing a fix, when most developers lack accessibility expertise, is time-consuming. A developer must translate each finding into a code change in the right parts of the code repository. That translation step is the slow, expensive, morale-draining part of accessibility work. A few years ago, to get assistance a developer might have searched Stack Overflow, or asked some peers down the hall. More recently, they’ve asked LLMs, but those systems are spotty at best, as we’ve shown.
- Detection with legacy systems was weak. At least before Evinced, the legacy tools that might have been rolled into a CI process could only detect approximately 25% of defects, and very few of those were critical ones. Even if the first two of these issues above could be overcome, this still left the bulk of the detection work to slow, expensive manual processes that usually ran only after the bugs had hit production. Better than nothing, perhaps – but not a lot better.
The transformational opportunity
What we are most excited about at Evinced is that these three reasons that have held accessibility back are no longer obstacles.
While we have described Evinced Autopilot elsewhere, it does overcome all three of these problems directly. It fixes as it runs in CI, it only asks of developers that they review already-fixed and already-tested code, and it can detect (and therefore test and fix) approximately 3X more issues than legacy open-source tools like axe-core.
Here is how it works: Autopilot runs inside the CI pipeline, on your infrastructure, with your existing LLM. When code lands, it detects accessibility defects with the Evinced engine against the rendered application, including interactive and dynamic states like open menus, modals, and logged-in views. It generates a fix for each defect.
The difference between Autopilot and a bare coding agent is what happens next. Autopilot verifies every fix independently by re-testing the original symptom against the rebuilt app. A fix that does not pass does not get counted; it gets iterated until it does pass or the system, for some rare reason, believes it can’t be done. The results arrive as an ordinary pull request, reasoning attached, for a developer to review and approve.
Benchmark results
Of course, it all has to work.
To test that, we turned Autopilot loose in CI (GitHub Actions) on three real open-source applications with deliberately different stacks.
- A multipage e-commerce storefront in vanilla JavaScript and CSS
- A React and Material UI admin dashboard
- A Vue 3 and Tailwind admin dashboard
The idea here was to imagine what would have happened if each of the engineering teams that built these projects had had Autopilot running inside their CI system at the time that they committed all the code for the entire project. (Note that Autopilot is smart enough to only analyze code that hasn’t already been checked for accessibility, but here, of course, the entire projects were new to Autopilot.)
We let Autopilot run, using Claude 4.6 Sonnet as our partner model, without stopping or touching anything between the initial commit and the completed pull requests. Exhibit 1 shows the results.
Exhibit 1. Autopilot Performance on Three Test Runs
| App | Defects Detected | Total Fixed (&Verified) | CI Clock Time | Token Cost | Cost per Verified Fix | Addressable Fixed |
|---|---|---|---|---|---|---|
| Vanilla JS + CSS storefront | 208 | 208 (100%) | 22:22 | $61 | $0.30 | 208/208 (100%) |
| React + Material UI dashboard | 179 | 179 (100%) | 28:03 | $41 | $0.23 | 179/179 (100%) |
| Vue 3 + Tailwind dashboard | 291 | 239 (82%) | 46:15 | $83 | $0.35 | 121/141 (86%) |
| Total | 678 | 626 (92%) | 96:40 | $185 | $0.30 (overall) | 508/528 (96%) |
Source: Evinced Autopilot running unattended in CI, one run per app, Claude 4.6 Sonnet. Defects are “critical” or “serious” bugs in Evinced’s classification system. Multiple defects due to an individual component are only counted once. “Addressable” excludes defects inside third-party framework and library code that the tested app’s own developers could not change.
The headline here is that Autopilot detected hundreds of bugs, and of those that were possible to fix, it fixed 96% of them, for about $0.30 each.
In terms of clock time, adding these three CI jobs together meant it took just under 97 minutes to fix 626 accessibility bugs, completely unattended. That’s a bug fixed – not just fixed, remember, but fixed, tested, and verified as fixed by a world leader in accessibility – every 11 seconds.
There is no hiding that these numbers are impressive, but there’s a benefit here that does need calling out.
Because each of these bugs, had they made it to production, would have cost hours and hours of team-wide communication and work to fix. Even if you agreed that the cost of fixing one of these bugs manually was, say $500, which we would say is optimistic, the net savings to the company owning these tested projects would be in the hundreds of thousands of dollars. Assuming, of course, that it even had the bandwidth to fix them all.
The control run
To be thorough, we wanted to consider whether a team could simply use an LLM out of the box inside a CI system, and how that would compare to Autopilot.
The obvious worry is that an unguarded agent that fixes an issue and then checks its own work is grading its own homework. LLMs are after all trained on inaccessible code, and are famously overconfident.
To find out how much that matters, we pointed the same model, Claude 4.6 Sonnet, at the same three apps. We gave it a standard accessibility prompt about how to detect and fix issues, and ran it as a comparison. The results are in Exhibit 2 below.
Exhibit 2. Autopilot vs. Prompted LLM, on Three Test Runs
| App | Defects Detected | Autopilot Fixes | LLM Fixes | Autopilot Fixes/min | LLM Fixes/min | Autopilot $/fix | LLM $/fix |
|---|---|---|---|---|---|---|---|
| Vanilla JS + CSS storefront | 208 | 208/208 (100%) | 86/208 (41%) | 9.3 | 12.3 | $0.30 | $0.14 |
| React + Material UI dashboard | 179 | 179/179 (100%) | 12/179 (7%) | 6.4 | 0.9 | $0.23 | $0.25 |
| Vue 3 + Tailwind dashboard | 291 | 239/291 (82%) | 28/291 (10%) | 5.2 | 2.1 | $0.35 | $1.01 |
| Total | 678 | 626 (92%) | 126 (19%) | 6.5 | 3.7 | $0.30 | $0.35 |
Source: Same three apps, same CI harness, Claude 4.6 Sonnet, Evinced Autopilot vs. accessibility-aware LLM prompt. Fixed counts for both are confirmed by the same post-run rescan. Counts are raw instances of critical and serious defects across all pages; a repeated component is counted once per instance.
As a top-line comparison, the LLM fixed dramatically fewer bugs than Autopilot did. Across the three three runs, Autopilot fixed 92% of all the detectable issues, and the LLM – prompted specifically to look for and fix accessibility issues – managed to only fix 19%, which is 5X less.
Looking more deeply, we note that with Autopilot, two of the three projects actually had 100% fix rates. Remember, these defects are classified in our system as critical or serious, so missing even a few of these issues could translate into a roadblock for assistive technology users. That’s particularly troubling in the LLM case, where none of the projects saw their defects drop by even 50%.
Keep in mind, too, that Autopilot fixes are verified by Evinced, and LLM fixes are… not really verified at all. The LLM simply developed a solution to the perceived problem, implemented it, and reported the item fixed. There was no built-in verification process at all.
Autopilot was also in the main cheaper per fix and faster per fix, even though it’s doing much more work – testing, iterating, and verifying.
But the important part is what did not happen: Autopilot never marked any of the leftovers done. Unverified fixes stayed unverified, which is the desirable behavior; nothing disappeared behind a green checkmark. The leftovers stayed flagged and visible, queued for the next Autopilot pass or a developer.
The takeaway
What we find so exciting about this opportunity is that it’s a chance for accessibility teams to steer the ship in the right direction.
Accessibility teams are always quite small relative to the number of developers at the same company, and reviewing releases prior to production more than once or twice a month is simply not in the cards.
That has left accessibility teams with three choices:
- Pick your battles. Be very selective about what gets reviewed prior to release. As some teams release code thousands of times a week, this would need to be very selective indeed.
- Train thousands. Many of our customers have more than 2,000 front-end developers. Even with online courseware, that’s a challenge.
- Change habits. Absent training, the only way to change developer coding habits is by introducing tools. These tools have learning curves and developers have individualized preferences. So internal promotional efforts are required, and they are hard to execute well.
Autopilot offers accessibility teams the chance to, in effect, review every single release, at high Evinced standards, without asking developers to change their behavior. And it improves itself, because every pull request reviewed by a developer teaches them something about accessibility, on their own code, at exactly the right teachable moment. That will ultimately translate into fewer bugs put into the system, and into a codebase that stays accessible.
Pipelines protect what you ship next. But most organizations are also staring at what they shipped already: an audit report, a scanner export, a spreadsheet of findings with a legal deadline attached. Turning that backlog into fixed, verified code is a different job, and in this series we’ll turn to that next.






