What Oxylabs learned from its AI parser agent
Oxylabs built an AI agent to fix its parsers, and at first its numbers looked good. It made 251 changes in under a year for $847. Then its engineers quietly stopped using it.
An Oxylabs engineering leader told the story at OxyCon 2026, along with a second agent that did succeed. Companies rarely share a failed project with numbers this openly, and the comparison is useful for any team adding AI to a scraping pipeline.
The cost of “just use AI”
The talk opened with a request the speaker’s team kept hearing from sales and product managers, to “just use AI”. He answered with Oxylabs’ own numbers.
The company scrapes about 7 billion pages a day. Parsing only 10% of them with an LLM, at about 1 cent a page for the cheapest model, would cost about $7 million a day. A mid-priced model at 5 cents a page would cost 5 times more.
So the question wasn’t whether to use AI. It was where.
The parser agent
The first agent handled parsing tasks, which the speaker said nobody on the team wanted to do anymore. It took a Jira ticket, changed the code in the parsing repository, and handed the change to an engineer to approve and deploy.
At first, it looked like a success. The team set an OKR for it, and in Q4, 51% of parser changes came from the agent.
When the OKR ended, the agent’s share fell to 14%. As the speaker noted, this happened “while the models were improving”.
The detailed numbers showed why.
| Measure | Agent | Engineers |
|---|---|---|
| Merge requests that reached production | 59% | 87% |
| Changes that passed CI tests | 76% | 91% |
| Median time to production | 7 days | 20 hours |
In addition, engineers made 3.6 changes on average to make each agent change ready for production. The agent created more work than it saved. In the speaker’s words, “Nobody quit their jobs. Nobody, you know, uh screamed or burned out. They just quietly stopped calling the agent”.
The team had measured the wrong thing. It counted merge requests and ignored what it took to get each one into production.
The agent that worked
The second project started with a different problem. Scraping engineers had to make fast decisions from a huge amount of data, spread across many dashboards. Oxylabs’ systems were receiving 1.3 million metrics a second, anti-bot systems changed all the time, and there was no documentation for them.
The team built a natural-language interface to its production systems in 3 steps.
- Access through a command line. One CLI gave the agent access to production and monitoring systems and to the strategy engine, with limits on how many jobs it could run.
- Written team knowledge. The team wrote down how it works as “skills” for the agent. This also forced the team to document its methods.
- Guardrails. For example, the agent must not draw conclusions from too few jobs, and tests need a 95% significance level.
In the demo, the speaker told the agent that a target didn’t work and asked for the reason. It checked the monitoring systems, collected sample HTML, found a block page that was missing from the site’s strategy, added it, and ran a test.
The results, as the speaker reported them.
- 500 downloads of the tool since April 2026.
- Daily active users reached 49 in one week, in a department of 16 engineers. Not everyone in that department is a scraping engineer.
- The median time to fix a scraping problem fell from 5 hours to 2, for the problems the team could solve.
Why one failed and the other worked
The speaker’s conclusion was that “enhancing and not replacing engineers is actually a much better strategy”. The details show 3 more specific differences.
The first agent did the work, and the second one helped with judgment. The parser agent produced code that engineers then had to check and fix. The second agent collected evidence and ran tests, and engineers decided what to do.
The second agent carried the team’s standards. The skills and guardrails made it work the way the team works. The speaker added that the same guardrails raised the standard for the engineers who used it.
The first project measured activity, and the second measured outcomes. Merge request counts looked good while the real cost stayed hidden. Time to fix a problem is harder to fake.
What other teams can learn from this
This talk fits a pattern across the 2026 conferences, which I covered in a longer post. Teams get the most from AI when it builds, repairs and diagnoses, and the least when it replaces an engineer’s output directly.
If you’re adding an agent to a scraping team, 4 checks from this story are worth keeping.
- Measure the full cost of each change, including review time and the fixes that follow, not only how many changes the agent makes.
- Watch usage after the launch period ends. This agent’s failure was silent. Its share of the work fell once no one was measuring it.
- Write your team’s knowledge down first. An agent with your methods and limits is more useful, and the written knowledge helps new engineers too.
- Give the agent evidence to collect, not decisions to make. Research and testing suit an agent well, and the final decision stays with an engineer.