Benchmarking Bot Detection Systems Against Modern AI Agents in 2026

We revisit our bot detection benchmark with stealth browsers and modern AI agents, and find detection rates ranging from 9% to 99%.

A year ago, we evaluated five leading bot detection systems (Proof of Human, Google reCAPTCHA v3, hCaptcha, Fingerprint Pro, and Cloudflare Turnstile) against automated browsers programmed to complete a set of predefined tasks. Detection rates varied substantially, ranging from 33% to 87%. Systems that incorporated behavioral signals alongside device-level signals generally outperformed those relying on device information alone.

Since then, browser automation has advanced considerably. Modern AI agents can interpret webpages, plan actions, and interact with interfaces through increasingly capable browser environments. We therefore revisited our benchmark with an expanded set of adversaries, including conventional browser automation, stealth browsers, and AI-powered browsing agents, and tested a similar set of detection systems: Proof of Human, Google reCAPTCHA v3, hCaptcha, Cloudflare Turnstile, and Cloudflare Bot Management. The results reveal similarly wide variation in detection performance.

In this post, we examine how each system performed, where it failed, and what these results suggest about the challenges of distinguishing humans from increasingly capable AI agents.

Methodology

Our task used a five-question survey combining multiple-choice items with open-ended responses. Each of the five adversaries (i.e., AI agents) attempted the survey 20 times, yielding up to 100 sessions per detection system. Some detection systems failed or crashed before session completion, leaving between 89 and 100 evaluable sessions for each system.

The table below summarizes the detection systems included in the benchmark and the signals each uses to identify automated activity.

Detection system Description
Proof of Human Uses a machine-learning model to classify sessions as human or automated based on both behavioral and device signals, including typing, mouse movements, clicks, and scrolling.
Google reCAPTCHA v3 Assigns each session a score from 0 to 1 indicating the likelihood that the user is human. It operates without a visible challenge and considers browser-integrity and Google-side reputation signals; the site determines the threshold for accepting a session.
hCaptcha Presents an “I am human” checkbox and evaluates the browser's risk profile. Depending on the result, it either verifies the session immediately or presents an image-based challenge that must be completed.
Cloudflare Turnstile Performs browser-integrity checks in the background and issues a verification token when a session passes. It generally operates without a puzzle but may request user interaction when signals are ambiguous.
Cloudflare Bot Management Evaluates requests using heuristics and machine-learning models trained on traffic across Cloudflare's network. Although the full system assigns scores from 1 (automated) to 99 (human), the Pro-tier implementation returns broader classifications such as likely_human, likely_automated, automated, or verified_bot.

Our set of adversaries covers increasingly sophisticated forms of automation: conventional scripted browsers, stealth browsers designed to conceal automation signals, and AI agents capable of navigating webpages and completing tasks autonomously.

Uses
Adversary Description
Chromium A standard Chromium browser automated through Playwright and the Chrome DevTools Protocol, without stealth modifications. It retains common indicators of automation, such as navigator.webdriver and CDP-related artifacts, and serves as the baseline “obvious bot.”
CloakBrowser A stealth-focused Chromium fork that modifies browser fingerprints, including GPU, canvas, WebGL, audio, fonts, user-agent, and TLS signals, at the source-code level. We ran it through a residential proxy with its optional humanization features for mouse and keyboard input enabled.
Browser Use An open-source AI-agent framework that uses an LLM (Claude Opus in our setup) and visual page information to plan actions and operate an unmodified browser from cloud infrastructure.
Perplexity CometPerplexity's AI-powered browser, which uses an LLM and visual information to navigate webpages and complete tasks.
Grok xAI's browser-use agent, which uses an LLM and visual information to navigate and complete web tasks.

Results

We found that detection performance varied substantially across systems. Aggregated across all adversaries, Proof of Human identified 99% of automated sessions, followed by reCAPTCHA v3 at 69.4%, hCaptcha at 44.3%, Cloudflare Turnstile at 42.1%, and Cloudflare Bot Management at 9.1%.

Bot benchmarking results
Figure 1: Bot benchmarking results. Each bar shows the percentage of automated sessions identified by the given detection system, aggregated across all five adversaries.

Performance differed considerably by adversary. The clearest split was between scripted browsers and AI agents. hCaptcha and Cloudflare Turnstile caught nearly every scripted Chromium and CloakBrowser session but few of the AI-agent sessions. reCAPTCHA v3 showed the opposite pattern, identifying most agent sessions but none from CloakBrowser. Cloudflare Bot Management caught few sessions from any adversary, while Proof of Human detected at least 95% of sessions overall.

Detection rates by detection system and adversary
Figure 2: Percentage of sessions identified as automated for each combination of detection system and adversary. Green cells indicate adversaries that were caught; red cells indicate adversaries that fooled the detection system.

As most detection systems are black-box, we cannot directly determine which signals produced each result. However, the systems' observed behavior, together with publicly available documentation, suggests several likely sources for this failure:

Limitations

Several limitations should be considered when interpreting these results:

Conclusion

Our results suggest that bot detection is entering a new phase. Traditional automation signals remain effective against scripted browsers, but they are less reliable against AI agents that genuine browser environments intelligently. Interactive challenges may offer limited protection when vision-capable agents can interpret and complete them. And as browser-level indicators become easier to avoid, the way a user behaves throughout a session becomes increasingly important for distinguishing human activity from automation.

Future benchmarks should test larger volumes of traffic, include human participants to measure false-positive rates, and evaluate more complex tasks and emerging agents. At Proof of Human, we will continue tracking these changes and developing methods in this evolving threat landscape. To test Proof of Human on your own application, visit poh.org.

Milena Rmus, Mayank Agrawal, and Mathew Hardy work at Proof of Human, PBC, where they are building Proof of Human, an invisible authentication system for the web. Previously, they completed PhDs in cognitive science at the University of California, Berkeley (Milena) and Princeton University (Mayank and Matt).