We revisit our bot detection benchmark with stealth browsers and modern AI agents, and find detection rates ranging from 9% to 99%.
A year ago, we evaluated five leading bot detection systems (Proof of Human, Google reCAPTCHA v3, hCaptcha, Fingerprint Pro, and Cloudflare Turnstile) against automated browsers programmed to complete a set of predefined tasks. Detection rates varied substantially, ranging from 33% to 87%. Systems that incorporated behavioral signals alongside device-level signals generally outperformed those relying on device information alone.
Since then, browser automation has advanced considerably. Modern AI agents can interpret webpages, plan actions, and interact with interfaces through increasingly capable browser environments. We therefore revisited our benchmark with an expanded set of adversaries, including conventional browser automation, stealth browsers, and AI-powered browsing agents, and tested a similar set of detection systems: Proof of Human, Google reCAPTCHA v3, hCaptcha, Cloudflare Turnstile, and Cloudflare Bot Management. The results reveal similarly wide variation in detection performance.
In this post, we examine how each system performed, where it failed, and what these results suggest about the challenges of distinguishing humans from increasingly capable AI agents.
Our task used a five-question survey combining multiple-choice items with open-ended responses. Each of the five adversaries (i.e., AI agents) attempted the survey 20 times, yielding up to 100 sessions per detection system. Some detection systems failed or crashed before session completion, leaving between 89 and 100 evaluable sessions for each system.
The table below summarizes the detection systems included in the benchmark and the signals each uses to identify automated activity.
| Detection system | Description |
|---|---|
| Proof of Human | Uses a machine-learning model to classify sessions as human or automated based on both behavioral and device signals, including typing, mouse movements, clicks, and scrolling. |
| Google reCAPTCHA v3 | Assigns each session a score from 0 to 1 indicating the likelihood that the user is human. It operates without a visible challenge and considers browser-integrity and Google-side reputation signals; the site determines the threshold for accepting a session. |
| hCaptcha | Presents an “I am human” checkbox and evaluates the browser's risk profile. Depending on the result, it either verifies the session immediately or presents an image-based challenge that must be completed. |
| Cloudflare Turnstile | Performs browser-integrity checks in the background and issues a verification token when a session passes. It generally operates without a puzzle but may request user interaction when signals are ambiguous. |
| Cloudflare Bot Management | Evaluates requests using heuristics and machine-learning models trained on traffic across
Cloudflare's network. Although the full system assigns scores from 1 (automated) to 99 (human), the
Pro-tier implementation returns broader classifications such as likely_human,
likely_automated, automated, or verified_bot. |
Our set of adversaries covers increasingly sophisticated forms of automation: conventional scripted browsers, stealth browsers designed to conceal automation signals, and AI agents capable of navigating webpages and completing tasks autonomously.
| Adversary | Description |
|---|---|
| Chromium | A standard Chromium browser automated through Playwright and the Chrome DevTools Protocol, without
stealth modifications. It retains common indicators of automation, such as
navigator.webdriver and CDP-related artifacts, and serves as the baseline “obvious
bot.” |
| CloakBrowser | A stealth-focused Chromium fork that modifies browser fingerprints, including GPU, canvas, WebGL, audio, fonts, user-agent, and TLS signals, at the source-code level. We ran it through a residential proxy with its optional humanization features for mouse and keyboard input enabled. |
| Browser Use | An open-source AI-agent framework that uses an LLM (Claude Opus in our setup) and visual page information to plan actions and operate an unmodified browser from cloud infrastructure. |
| Perplexity Comet | UsesPerplexity's AI-powered browser, which uses an LLM and visual information to navigate webpages and complete tasks. |
| Grok | xAI's browser-use agent, which uses an LLM and visual information to navigate and complete web tasks. |
We found that detection performance varied substantially across systems. Aggregated across all adversaries, Proof of Human identified 99% of automated sessions, followed by reCAPTCHA v3 at 69.4%, hCaptcha at 44.3%, Cloudflare Turnstile at 42.1%, and Cloudflare Bot Management at 9.1%.
Performance differed considerably by adversary. The clearest split was between scripted browsers and AI agents. hCaptcha and Cloudflare Turnstile caught nearly every scripted Chromium and CloakBrowser session but few of the AI-agent sessions. reCAPTCHA v3 showed the opposite pattern, identifying most agent sessions but none from CloakBrowser. Cloudflare Bot Management caught few sessions from any adversary, while Proof of Human detected at least 95% of sessions overall.
As most detection systems are black-box, we cannot directly determine which signals produced each result. However, the systems' observed behavior, together with publicly available documentation, suggests several likely sources for this failure:
Several limitations should be considered when interpreting these results:
Our results suggest that bot detection is entering a new phase. Traditional automation signals remain effective against scripted browsers, but they are less reliable against AI agents that genuine browser environments intelligently. Interactive challenges may offer limited protection when vision-capable agents can interpret and complete them. And as browser-level indicators become easier to avoid, the way a user behaves throughout a session becomes increasingly important for distinguishing human activity from automation.
Future benchmarks should test larger volumes of traffic, include human participants to measure false-positive rates, and evaluate more complex tasks and emerging agents. At Proof of Human, we will continue tracking these changes and developing methods in this evolving threat landscape. To test Proof of Human on your own application, visit poh.org.
Milena Rmus, Mayank Agrawal, and Mathew Hardy work at Proof of Human, PBC, where they are building Proof of Human, an invisible authentication system for the web. Previously, they completed PhDs in cognitive science at the University of California, Berkeley (Milena) and Princeton University (Mayank and Matt).