Skip to content
Powered by Crawl4AI badge

Integrations

Crawl4AI proxy configuration and rotation

Pass a ProxyConfig in CrawlerRunConfig with the TrueProxies gateway as the server; put the session ID on the username for one exit per crawl, or leave it off to rotate per request.

Tested October 1, 2026 with Crawl4AI 0.9.4 through the TrueProxies residential gateway. Reviewed by John Hale. Last updated October 1, 2026.

Prerequisites

  • Python 3.10 or newer, pip install crawl4ai, then crawl4ai-setup to install the Chromium build it drives.
  • A Residential IPv4 GB-based service (Traffic pack) and its credentials from the dashboard, exported as PROXY_HOST, PROXY_USER and PROXY_PASS.
  • An allow-list of targets whose robots.txt permits crawling; Crawl4AI does not read robots.txt for you.

Basic proxy setup

The gateway host, your username and password are in the dashboard. Keep them in environment variables; never in a prompt, a repo or a log. The proxy goes on the run config, so one crawler can use different exits for different crawls. HTTP on port 8080 is the simplest; HTTPS (8443) and SOCKS5 (1080) also serve.

crawl4ai_basic.py
import asyncio, os
from crawl4ai import AsyncWebCrawler, BrowserConfig, CacheMode, CrawlerRunConfig, ProxyConfig

proxy = ProxyConfig(
    server=f"http://{os.environ['PROXY_HOST']}:8080",
    username=os.environ["PROXY_USER"],        # rotates per request: no session on the username
    password=os.environ["PROXY_PASS"],
)
run = CrawlerRunConfig(proxy_config=proxy, cache_mode=CacheMode.BYPASS, page_timeout=30_000)

async def main():
    async with AsyncWebCrawler(config=BrowserConfig(headless=True)) as crawler:
        probe = await crawler.arun(url="https://api.ipify.org?format=json", config=run)
        print("exit:", probe.html)                 # confirm the residential exit before real pages
        result = await crawler.arun(url="https://example.com/", config=run)
        print(result.success, result.status_code, len(result.markdown))

asyncio.run(main())

Sticky sessions and rotation

Sessions live on the username: USER-session-crawl01-lifetime-600 keeps one exit for ten minutes under the ID crawl01; add -country-us or -city-berlin for location. Leave the session off to rotate per request. For a crawl that must keep one IP across many pages (a paginated listing, a site that keys content to the first visit), build one ProxyConfig per crawl with its own session ID. To spread independent crawls over several exits, hand a list of such configs to Crawl4AI's RoundRobinProxyStrategy; each entry is a different session, not a different provider.

crawl4ai_sessions.py
import asyncio, os
from crawl4ai import AsyncWebCrawler, BrowserConfig, CacheMode, CrawlerRunConfig, ProxyConfig
from crawl4ai.proxy_strategy import RoundRobinProxyStrategy

HOST, USER, PASS = os.environ["PROXY_HOST"], os.environ["PROXY_USER"], os.environ["PROXY_PASS"]

def session(session_id: str, lifetime: int = 600, country: str | None = None) -> ProxyConfig:
    suffix = (f"-country-{country}" if country else "") + f"-session-{session_id}-lifetime-{lifetime}"
    return ProxyConfig(server=f"http://{HOST}:8080", username=f"{USER}{suffix}", password=PASS)

# one exit per crawl: every page of this crawl leaves through the same IP for 10 minutes
one_crawl = CrawlerRunConfig(proxy_config=session("crawl01", country="us"), cache_mode=CacheMode.BYPASS)

# several independent crawls spread over four exits, round robin
pool = RoundRobinProxyStrategy([session(f"crawl{i:02d}") for i in range(1, 5)])
spread = CrawlerRunConfig(proxy_rotation_strategy=pool, cache_mode=CacheMode.BYPASS)

async def main():
    async with AsyncWebCrawler(config=BrowserConfig(headless=True)) as crawler:
        pages = await crawler.arun_many(urls=["https://example.com/a", "https://example.com/b"], config=one_crawl)
        print([(p.url, p.status_code) for p in pages])

asyncio.run(main())

Verify traffic goes through the proxy

Crawl https://api.ipify.org?format=json with the same run config before the first real page and compare it with your direct address. The exit should be a residential IP in the country you asked for; if it is your own address, the proxy config did not apply to that run.

Measured results

Crawl4AI 0.9.4, October 1, 2026, through the TrueProxies residential gateway. Exit IPs seen: residential, one per request or per session. The gateway billed 41.2 MB for 40 page loads, about 1.03 MB per page.

Runs through the gateway
ConfigurationSuccessfulMedian time per page
ProxyConfig, rotating per request15 of 203.6 s
One session per crawl14 of 203.4 s

Errors we hit, and the fixes

Failed on navigation: Page.goto timeout 30000ms exceeded

Through a residential exit, ad-heavy pages take longer to reach the load event than from a datacenter. Five of twenty pages hit the default 30-second limit in our run. Raise page_timeout in CrawlerRunConfig to 60000 for those targets, or set wait_until to domcontentloaded when you only need the document text; the markdown is usually complete long before load fires.

Proxy error codes explained

HTTP 503 reported as anti-bot protection

Crawl4AI labels a 503 with an HTML body as anti-bot protection. In our run the only 503 came from docs.python.org, whose edge rate-limits bursts of requests from any client. Retry with a longer delay between requests to the same host; a different exit does not help when the limit is per request rate.

Session ID rejected (bad_params)

Session IDs are 1 to 32 letters or digits and the lifetime is in seconds between 60 and 86400. A hyphen or underscore inside the ID, or a lifetime outside that range, makes the gateway refuse the username. Keep IDs alphanumeric: crawl01, not crawl-01.

FAQ

Questions about Crawl4AI and proxies

How do I rotate proxies in Crawl4AI?

With TrueProxies you do not need several providers to rotate: leave the session ID off the username and the gateway gives each request a fresh exit. To rotate across a fixed set of sticky exits, build one ProxyConfig per session ID and pass them to RoundRobinProxyStrategy.

How do I keep one IP per crawl session?

Put a session ID and a lifetime on the username of that crawl's ProxyConfig, for example `USER-session-crawl01-lifetime-600`, and reuse the same run config for every page of the crawl. Lifetimes run from 1 minute to 24 hours.

Does Crawl4AI support SOCKS5 proxies?

Crawl4AI drives Chromium through Playwright, which accepts HTTP, HTTPS and SOCKS5 proxy servers. The gateway serves all three (ports 8080, 8443 and 1080). We tested HTTP on 8080; SOCKS5 with authentication depends on the Chromium build, so test it against api.ipify.org first.

How much bandwidth does Crawl4AI use per page?

Crawl4AI renders pages in Chromium, so expect browser-level traffic: the measured results above give the billed MB per page for this run, and the AI agents page gives the general figures for HTML fetches, asset-blocked browsers and full renders.

Cost per 1,000 page loads

Run your own targets through the gateway before you buy.