Integrations
Crawl4AI proxy configuration and rotation
Pass a ProxyConfig in CrawlerRunConfig with the TrueProxies gateway as the server; put the session ID on the username for one exit per crawl, or leave it off to rotate per request.
Tested October 1, 2026 with Crawl4AI 0.9.4 through the TrueProxies residential gateway. Reviewed by John Hale. Last updated October 1, 2026.
Prerequisites
- Python 3.10 or newer,
pip install crawl4ai, thencrawl4ai-setupto install the Chromium build it drives. - A Residential IPv4 GB-based service (Traffic pack) and its credentials from the dashboard, exported as PROXY_HOST, PROXY_USER and PROXY_PASS.
- An allow-list of targets whose robots.txt permits crawling; Crawl4AI does not read robots.txt for you.
Basic proxy setup
The gateway host, your username and password are in the dashboard. Keep them in environment variables; never in a prompt, a repo or a log. The proxy goes on the run config, so one crawler can use different exits for different crawls. HTTP on port 8080 is the simplest; HTTPS (8443) and SOCKS5 (1080) also serve.
import asyncio, os
from crawl4ai import AsyncWebCrawler, BrowserConfig, CacheMode, CrawlerRunConfig, ProxyConfig
proxy = ProxyConfig(
server=f"http://{os.environ['PROXY_HOST']}:8080",
username=os.environ["PROXY_USER"], # rotates per request: no session on the username
password=os.environ["PROXY_PASS"],
)
run = CrawlerRunConfig(proxy_config=proxy, cache_mode=CacheMode.BYPASS, page_timeout=30_000)
async def main():
async with AsyncWebCrawler(config=BrowserConfig(headless=True)) as crawler:
probe = await crawler.arun(url="https://api.ipify.org?format=json", config=run)
print("exit:", probe.html) # confirm the residential exit before real pages
result = await crawler.arun(url="https://example.com/", config=run)
print(result.success, result.status_code, len(result.markdown))
asyncio.run(main())Sticky sessions and rotation
Sessions live on the username: USER-session-crawl01-lifetime-600 keeps one exit for ten minutes under the ID crawl01; add -country-us or -city-berlin for location. Leave the session off to rotate per request. For a crawl that must keep one IP across many pages (a paginated listing, a site that keys content to the first visit), build one ProxyConfig per crawl with its own session ID. To spread independent crawls over several exits, hand a list of such configs to Crawl4AI's RoundRobinProxyStrategy; each entry is a different session, not a different provider.
import asyncio, os
from crawl4ai import AsyncWebCrawler, BrowserConfig, CacheMode, CrawlerRunConfig, ProxyConfig
from crawl4ai.proxy_strategy import RoundRobinProxyStrategy
HOST, USER, PASS = os.environ["PROXY_HOST"], os.environ["PROXY_USER"], os.environ["PROXY_PASS"]
def session(session_id: str, lifetime: int = 600, country: str | None = None) -> ProxyConfig:
suffix = (f"-country-{country}" if country else "") + f"-session-{session_id}-lifetime-{lifetime}"
return ProxyConfig(server=f"http://{HOST}:8080", username=f"{USER}{suffix}", password=PASS)
# one exit per crawl: every page of this crawl leaves through the same IP for 10 minutes
one_crawl = CrawlerRunConfig(proxy_config=session("crawl01", country="us"), cache_mode=CacheMode.BYPASS)
# several independent crawls spread over four exits, round robin
pool = RoundRobinProxyStrategy([session(f"crawl{i:02d}") for i in range(1, 5)])
spread = CrawlerRunConfig(proxy_rotation_strategy=pool, cache_mode=CacheMode.BYPASS)
async def main():
async with AsyncWebCrawler(config=BrowserConfig(headless=True)) as crawler:
pages = await crawler.arun_many(urls=["https://example.com/a", "https://example.com/b"], config=one_crawl)
print([(p.url, p.status_code) for p in pages])
asyncio.run(main())Verify traffic goes through the proxy
Crawl https://api.ipify.org?format=json with the same run config before the first real page and compare it with your direct address. The exit should be a residential IP in the country you asked for; if it is your own address, the proxy config did not apply to that run.
Measured results
Crawl4AI 0.9.4, October 1, 2026, through the TrueProxies residential gateway. Exit IPs seen: residential, one per request or per session. The gateway billed 41.2 MB for 40 page loads, about 1.03 MB per page.
| Configuration | Successful | Median time per page |
|---|---|---|
| ProxyConfig, rotating per request | 15 of 20 | 3.6 s |
| One session per crawl | 14 of 20 | 3.4 s |
Errors we hit, and the fixes
Failed on navigation: Page.goto timeout 30000ms exceeded
Through a residential exit, ad-heavy pages take longer to reach the load event than from a datacenter. Five of twenty pages hit the default 30-second limit in our run. Raise page_timeout in CrawlerRunConfig to 60000 for those targets, or set wait_until to domcontentloaded when you only need the document text; the markdown is usually complete long before load fires.
HTTP 503 reported as anti-bot protection
Crawl4AI labels a 503 with an HTML body as anti-bot protection. In our run the only 503 came from docs.python.org, whose edge rate-limits bursts of requests from any client. Retry with a longer delay between requests to the same host; a different exit does not help when the limit is per request rate.
Session ID rejected (bad_params)
Session IDs are 1 to 32 letters or digits and the lifetime is in seconds between 60 and 86400. A hyphen or underscore inside the ID, or a lifetime outside that range, makes the gateway refuse the username. Keep IDs alphanumeric: crawl01, not crawl-01.
FAQ
Questions about Crawl4AI and proxies
How do I rotate proxies in Crawl4AI?
With TrueProxies you do not need several providers to rotate: leave the session ID off the username and the gateway gives each request a fresh exit. To rotate across a fixed set of sticky exits, build one ProxyConfig per session ID and pass them to RoundRobinProxyStrategy.
How do I keep one IP per crawl session?
Put a session ID and a lifetime on the username of that crawl's ProxyConfig, for example `USER-session-crawl01-lifetime-600`, and reuse the same run config for every page of the crawl. Lifetimes run from 1 minute to 24 hours.
Does Crawl4AI support SOCKS5 proxies?
Crawl4AI drives Chromium through Playwright, which accepts HTTP, HTTPS and SOCKS5 proxy servers. The gateway serves all three (ports 8080, 8443 and 1080). We tested HTTP on 8080; SOCKS5 with authentication depends on the Chromium build, so test it against api.ipify.org first.
How much bandwidth does Crawl4AI use per page?
Crawl4AI renders pages in Chromium, so expect browser-level traffic: the measured results above give the billed MB per page for this run, and the AI agents page gives the general figures for HTML fetches, asset-blocked browsers and full renders.
Cost per 1,000 page loadsRun your own targets through the gateway before you buy.