Skip to content
Firecrawl logo

Integrations

How to use a proxy with self-hosted Firecrawl

Set PROXY_SERVER, PROXY_USERNAME and PROXY_PASSWORD in the self-hosted .env to the TrueProxies gateway and your credentials; every scrape and crawl then leaves through the residential pool.

Tested October 1, 2026 with Firecrawl v2.11.444 through the TrueProxies residential gateway. Reviewed by John Hale. Last updated October 1, 2026.

Firecrawl's cloud plans bundle their own proxies, and the hosted API has no setting for yours. This page is for the self-hosted stack, where the outbound proxy is yours to choose.

Prerequisites

  • Docker and Docker Compose, and a checkout of the Firecrawl repository at the release you intend to run.
  • A Residential IPv4 GB-based service (Traffic pack) and its credentials from the dashboard.
  • Enough memory for the API, worker and Playwright services; the stack renders pages in a browser.

Environment config

The gateway host, your username and password are in the dashboard. Keep them in environment variables; never in a prompt, a repo or a log. Firecrawl reads one proxy for the whole instance from the .env file next to docker-compose.yaml (the template is apps/api/.env.example). PROXY_SERVER takes a full URL or host:port; the username and password are separate variables.

.env
# Outbound proxy for every scrape and crawl this instance runs
PROXY_SERVER=http://YOUR_GATEWAY_HOST:8080
PROXY_USERNAME=YOUR_USERNAME            # add -session-<id>-lifetime-<seconds> to keep one exit
PROXY_PASSWORD=YOUR_PASSWORD

# then
# docker compose up -d --build
# curl -s http://localhost:3002/v1/health

Sticky sessions and rotation

Firecrawl applies the proxy per instance, so the session policy is set once, in PROXY_USERNAME. Sessions live on the username: USER-session-crawl01-lifetime-600 keeps one exit for ten minutes under the ID crawl01; add -country-us or -city-berlin for location. Leave the session off to rotate per request. For an instance that scrapes unrelated URLs, leave the session off and let every request rotate. For an instance dedicated to one site that needs a stable visitor, add a session ID and a lifetime; restart the stack to change it.

scrape.py
import httpx

API = "http://localhost:3002"

# confirm the exit the instance uses before real pages
probe = httpx.post(f"{API}/v1/scrape", json={"url": "https://api.ipify.org?format=json", "formats": ["markdown"]}, timeout=60).json()
print(probe["data"]["markdown"])            # the residential exit, not your server's IP

page = httpx.post(f"{API}/v1/scrape", json={"url": "https://example.com/", "formats": ["markdown"]}, timeout=60).json()
print(page["success"], len(page["data"]["markdown"]))

Verify traffic goes through the proxy

Scrape https://api.ipify.org?format=json through the instance and compare the address with your server's own. If the probe returns your server's IP, the variables were not read: check that .env sits next to the compose file and that the api and worker services restarted after the change.

Measured results

Firecrawl v2.11.444, October 1, 2026, through the TrueProxies residential gateway. Exit IPs seen: residential, one per request. The gateway billed 27.5 MB for 20 page loads, about 1.38 MB per page.

Runs through the gateway
ConfigurationSuccessfulMedian time per page
/v1/scrape, markdown20 of 2011.2 s

Errors we hit, and the fixes

DATABASE_URL is not configured; every worker exits after 12 seconds

The example .env ships with USE_DB_AUTHENTICATION=true, the cloud setting, so the API and the queue workers look for a Supabase database, log DATABASE_URL is not configured, and the api container stops before it ever listens on port 3002. Set USE_DB_AUTHENTICATION=false in .env for a self-hosted stack and start it again.

Image pull fails with an unexpected status from Docker Hub

docker compose up --build pulls the RabbitMQ, Postgres and FoundationDB images from Docker Hub; one of those pulls failed mid-blob on our first attempt. It was a registry hiccup, not a configuration problem: the same command succeeded on the retry.

Scrapes are slower than a direct fetch

Through a residential exit the median scrape took 11 seconds and the slowest news page 29, because the Playwright service renders the page and waits for it to settle. Raise the request timeout in the scrape body for heavy sites, and keep formats to markdown when you do not need the raw HTML.

FAQ

Questions about Firecrawl and proxies

Does Firecrawl cloud need my own proxies?

No. The hosted service includes its own proxy layer and offers no setting for a third-party proxy. Bring your own proxies only on the self-hosted stack.

Which environment variables set the proxy?

PROXY_SERVER (a full URL such as http://host:8080, or host:port), PROXY_USERNAME and PROXY_PASSWORD, in the .env file the compose stack reads. Leave the username and password unset only for an unauthenticated proxy.

How do I rotate IPs?

Leave the session ID off PROXY_USERNAME and the TrueProxies gateway gives every request a fresh residential exit. Firecrawl itself has no per-request proxy switch on the self-hosted stack.

How much bandwidth does Firecrawl use per page?

Firecrawl renders pages in a browser service, so expect browser-level traffic rather than plain HTML fetches. The measured results above give the billed MB per page for this run.

Cost per 1,000 page loads

Run your own targets through the gateway before you buy.