150+ countries, local editions
Country, region, city and ASN targeting on the username, so each source is fetched from the audience it serves.
Proxies for AI training data are worth paying for only on the sources open datasets do not cover: fresh pages, country-specific sites and the long tail. Measured through the TrueProxies gateway, a page held about 990 tokens of main text at the median, so a million tokens cost $0.06 of traffic with an HTML fetch at the 1,000 GB tier and $1.87 with a full browser render. Start with Common Crawl and the official dumps, then proxy the gaps with rotating residential IPs on a Traffic pack.
Last updated: · Reviewed by John Hale
| Proxy type | Residential IPv4 GB-based, rotating per request |
|---|---|
| Pool | 50M+ residential IPs in 150+ countries (GB-based) |
| Targeting | Country, region, city, ASN on the username |
| Measured | 990 tokens per page; 0.073 GB per million tokens as HTML |
| Price | Traffic packs from $1.64 for 1 GB down to $0.60/GB |
| Packs | 1 GB to 5,000 GB, valid 30 days from activation |
| Connections | Up to 3,000 open connections per service |
| Trial | Free hour of Residential IPv4 Unlimited, country targeting; GB-based plans have no trial |
Strategy
Most of the public web you would crawl for a training set already exists as a download. Common Crawl publishes monthly archives of billions of pages, Wikipedia and Wikimedia ship full dumps, and many sources offer official APIs or bulk exports under clear licences. Those cost no proxy traffic, carry provenance, and are what the large models were built on.
Proxies earn their price on what the dumps miss: pages published since the last archive, sites that serve different content by country or language, sources behind geographic restrictions, and the long tail that crawlers never reached. Use residential exits for those, record every fetch, and keep the proxied share of the corpus small and deliberate.
Cost
We fetched 50 public pages (news, documentation, product pages, blog posts and forum threads) through the TrueProxies residential gateway on October 1, 2026, extracted the main text with trafilatura and counted it with tiktoken's o200k_base encoding. The median page held 990 tokens, so a million tokens is about 1,010 pages. Multiply by the MB each pipeline moves per page and you have the traffic bill.
| Pipeline | GB per 1M tokens | 1 GB tier | 100 GB tier | 1,000 GB tier |
|---|---|---|---|---|
| HTTP fetch, HTML only | 0.073 | $0.12 | $0.07 | $0.06 |
| Browser, assets blocked | 0.943 | $1.55 | $0.91 | $0.76 |
| Browser, full render | 2.307 | $3.78 | $2.24 | $1.87 |
Billed bytes, scaled from client-side accounting by the measured factor of 1.381. Pack prices: 1 GB $1.64, 100 GB $0.97/GB, 1,000 GB $0.81/GB. Token counts vary by site and language; measure your own sources before sizing a pack.
Method
A plain HTTP fetch gets the HTML and nothing else: 0.072 MB per page at the median in our test. A full browser render loads scripts, styles, images, fonts, trackers and ads, 2.284 MB per page, or 31.7x more, for the same text. Render only when the text is built by JavaScript, and when you do, abort image, font and media requests: that alone cut the browser figure by 59%.
Check a sample of each source with both methods before the crawl. If the HTML fetch already contains the article body, the browser is wasted traffic and wasted time; our browser loads took several times longer than the fetches through the same exits.
The problem
A crawler from one cloud range hits every site from the same neighbourhood of addresses. Sites rate-limit the range, serve the foreign-visitor edition, or block it, and the corpus ends up skewed toward whatever let the crawler in. Residential exits in the right countries fix the vantage point and spread the load.
They do not make a site's terms, robots rules or opt-outs go away, and they do not turn a scraped page into a licensed one. Budget bytes with the measured figures above, keep the proxied share of the corpus to the sources that need it, and log everything.
If you're comparing workflows before you commit, explore all proxy use cases to see which setup fits best.
Coverage
GB-based plans exit from 50M+ residential IPs in 150+ countries, with country, region, city and ASN targeting on the username. That is how a multilingual corpus gets the local edition of a site rather than the version served to foreign datacenters: a German exit for the German pages, a Japanese exit for the Japanese ones, each recorded with the exit country in the provenance log.
Throughput
Concurrent connections aren't billed separately. Plans allow up to 3,000 connections (Datacenter IPv6: 3,000). A training crawl is the rotating-session case: no session ID on the username, so every request can take a fresh exit and the load spreads across the pool. Keep per-domain rates polite regardless; a proxy does not change a site's rate limits.
For volume, the 100 GB and 1,000 GB Traffic packs bring the price to $0.97/GB and $0.81/GB, and the ladder continues to $0.60/GB at 5,000 GB. Packs are valid 30 days from activation, so buy for the crawl you will run this month.
Hygiene
Deduplicate at the document and near-duplicate level, strip boilerplate before counting tokens, and keep a provenance record for every page: URL, fetch time, HTTP status, exit country and the extraction method. The record is what lets you honour an opt-out or takedown later without rebuilding the corpus, and it is what makes a dataset auditable.
Compliance
Not legal advice. Follow robots.txt, honour text-and-data-mining opt-outs (including the machine-readable kinds) and site terms, and handle personal data under the rules of the jurisdictions you collect in and serve from. Providers of general-purpose AI models in the EU carry a copyright-policy duty under the AI Act, which includes respecting rights reservations. Keep the provenance log, and review the pipeline with counsel before it feeds a model.
How TrueProxies solves it
Country, region, city and ASN targeting on the username, so each source is fetched from the audience it serves.
Traffic packs from $1.64 for 1 GB down to $0.60/GB, valid 30 days from activation. A crawl that fetches HTML pays for kilobytes per page.
Tokens per page, MB per page and the billed-bytes factor were measured through the gateway on October 1, 2026, with the method published.
Recommended
1 GB for $1.64, down to $0.60/GB
Traffic packs for rotating crawls of the sources open datasets miss.
$3.99/day at 25 Mbps
A private /48 for the sources that serve IPv6.
Integration
import os, time, httpx, trafilatura
from urllib import robotparser
# pip install "httpx[http2,brotli]" trafilatura. No session on the username: every request rotates.
proxy = f"http://{os.environ['PROXY_USER']}:{os.environ['PROXY_PASS']}@{os.environ['PROXY_HOST']}:8080"
rp = robotparser.RobotFileParser("https://example.com/robots.txt"); rp.read()
with httpx.Client(proxy=proxy, follow_redirects=True, timeout=30.0,
headers={"Accept-Encoding": "gzip, br", "User-Agent": "Mozilla/5.0"}) as client:
for url in ["https://example.com/", "https://example.com/about"]:
if not rp.can_fetch("*", url):
continue
r = client.get(url)
text = trafilatura.extract(r.text) or ""
print({"url": url, "status": r.status_code, "bytes": r.num_bytes_downloaded,
"chars": len(text), "fetched_at": time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime())})
time.sleep(1.0) # per-domain politenessTested tools
httpxtrafilatura
Firecrawl (self-hosted)FAQ
About 0.073 GB with a plain HTML fetch and 2.307 GB with a full browser render, at the median page we measured (990 tokens of main text, 0.072 MB as HTML). At the 1,000 GB tier that is $0.06 and $1.87 per million tokens.
Rotating, per request, unless a source needs one exit across several pages (a paginated listing, a site that keys content to the first visit). Leave the session ID off the username for rotation and add one only for those sources.
Sticky vs rotating sessions explainedNot for open datasets or sites that serve everyone the same page; download those directly. Residential IPs are for sources that serve different content by country, filter datacenter ranges, or rate-limit shared cloud addresses.
Yes. GB-based plans exit from 150+ countries with country, region, city and ASN targeting written on the username, so each source is fetched from the audience it serves and the exit country is part of the provenance record.
Browse locationsYes. Each Traffic pack is valid 30 days from activation; a top-up carries its own 30 days and does not extend older traffic, and unused traffic does not roll over. Buy for the crawl you will run this month.
Traffic pack termsThis is general information, not legal advice. The IP you use does not change the rules: site terms, robots.txt, text-and-data-mining opt-outs, personal-data law and copyright all apply, and EU general-purpose model providers carry a copyright-policy duty. Keep a provenance log, honour opt-outs and review the pipeline with counsel.
Read the acceptable use policyOther ways teams use residential proxies.
Free trial
Start a trial, fetch a sample, and size the pack from measured bytes.
Start trial