Skip to content

Use cases

Proxies for AI training data collection

Proxies for AI training data are worth paying for only on the sources open datasets do not cover: fresh pages, country-specific sites and the long tail. Measured through the TrueProxies gateway, a page held about 990 tokens of main text at the median, so a million tokens cost $0.06 of traffic with an HTML fetch at the 1,000 GB tier and $1.87 with a full browser render. Start with Common Crawl and the official dumps, then proxy the gaps with rotating residential IPs on a Traffic pack.

Last updated: · Reviewed by John Hale

Key facts
Proxy typeResidential IPv4 GB-based, rotating per request
Pool50M+ residential IPs in 150+ countries (GB-based)
TargetingCountry, region, city, ASN on the username
Measured990 tokens per page; 0.073 GB per million tokens as HTML
PriceTraffic packs from $1.64 for 1 GB down to $0.60/GB
Packs1 GB to 5,000 GB, valid 30 days from activation
ConnectionsUp to 3,000 open connections per service
TrialFree hour of Residential IPv4 Unlimited, country targeting; GB-based plans have no trial

Strategy

Start with open datasets, proxy only the gaps

Most of the public web you would crawl for a training set already exists as a download. Common Crawl publishes monthly archives of billions of pages, Wikipedia and Wikimedia ship full dumps, and many sources offer official APIs or bulk exports under clear licences. Those cost no proxy traffic, carry provenance, and are what the large models were built on.

Proxies earn their price on what the dumps miss: pages published since the last archive, sites that serve different content by country or language, sources behind geographic restrictions, and the long tail that crawlers never reached. Use residential exits for those, record every fetch, and keep the proxied share of the corpus small and deliberate.

Cost

Cost per million tokens

We fetched 50 public pages (news, documentation, product pages, blog posts and forum threads) through the TrueProxies residential gateway on October 1, 2026, extracted the main text with trafilatura and counted it with tiktoken's o200k_base encoding. The median page held 990 tokens, so a million tokens is about 1,010 pages. Multiply by the MB each pipeline moves per page and you have the traffic bill.

Traffic and cost per million tokens of main text, measured October 1, 2026
PipelineGB per 1M tokens1 GB tier100 GB tier1,000 GB tier
HTTP fetch, HTML only0.073$0.12$0.07$0.06
Browser, assets blocked0.943$1.55$0.91$0.76
Browser, full render2.307$3.78$2.24$1.87

Billed bytes, scaled from client-side accounting by the measured factor of 1.381. Pack prices: 1 GB $1.64, 100 GB $0.97/GB, 1,000 GB $0.81/GB. Token counts vary by site and language; measure your own sources before sizing a pack.

Scripts and raw results on GitHub

Method

HTTP fetch or headless browser

A plain HTTP fetch gets the HTML and nothing else: 0.072 MB per page at the median in our test. A full browser render loads scripts, styles, images, fonts, trackers and ads, 2.284 MB per page, or 31.7x more, for the same text. Render only when the text is built by JavaScript, and when you do, abort image, font and media requests: that alone cut the browser figure by 59%.

Check a sample of each source with both methods before the crawl. If the HTML fetch already contains the article body, the browser is wasted traffic and wasted time; our browser loads took several times longer than the fetches through the same exits.

The problem

What proxies change for a training crawl, and what they don't

A crawler from one cloud range hits every site from the same neighbourhood of addresses. Sites rate-limit the range, serve the foreign-visitor edition, or block it, and the corpus ends up skewed toward whatever let the crawler in. Residential exits in the right countries fix the vantage point and spread the load.

They do not make a site's terms, robots rules or opt-outs go away, and they do not turn a scraped page into a licensed one. Budget bytes with the measured figures above, keep the proxied share of the corpus to the sources that need it, and log everything.

If you're comparing workflows before you commit, explore all proxy use cases to see which setup fits best.

Coverage

Geo and language coverage

GB-based plans exit from 50M+ residential IPs in 150+ countries, with country, region, city and ASN targeting on the username. That is how a multilingual corpus gets the local edition of a site rather than the version served to foreign datacenters: a German exit for the German pages, a Japanese exit for the Japanese ones, each recorded with the exit country in the provenance log.

Browse proxy locations by country

Throughput

Throughput for crawlers

Concurrent connections aren't billed separately. Plans allow up to 3,000 connections (Datacenter IPv6: 3,000). A training crawl is the rotating-session case: no session ID on the username, so every request can take a fresh exit and the load spreads across the pool. Keep per-domain rates polite regardless; a proxy does not change a site's rate limits.

For volume, the 100 GB and 1,000 GB Traffic packs bring the price to $0.97/GB and $0.81/GB, and the ladder continues to $0.60/GB at 5,000 GB. Packs are valid 30 days from activation, so buy for the crawl you will run this month.

Hygiene

Data hygiene

Deduplicate at the document and near-duplicate level, strip boilerplate before counting tokens, and keep a provenance record for every page: URL, fetch time, HTTP status, exit country and the extraction method. The record is what lets you honour an opt-out or takedown later without rebuilding the corpus, and it is what makes a dataset auditable.

Compliance

Compliance

Not legal advice. Follow robots.txt, honour text-and-data-mining opt-outs (including the machine-readable kinds) and site terms, and handle personal data under the rules of the jurisdictions you collect in and serve from. Providers of general-purpose AI models in the EU carry a copyright-policy duty under the AI Act, which includes respecting rights reservations. Keep the provenance log, and review the pipeline with counsel before it feeds a model.

Acceptable use policy

How TrueProxies solves it

Rotating residential exits, billed by the gigabyte

150+ countries, local editions

Country, region, city and ASN targeting on the username, so each source is fetched from the audience it serves.

Pay per gigabyte, not per hour

Traffic packs from $1.64 for 1 GB down to $0.60/GB, valid 30 days from activation. A crawl that fetches HTML pays for kilobytes per page.

Measured, not estimated

Tokens per page, MB per page and the billed-bytes factor were measured through the gateway on October 1, 2026, with the method published.

Integration

Set up a training crawl through the gateway

  1. 1Inventory the sources. Pull everything available from Common Crawl, official dumps and APIs first, and list only the gaps for the proxied crawl.
  2. 2Sample each gap source with a plain HTTP fetch and with a headless browser; keep the browser only where the text needs JavaScript.
  3. 3Generate credentials in the dashboard, keep them in environment variables, and leave the session off the username so requests rotate.
  4. 4Set per-domain rate limits, honour robots.txt and opt-outs in code, and write the provenance record for every fetch.
  5. 5Run a pilot of a few thousand pages, read the service usage in the dashboard, and buy the Traffic pack whose per-GB price matches the measured volume.
Python (httpx, rotating)
import os, time, httpx, trafilatura
from urllib import robotparser

# pip install "httpx[http2,brotli]" trafilatura. No session on the username: every request rotates.
proxy = f"http://{os.environ['PROXY_USER']}:{os.environ['PROXY_PASS']}@{os.environ['PROXY_HOST']}:8080"
rp = robotparser.RobotFileParser("https://example.com/robots.txt"); rp.read()

with httpx.Client(proxy=proxy, follow_redirects=True, timeout=30.0,
                  headers={"Accept-Encoding": "gzip, br", "User-Agent": "Mozilla/5.0"}) as client:
    for url in ["https://example.com/", "https://example.com/about"]:
        if not rp.can_fetch("*", url):
            continue
        r = client.get(url)
        text = trafilatura.extract(r.text) or ""
        print({"url": url, "status": r.status_code, "bytes": r.num_bytes_downloaded,
               "chars": len(text), "fetched_at": time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime())})
        time.sleep(1.0)  # per-domain politeness

Tested tools

httpx logohttpxtrafilaturaPowered by Crawl4AI badgeCrawl4AIFirecrawl logoFirecrawl (self-hosted)

FAQ

Frequently asked questions

How much traffic does a million tokens of web text take?

About 0.073 GB with a plain HTML fetch and 2.307 GB with a full browser render, at the median page we measured (990 tokens of main text, 0.072 MB as HTML). At the 1,000 GB tier that is $0.06 and $1.87 per million tokens.

Rotating or sticky sessions for training crawls?

Rotating, per request, unless a source needs one exit across several pages (a paginated listing, a site that keys content to the first visit). Leave the session ID off the username for rotation and add one only for those sources.

Sticky vs rotating sessions explained
Do I need residential IPs for AI training data?

Not for open datasets or sites that serve everyone the same page; download those directly. Residential IPs are for sources that serve different content by country, filter datacenter ranges, or rate-limit shared cloud addresses.

Can I target specific countries and languages?

Yes. GB-based plans exit from 150+ countries with country, region, city and ASN targeting written on the username, so each source is fetched from the audience it serves and the exit country is part of the provenance record.

Browse locations
Do Traffic packs expire?

Yes. Each Traffic pack is valid 30 days from activation; a top-up carries its own 30 days and does not extend older traffic, and unused traffic does not roll over. Buy for the crawl you will run this month.

Traffic pack terms

Free trial

Proxy the gaps, not the whole corpus.

Start a trial, fetch a sample, and size the pack from measured bytes.

Start trial