Skip to content
By John Hale2 min readUpdated

How to scrape websites with Python and residential proxies

A practical Python setup for permitted web data collection, including proxy configuration, session handling, retries, and choosing a pricing plan.

Configure Python requests for your data collection workflow

A reliable Python workflow starts with connection setup, sensible timeouts, and clear error handling. Check that you have permission to collect the data, follow the site’s usage rules, and use an official API when one is available.

Residential proxies route your requests through residential IP addresses. They can help you check location-specific responses and keep a session on a consistent route. They do not grant permission to access restricted data or override a website’s access controls.

Start with a simple requests setup

Basic authenticated proxy requestpython
import requests

proxy_url = "http://USER:PASS@YOUR_HOST:YOUR_HTTP_PORT"
proxies = {"http": proxy_url, "https": proxy_url}

response = requests.get("https://httpbin.org/ip", proxies=proxies, timeout=30)
print(response.status_code)
print(response.text[:200])

This is the same connection model highlighted on the homepage and in our proxies for web scraping page. If this request works, your credentials and routing are in place.

Add retries before you scale

Requests session with retriespython
import requests
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry

proxy_url = "http://USER:PASS@YOUR_HOST:YOUR_HTTP_PORT"
retry = Retry(total=5, backoff_factor=1, status_forcelist=[429, 500, 502, 503, 504])
session = requests.Session()
session.mount("http://", HTTPAdapter(max_retries=retry))
session.mount("https://", HTTPAdapter(max_retries=retry))

response = session.get(
    "https://example.com/data",
    proxies={"http": proxy_url, "https": proxy_url},
    timeout=30,
)
print(response.status_code)

Use bounded retries for temporary connection failures and server errors. Respect Retry-After and the site’s rate limits, and stop when access is denied.

When to rotate and when to stay sticky

For permitted catalog and search-result collection, choose rotating sessions when each request can be independent. For authorized logged-in workflows and multistep sessions, keep the same IP long enough to finish the action. Changing IP addresses does not remove the site’s rate limits or access requirements.

If your Python workflow includes browser automation as well as API-style requests, use the same rule: stability for authenticated sessions, distribution for crawl breadth.

When to choose GB-based vs. unlimited

Use GB-based residential proxies when you are testing selectors, validating response quality, or scraping a modest number of pages each day. Move to Residential IPv4 Unlimited when the crawl becomes continuous and you want cost predictability.

If you are still comparing pricing models, read our unlimited vs. GB-based guide after this article.

A production checklist

  1. Validate proxy credentials against a simple endpoint first.
  2. Log response status, timing, and retry count for each target.
  3. Control request rate before you increase concurrency.
  4. Use location targeting when the target changes content by market.
  5. Upgrade the billing model before growth makes cost unpredictable.

John Hale

Related articles

6 min readMarch 8, 2026

Residential vs. datacenter proxies: how to choose the right type

Residential proxies send traffic from IP addresses that consumer internet providers assign to homes; datacenter proxies use addresses owned by hosting companies. Websites and IP databases can tell the two apart, so residential suits targets that treat hosting traffic differently, while datacenter is often faster and cheaper per connection. Test your own targets before choosing.

Read article

Ready to compare current options?

Try 1 hour of Residential IPv4 Unlimited at 10 Mbps, free. No card required.

Test the network before you buy. No card required.