Setup and ConfigurationAutomation Frameworks

Scrapy

Proxies for Scrapy let the tool connect through a different IP address instead of always using the IP of the machine running it. This is useful for location-based tasks, testing, monitoring, and other jobs where the connection source matters.

Scrapy still handles the main task. The proxy handles where the connection comes from.

What happens without one: IP bans, rate limits, geo walls

Without a proxy, websites keep seeing the same IP. As activity increases, that IP may hit a rate limit or get blocked.

Location can matter too. A page opened from Germany may not return the same content as one opened from the US.

A proxy for Scrapy makes it possible to connect from different IPs and locations when needed.

Which proxy type to use with Scrapy

The right proxy type depends on the job.

Proxy typeScrapingAccount managementTestingMonitoring
ResidentialGood when location and IP source matterUseful when accounts need residential IPsGood for location-based testingUseful for checking content from different locations
ISPGood when a stable IP is neededUseful for longer sessions on one IPGood for testing with a consistent IPUseful for ongoing checks from the same IP
DatacenterGood for speed and larger request volumesBetter when residential IPs are not requiredGood for general testingGood for frequent automated checks

Residential proxies are useful when the IP source and location matter. ISP proxies work well when the same IP needs to stay active for longer. Datacenter proxies make sense when the main focus is speed and scale.

The best option also depends on whether the task needs geo-targeting, a stable connection, or access to a larger pool of IPs.

How to connect a proxy in Scrapy

The Scrapy proxy setup needs the Geonode host, port, username, and password. Keep the credentials in environment variables, then load them when setting up the proxy in code.

Using Geonode? Get the username and password from Access Credentials and the host, port, and protocol from Proxy Server Information.

Settings path or code

For our test, we added a downloader middleware:

DOWNLOADER_MIDDLEWARES = {
    "geonode_scrapy.middlewares.GeonodeProxyMiddleware": 350,
    "scrapy.downloadermiddlewares.httpproxy.HttpProxyMiddleware": 400,
}

The middleware passes the proxy URL through request.meta["proxy"], which Scrapy's HttpProxyMiddleware then uses:

request.meta["proxy"] = proxy_url

You can also pass the proxy directly to an individual request:

yield scrapy.Request(
    url,
    meta={"proxy": proxy_url},
)

Keep the actual username and password in environment variables instead of putting them directly in the Python file. This setup successfully returned HTTP 200 through the proxy during our test.

Authentication format host:port:user:pass

Geonode gives you four values:

proxy.geonode.io:PORT:USERNAME:PASSWORD

Scrapy needs the same details as a proxy URL:

http://USERNAME:PASSWORD@proxy.geonode.io:PORT

Store the credentials or complete proxy value in an environment variable:

GEONODE_PROXY_URL=http://USERNAME:PASSWORD@proxy.geonode.io:PORT

Then load it in Python:

import os

proxy_url = os.environ["GEONODE_PROXY_URL"]

For our HTTP test, the resulting format was:

http://USERNAME:PASSWORD@proxy.geonode.io:9000

To check authentication, we also ran the request with incorrect credentials. Scrapy received HTTP 407 with the response:

Authentication error. Please check your authentication settings.

Scrapy surfaced the response through HttpErrorMiddleware as HttpError('Ignoring non-200 response').

Rotation vs sticky session in Scrapy

If the requests do not need to keep the same IP, use rotation.

For our test, we passed the rotating Geonode proxy to Scrapy:

proxy_url = "http://USERNAME:PASSWORD@proxy.geonode.io:9000"

yield scrapy.Request(
    url,
    meta={"proxy": proxy_url},
)

We made five sequential requests:

RequestIP
Request 1101.98.231.191
Request 2101.98.231.191
Request 3101.98.231.191
Request 4122.164.83.189
Request 5122.164.83.189

Unique IPs: 2
Rotation: Yes

The test confirms that rotation occurred, but the IP did not change on every request.

If related requests need to stay on the same IP, use a sticky session instead:

proxy_url = (
    "http://USERNAME-session-blogtest01-lifetime-10:"
    "PASSWORD@proxy.geonode.io:10000"
)

We made three requests using the same session:

RequestIP
Request 1177.73.202.78
Request 2177.73.202.78
Request 3177.73.202.78

Same IP: Yes

All three requests used the same IP in our test.

See the Geonode guides for rotating proxies and sticky sessions for the available settings.

Common proxy errors in Scrapy and how to fix them

If the proxy is not working, start with the connection details. Check the username and password, then the host, port, and protocol. The error returned by Scrapy can usually help narrow down the problem.

407 Proxy Authentication Required

A 407 usually points to an authentication problem.

When we tested Scrapy with incorrect credentials, the proxy returned:

HTTP status: 407
Authentication error. Please check your authentication settings.

Check the username and password first. If they are correct, make sure Scrapy receives the complete proxy URL with the scheme, credentials, host, and port.

ERR_TUNNEL_CONNECTION_FAILED / connection refused

A connection error usually means Scrapy could not reach the proxy.

Check the host and port first. Then make sure the protocol matches the proxy server being used.

Scrapy does not necessarily return the browser-style ERR_TUNNEL_CONNECTION_FAILED. In our failed connection test, Scrapy retried the request and eventually returned DownloadFailedError containing a Twisted ConnectionLost error.

Timeouts and empty responses

A timeout means the request did not receive a response within the configured time.

First, check whether the target URL works and whether the proxy can connect. If both are working, check the timeout configuration in Scrapy.

For our connection-failure test, we used an invalid proxy port with DOWNLOAD_TIMEOUT=5. After the retries, Scrapy returned:

DownloadFailedError(
  ConnectionLost:
  Connection to the other side was lost in a non-clean fashion.
)

So a failed proxy connection may appear as a connection error rather than a simple timeout message.

One error specific to Scrapy

One easy Scrapy-specific mistake is passing an incomplete value to request.meta["proxy"].

In our test, we used:

proxy.geonode.io:9000

without the scheme or credentials. The request returned HTTP 407.

The fix was to pass the complete proxy URL:

http://USERNAME:PASSWORD@proxy.geonode.io:9000

or let the downloader middleware build it from environment variables.

Scrapy + Geonode: what you get

Scrapy handles the requests or connections in the application. Geonode handles the proxy connection.

You can choose residential, ISP, or datacenter proxies based on the job, then use rotation, sticky sessions, and geo-targeting where needed. This keeps the proxy setup separate from the rest of the application code.

Plans vary by proxy type, with some options priced per GB.

FAQ

If you encounter any issues, refer to the troubleshooting section or Geonode support.

Was this page helpful?

On this page