# Change Log (/docs/changelog)
After months of work, we’re excited to introduce our **updated dashboard** — a major release focused on simplicity, speed, and flexibility.\
Below is an overview of the key improvements.
***
## Redesigned Dashboard
The dashboard has been completely reworked for better usability and performance.
* Clean, minimal layout
* Faster navigation
* Intuitive access to essential tools
Everything is now more accessible and built with your feedback in mind.
***
## Proxy Membership
We’ve replaced the old **Proxy Residential Membership** with the new **Proxy Membership** — a unified and flexible way to manage all your proxies.
* Manage both **Residential** and **Datacenter** IPs
* Filter or combine IP types effortlessly
* Simplified structure for all users
For API integration details, see the [API documentation](https://docs.geonode.com/).
***
## Endpoint Generator
Our most requested tool is here — the **Endpoint Generator**.
### Create endpoints in multiple formats
Multiple formats are supported for easy integration.
### Copy or export instantly
Copy endpoints or export them to your preferred format in one click.
### Integrate with your favorite tools
No manual setup — just plug it in and start using it.
***
## Updated Host URL
We’ve updated our host structure to improve consistency.
**Old host:** `premium-residential.geonode.com`
**New host:**
`proxy.geonode.io`
The old hostname will continue to function, but only the new one will appear in the dashboard moving forward.
***
## Feature Request System
Your feedback drives our development.\
With the new **Feature Request Tool**, you can:
* Submit new feature ideas directly from your dashboard
* Vote for the improvements you want most
You’ll find it under **Profile → Feature Request**.
***
## Changelog Access
You’re reading it!\
From now on, every platform update and feature release will appear here, so you can always stay informed.
***
## Revamped API Documentation
The [API documentation](https://docs.geonode.com/) has been fully updated for better clarity and ease of use.
* Improved structure and examples
* Updated endpoints
* More guides coming soon
***
## Under-the-Hood Improvements
We’ve made significant backend enhancements to boost speed, reliability, and stability.
| Area | Improvement |
| --------------- | --------------------------------------------- |
| **Speed** | Faster performance across dashboard and APIs |
| **Latency** | Reduced response times for smoother operation |
| **Reliability** | Higher success rates for proxy connections |
| **Stability** | Optimized infrastructure and bug fixes |
These upgrades work silently to ensure your experience is faster and more dependable.
***
## Try It Out
[Log in now →](https://app.geonode.com/residential-proxies)\
Explore the new dashboard and let us know what you think.
***
*Thank you for being part of our journey — more great updates are on the way!*
# May 2025 (/docs/changelog/may2025)
After months of hard work, we're thrilled to introduce our **updated dashboard** and a major platform upgrade.\
This release focuses on speed, reliability, and a smoother user experience.
***
## New Website & Branding
The new **Geonode website and visual identity** mark a major milestone in our journey.\
This redesign reflects our growth and focuses on developers — with a cleaner layout, faster load times, and improved accessibility.
***
## Proxy Infrastructure 2.0
We’ve launched **Proxy Infrastructure 2.0** — a full rebuild of our global proxy network.
* Significantly reduced latency
* 99 % + success rate across all regions
* Major speed boost for **SOCKS5** users
* Expanded IP pool in key markets like the United States
This upgrade delivers faster, more stable connections and improved reliability worldwide.
***
## Upgraded SOCKS5 Infrastructure
Our SOCKS5 network has been completely re-engineered for improved throughput and connection stability.\
You’ll notice faster response times and fewer dropped sessions across all endpoints.
***
## New Pricing Plans
We’ve introduced three new plans to better fit different use cases:
| Plan | Monthly Price | Included GB | Rate per GB |
| ------------ | ------------- | ----------- | ----------- |
| **Starter** | $50 / month | 50 GB | $1.00 |
| **Growth** | $200 / month | 267 GB | $0.75 |
| **Business** | $500 / month | 1000 GB | $0.50 |
***
## 1 TB Free Trial in Dashboard
Qualified business users can now apply for the **1 TB free trial** directly inside the dashboard — no sales call required.\
Get access, test at scale, and start evaluating performance in minutes.
***
## User-Friendly Billing & Grace Period
Billing and subscription handling have been redesigned for flexibility:
* Failed payment? You now have **5 days** to resolve it before cancellation.
* Cancelled subscription? You can reactivate anytime before the billing period ends.
* Finished plan? You still have **5 days** to restore it and keep unused bandwidth.
***
## Improved Usage Graphs
Usage analytics now include:
* Bandwidth by **hour, day, week, or month**
* Look-back period of up to 31 days
* Clearer trend visualization and comparison
These updates make monitoring and optimization much easier.
***
## Error Handling and Service Codes
Error handling has been simplified and clarified:
* Cleaner error messages
* Fewer, more actionable error codes
See the [updated Error Handling Guide](https://docs.geonode.com/docs/proxies/api-reference/error-handling).
***
## Geonode SDK Program Launch
We’ve officially launched the **Geonode SDK Program** — enabling developers to monetize apps across iOS, Android, desktop, smart TVs, and more.\
[Learn more and apply here](https://geonode.com/app-sdk).
***
## Expanded Event-Based Email Notifications
You’ll now receive clear, event-based notifications for:
* Purchases and upgrades
* Downgrades and usage warnings
* Trial updates and billing issues
Stay informed about every important account event.
***
## Shape the Future of Geonode
Want to help define what comes next?\
[Vote on and suggest new features](https://feedback.geonode.com/) directly through our feedback portal.
# Feature Requests (/docs/getting-started/feature-request)
Have an idea for a new Geonode feature or an improvement to an existing one?\
You can easily submit your request and explore what other users are suggesting.
***
## Submit Your Feature Request
To request a new feature or enhancement, visit:\
➡️ [Geonode Feature Requests](https://feedback.geonode.com/)
Provide as much detail as possible about the feature, including what problem it solves and how it would improve your workflow.
***
## Why Submit a Feature Request?
* 🧠 **Influence future updates** – Help shape Geonode’s roadmap based on real user needs.
* ⚙️ **Improve functionality** – Suggest enhancements that boost performance, security, or usability.
* 💬 **Collaborate with the community** – Upvote and comment on ideas from other users.
***
## How It Works
1. **Browse existing requests**\
Someone may have already suggested your idea — check first before creating a new one.
2. **Submit a new request**\
Provide a clear title, description, and (if possible) your use case or expected benefit.
3. **Vote and comment**\
Upvote ideas you support and share feedback to help prioritize development.
4. **Track progress**\
Follow the status of your requests as the Geonode team reviews and implements new features.
***
## Final Notes
Geonode values community input — every feature request helps make the platform better.\
Your ideas directly contribute to improving **performance, reliability, and user experience** for everyone.
# Getting Started (/docs/getting-started)
Welcome to Geonode. Use this section to learn the product basics and set up your account.
## Start here
* [What is Geonode?](/docs/getting-started/what-is-geonode/what-is-geonode) — how Geonode works and who it is for
* [Account setting and billing](/docs/getting-started/account-setting-and-billing/profile-setting) — profile, subscriptions, payments, and wallet
## Next steps
* [Proxies](/docs/proxies) — set up proxies and use the Proxy API
* [Web Data](/docs/scraper-api) — extract, crawl, map, and search web content
# Service Status (/docs/getting-started/service-status)
Stay informed about the current operational status of **Geonode services**, including proxy availability, system uptime, and scheduled maintenance.
***
## Service Status Overview
🟢 **Coming Soon**\
A real-time **Service Status Dashboard** is in development.\
It will let you:
* Monitor **system health** in real time
* View **proxy uptime and latency metrics**
* Track **maintenance windows and past incidents**
* Subscribe to **status updates and alerts**
***
## Need Help Right Now?
If you’re currently experiencing connection issues or service interruptions, please reach out to our support team:
➡️ [Contact Geonode Support](https://geonode.com/contact)
Our team is available to help diagnose and resolve issues as quickly as possible.
***
## Stay Updated
Check back soon for the live **status.geonode.com** page —\
your one-stop hub for Geonode uptime, maintenance, and performance insights.
# Subscription, Cancellations, Reactivation and Grace Period (/docs/getting-started/subscription-cancellations-reactivation-and-grace-period)
This guide explains what happens when you **cancel your Geonode subscription**, how you can **reactivate it**, and what **grace periods** apply.\
Understanding these timelines helps prevent loss of access or unused bandwidth.
***
## Canceling During Your Billing Period
When you cancel your subscription, it remains active until the **end of your current billing cycle**.
* You can continue using the service (including proxy access and any remaining bandwidth) until the period ends.
* You can **reactivate anytime before the billing cycle ends** — this cancels your cancellation and keeps your subscription active without interruption.
***
## After Your Billing Period Ends
Once your billing cycle ends and the subscription is fully canceled:
* You automatically enter a **5-day grace period**.
* During this time, you can **reactivate your subscription** without losing data.
### If you reactivate within the grace period:
* Your subscription is **restored immediately**.
* Any **unused bandwidth** is reinstated.
* Proxy access resumes without delay.
### If you do **not** reactivate within 5 days:
* The subscription remains permanently canceled.
* Any **unused bandwidth is lost** and cannot be recovered.
***
## Failed Payments
If your payment fails:
* Access to proxies is **temporarily paused**.
* You have **5 days** to update your payment method and fix the issue.
* If resolved within 5 days, your subscription continues as normal.
* If not resolved, your subscription may be **canceled** and remaining bandwidth **forfeited**.
***
## Summary
| Scenario | Grace Period | Can Reactivate? | Bandwidth Restored? | Access Status |
| ------------------------------- | ------------ | --------------- | ------------------- | ------------------ |
| **Manual Cancellation** | ✅ 5 days | ✅ Yes | ✅ Yes | Temporarily Paused |
| **End of Billing Period** | ✅ 5 days | ✅ Yes | ✅ Yes | Inactive |
| **Failed Payment (Unresolved)** | ❌ | ❌ No | ❌ No | Canceled |
***
If you need help managing your subscription or reactivating access, please contact:\
➡️ [Geonode Support](https://geonode.com/contact)
# Support (/docs/getting-started/support)
## Geonode Community and Support
Join the Geonode Community to learn, share, and connect with proxy users at all experience levels.\
Get access to helpful guides, expert tips, and support from the Geonode team.
💬 [Join us on Discord](https://discord.com/invite/32RXgzgeAf) to get started, meet other users, and stay up to date with the latest in proxy services.
***
## Contact Methods
If you need help or want to reach our support team, you can contact us using one of the methods below.
### Email
Reach out to us directly at [hello@geonode.com](mailto:hello@geonode.com).\
Use this for general inquiries, technical issues, or billing questions.
***
### Contact Form
Submit your request using our [Contact Form](https://geonode.com/contact).\
Recommended for non-urgent questions, suggestions, or feedback.
***
### Book a Call
Schedule a 30-minute support call with our team through [Calendly](https://calendly.com/maria-geonode/30min).\
Perfect for detailed troubleshooting or guided onboarding.
***
Our support team is available around the clock to help with technical issues,
billing, or account management.
# Useful Links (/docs/getting-started/useful-links)
This page provides a quick overview of key Geonode resources — manage your proxies, stay updated, and get support whenever needed.
***
## Geonode Dashboard
Your main control panel for managing proxies, monitoring usage, and configuring account settings.\
➡️ [Access the Geonode Dashboard](https://app.geonode.com/proxies)
***
## Discord Community
Join the Geonode community to connect with other users, get real-time help, and receive product updates.\
➡️ [Join the Geonode Discord](https://discord.com/invite/32RXgzgeAf)
***
## Geonode Blog
Read articles, tutorials, and insights from the Geonode team to enhance your proxy experience.\
➡️ [Visit the Blog](https://geonode.com/blog)
***
## Contact Support
Need help? Reach out to the Geonode support team for personalized assistance.\
➡️ [Contact Support](https://geonode.com/contact)
***
## Subscription Plans
Compare available proxy plans and find the best option for your needs.\
➡️ [View Subscription Plans](https://geonode.com/1TB-program)
***
> \[!NOTE]\
> Keep this page bookmarked for quick access to key Geonode resources.\
> Our support team is available 24/7 to help with any questions or issues.
# Proxies (/docs/proxies)
Use this section to set up Geonode proxies and work with the Proxy API.
## Getting Started
Start here if you are new to Geonode proxies.
* [Quick Start Guide](/docs/proxies/getting-started/quick-start) — connect your first proxy
* [Prerequisites](/docs/proxies/getting-started/prerequisites/access-credentials) — credentials and proxy server details
* [Setup and Configuration](/docs/proxies/getting-started/setup_and_configuration/desktop-OS/windows) — desktop, mobile, browsers, and tools
* [Knowledge Base](/docs/proxies/getting-started/knowledge-base/overview) — proxy concepts and service basics
## Products
Choose the proxy product that matches your use case.
* [Residential Proxies](/docs/proxies/guides/residential-proxies/overview) — real residential IPs for geo targeting, sticky sessions, and rotating traffic
* [Unlimited Residential Proxies](/docs/proxies/guides/unlimited-residential-proxies/00_unlimited_residential_proxies) — speed-based residential plans with unlimited traffic
* [ISP Proxies](/docs/proxies/guides/isp-proxies/isp-proxies) — ISP-assigned proxy IPs you can manage, organize, and monitor from the dashboard
* Rotating Datacenter Proxies — high-speed datacenter IPs for large-scale rotating traffic
## API Reference
Use the Proxy API to target locations, manage sticky sessions, and track usage.
* [Introduction](/docs/proxies/api-reference) — authentication and basic workflow
* [Geo Targeting](/docs/proxies/api-reference/geo-targeting/geo-targeting-options) — country, state, city, ISP, and OS targeting
* [Sticky Sessions](/docs/proxies/api-reference/sticky-session/get-sticky-session/get-create) — create and release sticky sessions
* [Error Handling](/docs/proxies/api-reference/error-handling) — common API errors and responses
## Next Steps
* Start with the [Quick Start Guide](/docs/proxies/getting-started/quick-start)
* Pick a product guide above for your use case
# Overview (/docs/scraper-api)
Use this section to extract web content, discover URLs, run search jobs, and connect Geonode to AI tools with MCP.
## Getting Started
Start here if you are new to the Scraper API.
* [Quick Start Guide](/docs/scraper-api/quick-start) — get your API key and authenticate requests
* [Before You Start](/docs/scraper-api/getting-started/00_before_you_start) — prerequisites and setup basics
## Products
Choose the API that matches your workflow.
* [Extraction](/docs/scraper-api/guides/extraction/01_understanding_extraction) — extract Markdown or HTML from a webpage
* [Batch](/docs/scraper-api/guides/batch/00_understanding_batch) — process multiple URLs in one job
* [Crawl](/docs/scraper-api/guides/crawl/00_overview) — discover and extract pages across a website
* [Map](/docs/scraper-api/guides/map/00_understanding_map) — discover URLs under a base URL
* [Search](/docs/scraper-api/guides/search/01_search_overview) — submit search queries and retrieve results
## Dashboard Guides
Use the Geonode Dashboard when you want to run jobs without writing code.
* [Dashboard Overview](/docs/scraper-api/dashboard-guides/overview) — navigate Scraper, Map, and Search in the UI
## MCP
You can also use Geonode through MCP to connect the Scraper API to AI assistants and IDEs. See the [MCP Guides](/docs/scraper-api/guides/mcp/00_overview) to get started.
## API Reference
Use the API reference when you need endpoint details, parameters, and responses.
* [API Overview](/docs/scraper-api/v1) — Scraper API reference entry point
* [Extraction](/docs/scraper-api/v1/extraction/extract-content) — extract content endpoints
* [Batch](/docs/scraper-api/v1/batch/start-batch-job) — batch job endpoints
* [Crawl](/docs/scraper-api/v1/crawl/start-crawl-job) — crawl job endpoints
* [Map](/docs/scraper-api/v1/map/map-urls) — map job endpoints
* [Search](/docs/scraper-api/v1/search/start-search-job) — search job endpoints
## Next Steps
* Start with the [Quick Start Guide](/docs/scraper-api/quick-start) to authenticate
* Pick a product guide above for your use case
# Quick Start Guide (/docs/scraper-api/quick-start)
This guide helps you get started with the Geonode Scraper API in a few minutes.
## Get Your API Key
1. Sign in to your Geonode account.
2. Open the Dashboard.
3. Navigate to the API Keys section.
4. Create or copy an existing API key.
## Authentication
All Scraper API requests require the `X-Api-Key` header.
```bash title="request.sh"
curl -H "X-Api-Key: YOUR_API_KEY"
```
Replace `YOUR_API_KEY` with your actual API key.
## Choose an API
| API | Use When |
| ---------- | ------------------------------------------------------- |
| Extraction | You want to extract content from one or more webpages |
| Batch | You have multiple URLs to process in a single job |
| Crawl | You want to discover and extract pages across a website |
| Webhooks | You want notifications when jobs complete |
## Next Steps
Choose the API that matches your use case and follow the guides in that section.
# Active Subscriptions (/docs/getting-started/account-setting-and-billing/active-subscriptions)
## Step 1 — Open Profile Settings
1. Log in to your Geonode account.
2. Go to **Profile Settings** in the left-hand sidebar.
***
## Step 2 — View your active subscription
1. In the Profile Settings page, select **Active Subscriptions**.
2. You’ll see details of your current plan, including billing cycle and renewal date.
***
## Tips
* Keep your subscription active to maintain uninterrupted access.
* Review renewal dates regularly to avoid service pauses.
* To upgrade or cancel, use the **Billing** section in your dashboard.
* For any payment or renewal issues, contact [Geonode Support](https://geonode.com/contact).
Managing subscriptions in Geonode is quick and transparent — everything you need is in your profile settings.
# Payment Methods (/docs/getting-started/account-setting-and-billing/payments-method)
Geonode lets you add, update, and manage your payment methods for smooth and secure billing.
## Step 1 — Access the Payments section
1. Log in to your Geonode account.
2. From the left sidebar, open **Payments**.
***
## Step 2 — Add a new payment method
1. Click **Add Payment Method**.
2. A pop-up window will appear for card details.
### Required information
* **Card Number** (Visa, Mastercard, or AMEX)
* **Expiration Date** (MM/YY)
* **Security Code (CVC)** — 3-digit code on the back of your card
* **Billing Country**
3. Click **Add** to save your payment method.
***
## Step 3 — Update billing information
1. In the **Billing Information** section, click **Update Information**.
2. Update fields such as:
* Name
* Company (optional)
* Billing address
* Phone number
3. Click **Save Changes** to confirm updates.
***
## Step 4 — View payment history
In the **Payment History** section, you can track:
* Transaction date
* Transaction details
* Amount charged
* Invoice records
If no payments have been made, you’ll see *No Data Yet*.
***
## Tips
* Keep billing details accurate to avoid failed payments.
* Supported payment types: **Visa**, **Mastercard**, **AMEX**.
* For payment or billing issues, contact [Geonode Support](https://geonode.com/contact).
Managing payment methods in Geonode ensures secure and uninterrupted service access.
# Profile Settings (/docs/getting-started/account-setting-and-billing/profile-setting)
Your profile settings let you manage personal information, change passwords, and, if needed, delete your account.
## Step 1 — Access profile settings
1. Log in to your Geonode account.
2. Open **Profile** from the left-hand sidebar.
3. View detailed profile information.
***
## Step 2 — Update personal details
In **Personal Details**, you can:
* Edit your first and last name.
* Update your phone number.
* Note: your email address cannot be changed.
Click **Save changes** after updating.
***
## Step 3 — Change your password
1. Scroll to the **Password** section.
2. Enter and confirm a new password.
3. Click **Update password**.
***
## Step 4 — Delete your account
If you wish to permanently remove your account:
1. Scroll to **Delete Account**.
2. Click **Delete**.
3. Confirm the action — all data will be removed.
Deleting your account does **not** automatically cancel any active services.
Cancel subscriptions or contact Geonode Support before deletion to stop future
charges.
***
## Final tips
* Keep profile information current to avoid issues.
* Use a strong password for better security.
* Account deletion is irreversible.
* For billing or payment concerns, contact [Geonode Support](https://geonode.com/contact).
Managing your profile properly helps maintain a smooth, secure experience in Geonode.
# Referral Program (/docs/getting-started/account-setting-and-billing/referral-program)
The Geonode Referral Program lets you earn a 10% commission every time a new user signs up and makes a purchase through your unique affiliate link.
## How to get your referral link
Follow these steps to access and start sharing your link.
### Step 1 — Open Profile Settings
1. Log in to your Geonode account.
2. Go to **Profile Settings** in the left sidebar.
***
### Step 2 — Copy your referral link
1. In the **Profile Settings** page, find the **Referral Program** section.
2. Copy your unique affiliate link.
***
## How the referral program works
* Share your referral link anywhere: social media, websites, or direct messages.
* When a user registers and makes a purchase using your link, you earn **10% commission**.
* Commissions are tracked and displayed in your referral dashboard.
***
## Tips
* Share your link across different platforms for better reach.
* More referrals mean higher earnings.
* Monitor your performance and payouts in the referral dashboard.
* For any payment or tracking issues, contact [Geonode Support](https://geonode.com/contact).
Start earning today with Geonode’s Referral Program.
# Reset API Password (/docs/getting-started/account-setting-and-billing/reset-api-password)
If you need to reset your API password, follow these steps to generate a new one securely.\
Resetting your API password immediately deactivates the old one — don’t forget to update your integrations afterward.
## Step 1 — Access API credentials
1. Log in to your Geonode account.
2. Open the **Proxy Configuration** tab in the dashboard.
3. Locate the **API Credentials** section.
***
## Step 2 — Generate a new API password
1. Click the **reset icon** next to your current API password.
2. A confirmation prompt will appear.
3. Click **Confirm** to generate a new password.
***
## Tips
* Resetting your API password deactivates the old one instantly.
* Update all connected tools, scripts, and automations with the new password.
* You can reset it anytime if you forget or lose access.
* For any billing or payment issues, contact [Geonode Support](https://geonode.com/contact).
Managing your API credentials properly ensures secure and uninterrupted access to Geonode services.
# Wallet (/docs/getting-started/account-setting-and-billing/wallet-ocerview)
The Geonode Wallet lets you view your balance, review transactions, and add funds for seamless proxy usage.
## Step 1 — Access the Wallet
1. Log in to your Geonode account.
2. Open **Wallet** from the left sidebar.
***
## Step 2 — Check your balance
* Your current wallet balance appears at the top.
* Previous transactions are shown under **Transactions from wallet**.
If you haven’t made any payments yet, the transaction list will be empty.
***
## Step 3 — Add funds
1. In the **Add Funds** section, enter the desired amount.
2. Click **Top Up Wallet** to start the payment process.
3. Once the payment is confirmed, the amount will appear in your wallet balance.
***
## Tips
* Keep your wallet funded to prevent service interruptions.
* Review your transaction history regularly for clarity and tracking.
* For any billing or payment issues, contact [Geonode Support](https://geonode.com/contact).
Using the Geonode Wallet makes managing payments simple and ensures uninterrupted access to all services.
# Billing & Payments (/docs/getting-started/faqs/billing-and-payments)
***
{" "}
All invoices are in USD, we are unable to offer other currencies.
{" "}
If you fail to update your payment information before the next billing cycle,
you may encounter service interruptions.
{" "}
After completing the payment update process, your next billing cycle will be
automatically processed.
{" "}
To view your billing history and invoices: 1. Click on the **Profile** tab. 2.
Navigate to **Account Settings**. 3. Go to the **Payment and Wallet** section.
4\. You will find your **billing history** and invoices listed there.
{" "}
Yes, you can add your credit card details through the Billing tab.
{" "}
Yes, but you must ensure that your credit card is not set as the default card.
If you need extra support, our team is always available to help.
# General FAQs (/docs/getting-started/faqs/general-faqs)
***
{" "}
We do not currently offer static IPs. We are planning to release this service
in the coming months!
{" "}
No. Any hacking/cracking or illegal activity is strictly forbidden on our
platform.
{" "}
We do not recommend our current services for streaming, you are welcome to try
and use them for this purpose.
{" "}
We do not block any websites, including survey sites.
# Subscriptions & Cancellations (/docs/getting-started/faqs/subscriptions-and-cancellations)
***
{" "}
Yes, you can cancel anytime. You can still use our service until the end of
the billing period if you have unused bandwidth.
{" "}
Unfortunately, we can't extend subscriptions as we will encounter a break in
the system. However, we will find the right solution for you to compensate for
the downtime, so please reach out to our support team for manual assistance.
{" "}
Yes, you can add multiple cards, but only one will be set as the default card.
{" "}
You can easily re-activate your subscription through your dashboard or reach
out to our support team for help.
# Technical Issues & Support (/docs/getting-started/faqs/technical-issues-and-support)
***
{" "}
Our team is working on resolving any occurring issues as soon as possible. We
will inform our users once everything is functioning normally.
{" "}
If your issue is not listed or you need further assistance, please [contact
our support team](https://geonode.com/contact). We are available 24/7 to help
you with any issues.
# Proxy Related Queries (/docs/proxies/additional-resources/faqs)
***
{" "}
Sticky proxies have a maximum timeout of 60 minutes. Static proxies are the
type of proxies that do not expire, which we do not currently offer. We
recommend setting auto-replace to true and choosing a lengthy enough rotating
interval for sticky connections when clients bring up the offline issue.
{" "}
We do not recommend our current services for streaming, but you are welcome to
try.
{" "}
We do not block any websites, including survey sites.
You will see popups everywhere asking for your proxy username and password.
{" "}
You can whitelist **up to 150 IP addresses**.
{" "}
A proxy might slightly affect speed, but Geonode's high-speed servers minimize
this impact.
{" "}
Yes, but you must configure it separately for each network type.
{" "}
There are no restrictions on the number of simultaneous connections.
# Core Features of Geonode (/docs/getting-started/what-is-geonode/core-features-of-geonode)
When it comes to managing your online activity securely and efficiently, Geonode provides the tools you need.\
Here’s what it offers for both individual users and businesses.
***
## 1. Proxy products
Websites constantly track online activity — Geonode’s proxy network gives you control again.
With Geonode, you can:
* Browse anonymously — hide your IP and location to stay private
* Access blocked content — bypass geo-restrictions and visit region-limited websites
Geonode acts as a secure, private gateway to the web.
***
## 2. Scraper tools and data collection
Collecting data from multiple sources often leads to blocks. Geonode prevents that by rotating IPs automatically.
* Collect data efficiently — extract information without interruptions
* Avoid blocks — stay under detection limits through IP rotation
Ideal for businesses, researchers, and anyone working with large-scale web data.
***
## 3. Advanced tools
Managing multiple accounts or troubleshooting network issues can be complex.\
Geonode’s advanced tools simplify these tasks.
* Manage multiple accounts — operate several profiles safely at once
* Detect and fix issues quickly — identify and resolve connection problems easily
These tools are designed for both beginners and professionals.
***
## 4. Global reach
Geonode keeps you connected anywhere — from Singapore to New York or Tokyo.
* Access global content — open region-specific websites and services
* Choose IPs worldwide — select IPs from diverse countries for testing, marketing, or research
With Geonode, you don’t just use the internet — you control how you connect to it.
# Use Cases and Applications (/docs/getting-started/what-is-geonode/uses-cases-and-applications)
Geonode isn’t just a proxy provider — it’s a flexible tool for diverse online workflows.\
Here are a few key ways you can use it effectively:
***
## Web scraping and data collection
Extract valuable data from websites for research, business analytics, and market insights.\
Geonode ensures stable access and prevents IP bans, so your data collection runs smoothly.
***
## Digital marketing and SEO
Analyze competitor strategies, monitor performance, and gather location-specific results.\
Geonode helps marketers and SEO specialists test campaigns and view search results from any region.
***
## Cybersecurity and network testing
Test firewalls, identify vulnerabilities, and verify network configurations securely.\
Geonode assists in assessing your infrastructure without exposing your real IP or internal endpoints.
***
These are only a few ways Geonode supports technical and business workflows.\
[Explore more use cases →](https://geonode.com/use-cases)
# What is Geonode? (/docs/getting-started/what-is-geonode/what-is-geonode)
Geonode is a proxy service that helps you bypass internet restrictions by masking your real IP address. It makes your online activity private, secure, and unrestricted.
***
## How Geonode works
Geonode provides scalable proxy solutions to keep your connection stable and private.
Main proxy types:
* Residential proxies — real IPs from real devices for undetectable browsing
* Datacenter proxies — high-speed, cost-efficient options for large data operations
***
## Problems Geonode solves
Common challenges users face online:
* Access and data collection — scrape or gather data without getting blocked
* Privacy and security — protect sensitive information while browsing or automating tasks
***
## Why choose Geonode
Geonode stands out through reliability and flexibility.
* Unlimited proxy endpoints — generate and manage as many as needed
* Simple dashboard — intuitive design for fast configuration
* Global coverage — millions of IPs across worldwide locations
***
## Who should use this guide
This guide is suitable for all users — beginners and experienced alike.\
It explains how to use Geonode effectively to enhance privacy, access, and productivity online.
# Retrieve Available Geo-locations (/docs/proxies/api-reference/available-geo-locations)
Retrieve a list of available geo-locations, including countries, cities, states, and ASNs that you can use for geo-targeting your proxy requests.
## Supported Service Types
This endpoint supports the following proxy services:
| Service Type | Description |
| ----------------------- | ----------------------------- |
| `RESIDENTIAL-PREMIUM` | Premium residential proxies |
| `ROTATING-DATACENTER` | Rotating datacenter proxies |
| `UNLIMITED-RESIDENTIAL` | Unlimited residential proxies |
Replace the service type in the request URL with the one you want to query.
## Request
### Premium Residential
```bash
curl -X GET "https://monitor.geonode.com/services/RESIDENTIAL-PREMIUM/targeting-options" \
-u geonode_username:password
```
### Rotating Datacenter
```bash
curl -X GET "https://monitor.geonode.com/services/ROTATING-DATACENTER/targeting-options" \
-u geonode_username:password
```
### Unlimited Residential
```bash
curl -X GET "https://monitor.geonode.com/services/UNLIMITED-RESIDENTIAL/targeting-options" \
-u geonode_username:password
```
## Response
### 200 Success
Returns the available geo-targeting options for the selected proxy service.
### Response Structure
The response is an array of country objects.
| Field | Type | Description |
| ---------------- | ------ | -------------------------------------------------------- |
| `code` | string | ISO 3166-1 alpha-2 country code |
| `name` | string | Full country name |
| `cities` | object | Available city-level targeting options |
| `cities.prefix` | string | Prefix used for city targeting (for example, `-city-`) |
| `cities.options` | array | Available cities |
| `states` | object | Available state-level targeting options |
| `states.prefix` | string | Prefix used for state targeting (for example, `-state-`) |
| `states.options` | array | Available states |
| `asns` | object | Available ASN-level targeting options |
| `asns.prefix` | string | Prefix used for ASN targeting (for example, `-asn-`) |
| `asns.options` | array | Available ASNs |
### Example Response
```json
[
{
"code": "AF",
"name": "Afghanistan",
"cities": {
"prefix": "-city-",
"options": [
{
"code": "kabul",
"name": "Kabul"
}
]
},
"states": {
"prefix": "-state-",
"options": [
{
"code": "kabul",
"name": "Kabul"
}
]
},
"asns": {
"prefix": "-asn-",
"options": [
{
"code": "131284",
"name": "AS131284 Etisalat Afghan"
}
]
}
}
]
```
# Error Handling (/docs/proxies/api-reference/error-handling)
When using the Geonode Proxy API, you may encounter various HTTP status codes. Understanding these codes helps you handle errors gracefully and implement robust error handling in your applications.
## HTTP Status Codes
| Status Code | Description | Expanded Description |
| ----------- | -------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| 403 | Invalid request configuration. . | This error occurs when the request contains invalid or improperly formatted parameters. The placeholder will include details about which field caused the issue. Common causes include incorrect formatting or unsupported values for fields like country, city, state, or ISP. Please verify that all parameters follow the expected structure and use values listed in the API documentation or supported geo-targeting options. |
| 407 | Authentication error. Please check your authentication settings. | This error indicates that authentication with the proxy server failed. Authentication is required, but the credentials provided were missing, invalid, or insufficient. Common causes include incorrect or missing proxy username/password, misconfigured authentication headers, or using an IP address that hasn't been whitelisted. Please verify that your credentials are correct and that your IP address is authorized to access the proxy if applicable. Refer to the authentication section of the API documentation for setup instructions and troubleshooting tips. |
| 411 | Your account has been blocked. If you think it is a mistake, please contact customer support. | This error means that the user's account has been manually blocked by the system or an administrator. This action is typically taken due to violations of the terms of service, suspicious activity, billing issues, or abuse prevention measures. If you believe this is an error, please contact customer support to review your account status and resolve the issue. |
| 464 | Connection to the specified target is not permitted due to security policies. | This error occurs when the requested connection to a specific host, IP address, or port is blocked due to security or access control policies. Common reasons include attempting to connect to restricted destinations, using unsupported ports, or using protocols that are not allowed (e.g., FTP or SMTP). Ensure that the target address, port, and protocol are supported and comply with the platform's usage policies. If you're unsure or believe the restriction is incorrect, please contact support for clarification. |
| 465 | No proxies available in the selected location. Please choose a different targeting configuration or try again later. | This error indicates that the proxy server was unable to find any available IP addresses that meet the specific geo-targeting requirements (e.g., country, city, ISP) specified in the request. This can occur when demand exceeds supply in a particular region or when targeting criteria are too narrow. To resolve this issue, try adjusting your targeting configuration to be less restrictive or retry your request later. If consistent access to specific regions is critical, please contact support to explore custom proxy allocations or availability options. |
| 466 | You've reached your bandwidth limit. Buy more data or upgrade your plan to continue. | This error indicates that the user has consumed all of the bandwidth allocated in their current plan. When the bandwidth limit is reached, no further requests can be processed until more data is added. To resume usage, you can upgrade your plan or purchase an add-on directly in the user dashboard. |
| 467 | You have reached your session bandwidth limit. | Specified bandwidth limit for this session has been reached. Limit is defined by -limit- parameter of the first session request and can not be changed until session is expired or released. |
| 468 | Target Forbidden. | The proxy refused this destination because it matches an entry on your customer [Block list](/docs/proxies/guides/residential-proxies/block-list) (domain, IP, or wildcard). Remove the entry if you need the proxy to fetch that host again. This is different from **464**, which is a platform security policy rather than your own Block list. |
| 500 | Internal error. Please try again. | This error indicates that something went wrong on the server while processing the request. It is not caused by user input or configuration. Please try again. If the issue persists, contact support with your request details so we can investigate. |
| 517 | Internal error. Please try again. | This error indicates that something went wrong on the server while processing the request. It is not caused by user input or configuration. Please try again. If the issue persists, contact support with your request details so we can investigate. |
| 518 | Internal error. Please try again. | This error indicates that something went wrong on the server while processing the request. It is not caused by user input or configuration. Please try again. If the issue persists, contact support with your request details so we can investigate. |
| 560 | Internal error. Please try again. | This error indicates that something went wrong on the server while processing the request. It is not caused by user input or configuration. Please try again. If the issue persists, contact support with your request details so we can investigate. |
| 561 | Internal error. Please try again. | This error indicates that something went wrong on the server while processing the request. It is not caused by user input or configuration. Please try again. If the issue persists, contact support with your request details so we can investigate. |
| 562 | Internal error. Please try again. | This error indicates that something went wrong on the server while processing the request. It is not caused by user input or configuration. Please try again. If the issue persists, contact support with your request details so we can investigate. |
| 563 | Internal error. Please try again. | This error indicates that something went wrong on the server while processing the request. It is not caused by user input or configuration. Please try again. If the issue persists, contact support with your request details so we can investigate. |
| 564 | Internal error. Please try again. | This error indicates that something went wrong on the server while processing the request. It is not caused by user input or configuration. Please try again. If the issue persists, contact support with your request details so we can investigate. |
| 565 | Internal error. Please try again. | This error indicates that something went wrong on the server while processing the request. It is not caused by user input or configuration. Please try again. If the issue persists, contact support with your request details so we can investigate. |
| 566 | Internal error. Please try again. | This error indicates that something went wrong on the server while processing the request. It is not caused by user input or configuration. Please try again. If the issue persists, contact support with your request details so we can investigate. |
| 567 | Internal error. Please try again. | This error indicates that something went wrong on the server while processing the request. It is not caused by user input or configuration. Please try again. If the issue persists, contact support with your request details so we can investigate. |
| 569 | Internal error. Please try again. | This error indicates that something went wrong on the server while processing the request. It is not caused by user input or configuration. Please try again. If the issue persists, contact support with your request details so we can investigate. |
## Best Practices
1. **Always Implement Error Handling**
* Never assume requests will succeed
* Handle both expected and unexpected errors
2. **Use Retry Logic**
* Implement exponential backoff for 5xx errors
* Respect rate limits and Retry-After headers
3. **Log Errors Appropriately**
* Include relevant request details
* Don't log sensitive information
4. **User Feedback**
* Provide clear error messages to end users
* Include actionable steps for resolution
## Support
If you encounter persistent errors or need assistance, contact our support team at [Support](https://geonode.com/contact).
***
# Introduction (/docs/proxies/api-reference)
Welcome to the Geonode Proxy API! This API allows you to:
* Target specific geolocations (countries, states, cities, and even ISPs).
* Manage sticky sessions to keep a consistent IP across multiple requests.
* Track usage statistics (e.g., bandwidth, session counts).
* Filter proxy types (residential, data center, mobile).
This guide will help you make your first calls, retrieve essential information, and move on to more advanced features—without overwhelming you.
## Prerequisites:
* API Credentials: **Our API uses Basic Authentication, requiring a proxy username and password for access. You must include these credentials in every request using the Authorization header.** *(Base64-encoded string)*
* Service Name: Indicate which specific service plan or tier you're using.
> Note: *Some sections provide a static cURL example rather than a live, testable endpoint—so you won't be able to send requests directly from this page. If you want to try it out, simply copy the cURL command into your terminal (or HTTP client) and insert your real credentials or parameters.*
Please refer to our [user dashboard](https://app.geonode.com/) to find the information mentioned above.
## Basic Workflow:
* Authenticate: Include your proxy username and password in each request using Basic Authentication (via the Authorization header).
* Specify Your Service: Use the appropriate service name in each call so the API applies the correct proxy settings.
* Send Requests: Interact using standard HTTP methods (GET, POST, PUT, etc.).
* Parse JSON Responses: You'll generally receive JSON objects containing the requested data or error details.
For detailed information about handling API errors and implementing robust error handling, please refer to our [Error Handling Guide](/docs/proxies/api-reference/error-handling).
***
## Next Steps:
* Explore the Full Reference: Dive deeper to learn how to configure proxy sessions, target specific regions, or gather usage insights.
* Check Plan Limits: Monitor bandwidth and session counts to stay within your plan's capacity.
* Stay Informed: Check our Changelog for updates & and new features. Share your opinion on feature requests.
* Contact Support: If you have any questions or run into issues, reach out to our support team at [hello@geonode.com](mailto:hello@geonode.com).
***
# IP Type Filtering (/docs/proxies/api-reference/ip-filter)
Filter proxy IPs by type to match your specific requirements. You can choose between residential IPs, datacenter IPs, or a mix of both.
You can filter IPs by type using the `-type-` parameter in your username:
* **Residential**: Real IPs from home users
* **Datacenter**: IPs from data centers
* **Mixed**: Combination of both types
## Request
```bash
curl -x "http://proxy.geonode.io:" \
--user "-type--country-:" \
--url "http://ip-api.com/json" \
--header "Accept: application/json"
```
## Response
### 200 Success
Successfully filtered IPs based on type.
#### Response Fields
| Field | Type | Description |
| --------------- | ------- | ------------------------------------------------------ |
| `status` | string | The status of the request (e.g., "success") |
| `continent` | string | The continent where the IP is located |
| `continentCode` | string | The continent code |
| `country` | string | The country where the IP is registered |
| `countryCode` | string | The country code in ISO 3166-1 alpha-2 format |
| `region` | string | The regional subdivision (state/province) |
| `regionName` | string | The full name of the region |
| `city` | string | The city associated with the IP address |
| `district` | string | The district or subdivision of the city |
| `zip` | string | The postal or ZIP code of the location |
| `lat` | number | Latitude coordinate of the location |
| `lon` | number | Longitude coordinate of the location |
| `timezone` | string | Time zone in which the IP is located |
| `offset` | integer | Time offset from UTC in seconds |
| `currency` | string | Local currency used in the country |
| `isp` | string | The name of the Internet Service Provider (ISP) |
| `org` | string | The name of the organization associated with the IP |
| `as` | string | The Autonomous System (AS) number and name |
| `asname` | string | The full Autonomous System (AS) name |
| `mobile` | boolean | Indicates whether the IP is from a mobile network |
| `proxy` | boolean | Indicates whether the IP is being used as a proxy |
| `hosting` | boolean | Indicates whether the IP belongs to a hosting provider |
| `query` | string | The IP address queried in the request |
# Retrieve Usage Statistics (/docs/proxies/api-reference/usage-statistics)
Retrieve bandwidth usage statistics for your Geonode proxy service account. This endpoint provides detailed information about your data consumption.
## Request
```bash
curl -X GET "https://monitor.geonode.com/monitor-light/proxies" \
-H "Authorization: Basic base64(username:password)"
```
## Response
### 200 Success
Usage statistics retrieved successfully.
#### Response Fields
| Field | Type | Description |
| ------------------------------------------ | ------- | ------------------------------------------------------ |
| `data` | object | Container for usage statistics |
| `data.bandwidth` | object | Bandwidth usage information |
| `data.bandwidth.data` | object | Detailed bandwidth data |
| `data.bandwidth.data.default` | integer | Total bandwidth used in bytes |
| `data.bandwidth.data.currentFastBandwidth` | integer | Current fast bandwidth usage (unlimited services only) |
| `data.bandwidth.data.totalBandwidthInGB` | integer | Total bandwidth used in gigabytes |
#### Example Response
```json
{
"data": {
"bandwidth": {
"data": {
"default": 17930356,
"currentFastBandwidth": 0,
"totalBandwidthInGB": 2
}
}
}
}
```
# Quick Start Guide (/docs/proxies/getting-started/quick-start)
import BrowsersFaqs from "../../../snippets/browsers-faqs.mdx";
import ProxyInfoFromGeonode from "../../../snippets/get-proxy-info-from-geonode.mdx";
import SupportParagraph from "../../../snippets/support-paragraph.mdx";
***
## 1. Access the User Dashboard
Start by accessing your user dashboard — this is where you manage all proxy settings.
* Open your Dashboard.
* Scroll down to Proxy Configuration.
***
## 2. Get Proxy Credentials from Geonode
Obtain your authentication credentials directly from the dashboard.
* **Username:** Copy your unique API username.
* **Password:** Copy your API password.
***
## 3. Configure Proxy Parameters
Within the Proxy Configuration section, adjust the parameters as needed.
| Parameter | Example | Description |
| ------------------- | ------------------------------------------------ | --------------------------------------------------------- |
| **Endpoint Format** | `hostname:port:username:password` | Connection string format |
| **IP Type** | `-type-datacenter` | Choose between Datacenter or Residential |
| **Gateway** | `192.155.103.209` | Proxy gateway |
| **Geo-Targeting** | `-country-jp`, `-state-tokyo`, `-as-2501` | Specify region or ASN (state and city cannot be combined) |
| **Protocol** | HTTP / HTTPS | Default port 9000 |
| **Session Type** | — | Rotate or Sticky sessions |
| **Output Format** | — | Customize response format |
See the full guide:\
[How to use the Endpoint Generator to configure proxy parameters](/docs/proxies/getting-started/knowledge-base/geo-targeting)
***
## 4. Make Your First API Call
Once your endpoint is configured, generate your first API request.
1. Go to **API Code Generator** in the dashboard.
2. Select your configuration settings.
3. Copy the generated code snippet.
***
## 5. Configure Proxy in Different Browsers or OS
To integrate your proxy on specific platforms, follow one of the setup guides below.
### Desktop Operating Systems
* [Windows 10/11](/docs/proxies/getting-started/setup_and_configuration/desktop-OS/windows)
* [macOS](/docs/proxies/getting-started/setup_and_configuration/desktop-OS/macOS)
### Mobile Operating Systems
* [Android](/docs/proxies/getting-started/setup_and_configuration/mobile-OS/android)
* [iOS](/docs/proxies/getting-started/setup_and_configuration/mobile-OS/ios)
### Browsers
* [Chrome](/docs/proxies/getting-started/setup_and_configuration/browsers/chrome)
* [Edge](/docs/proxies/getting-started/setup_and_configuration/browsers/edge)
* [Brave](/docs/proxies/getting-started/setup_and_configuration/browsers/brave)
* [Firefox](/docs/proxies/getting-started/setup_and_configuration/browsers/firefox)
* [Safari](/docs/proxies/getting-started/setup_and_configuration/browsers/safari)
* [Incognito Mode](/docs/proxies/getting-started/setup_and_configuration/browsers/incognito-mode)
* [Incogniton](/docs/proxies/getting-started/setup_and_configuration/browsers/incogniton)
* [AdsPower](/docs/proxies/getting-started/setup_and_configuration/browsers/adspower)
* [Ghost Browser](/docs/proxies/getting-started/setup_and_configuration/browsers/ghost)
* [GoLogin](/docs/proxies/getting-started/setup_and_configuration/browsers/gologin)
* [Dolphin Anty](/docs/proxies/getting-started/setup_and_configuration/browsers/dolphin-anty)
* [ClonBrowser](/docs/proxies/getting-started/setup_and_configuration/browsers/clonbrowser)
* [MoreLogin](/docs/proxies/getting-started/setup_and_configuration/browsers/morelogin)
* [MultiLogin](/docs/proxies/getting-started/setup_and_configuration/browsers/multilogin)
### Extensions
* [FoxyProxy](/docs/proxies/getting-started/setup_and_configuration/extensions/foxyproxy)
* [Geonode Proxy Manager](/docs/proxies/getting-started/setup_and_configuration/extensions/geonode-proxy-manager)
***
## Conclusion
You’re all set!\
Your Geonode proxy is now configured and ready for use across browsers, operating systems, and API integrations.
***
***
# Quick Start Guide (/docs/proxies/guides)
import BrowsersFaqs from "../../../snippets/browsers-faqs.mdx";
import ProxyInfoFromGeonode from "../../../snippets/get-proxy-info-from-geonode.mdx";
import SupportParagraph from "../../../snippets/support-paragraph.mdx";
***
## 1. Access the User Dashboard
Start by accessing your user dashboard — this is where you manage all proxy settings.
* Open your Dashboard.
* Scroll down to Proxy Configuration.
***
## 2. Get Proxy Credentials from Geonode
Obtain your authentication credentials directly from the dashboard.
* **Username:** Copy your unique API username.
* **Password:** Copy your API password.
***
## 3. Configure Proxy Parameters
Within the Proxy Configuration section, adjust the parameters as needed.
| Parameter | Example | Description |
| ------------------- | ------------------------------------------------ | --------------------------------------------------------- |
| **Endpoint Format** | `hostname:port:username:password` | Connection string format |
| **IP Type** | `-type-datacenter` | Choose between Datacenter or Residential |
| **Gateway** | `192.155.103.209` | Proxy gateway |
| **Geo-Targeting** | `-country-jp`, `-state-tokyo`, `-as-2501` | Specify region or ASN (state and city cannot be combined) |
| **Protocol** | HTTP / HTTPS | Default port 9000 |
| **Session Type** | — | Rotate or Sticky sessions |
| **Output Format** | — | Customize response format |
See the full guide:\
[How to use the Endpoint Generator to configure proxy parameters](/docs/proxies/getting-started/knowledge-base/geo-targeting)
***
## 4. Make Your First API Call
Once your endpoint is configured, generate your first API request.
1. Go to **API Code Generator** in the dashboard.
2. Select your configuration settings.
3. Copy the generated code snippet.
***
## 5. Configure Proxy in Different Browsers or OS
To integrate your proxy on specific platforms, follow one of the setup guides below.
### Desktop Operating Systems
* [Windows 10/11](/docs/proxies/getting-started/setup_and_configuration/desktop-OS/windows)
* [macOS](/docs/proxies/getting-started/setup_and_configuration/desktop-OS/macOS)
### Mobile Operating Systems
* [Android](/docs/proxies/getting-started/setup_and_configuration/mobile-OS/android)
* [iOS](/docs/proxies/getting-started/setup_and_configuration/mobile-OS/ios)
### Browsers
* [Chrome](/docs/proxies/getting-started/setup_and_configuration/browsers/chrome)
* [Edge](/docs/proxies/getting-started/setup_and_configuration/browsers/edge)
* [Brave](/docs/proxies/getting-started/setup_and_configuration/browsers/brave)
* [Firefox](/docs/proxies/getting-started/setup_and_configuration/browsers/firefox)
* [Safari](/docs/proxies/getting-started/setup_and_configuration/browsers/safari)
* [Incognito Mode](/docs/proxies/getting-started/setup_and_configuration/browsers/incognito-mode)
* [Incogniton](/docs/proxies/getting-started/setup_and_configuration/browsers/incogniton)
* [AdsPower](/docs/proxies/getting-started/setup_and_configuration/browsers/adspower)
* [Ghost Browser](/docs/proxies/getting-started/setup_and_configuration/browsers/ghost)
* [GoLogin](/docs/proxies/getting-started/setup_and_configuration/browsers/gologin)
* [Dolphin Anty](/docs/proxies/getting-started/setup_and_configuration/browsers/dolphin-anty)
* [ClonBrowser](/docs/proxies/getting-started/setup_and_configuration/browsers/clonbrowser)
* [MoreLogin](/docs/proxies/getting-started/setup_and_configuration/browsers/morelogin)
* [MultiLogin](/docs/proxies/getting-started/setup_and_configuration/browsers/multilogin)
### Extensions
* [FoxyProxy](/docs/proxies/getting-started/setup_and_configuration/extensions/foxyproxy)
* [Geonode Proxy Manager](/docs/proxies/getting-started/setup_and_configuration/extensions/geonode-proxy-manager)
***
## Conclusion
You’re all set!\
Your Geonode proxy is now configured and ready for use across browsers, operating systems, and API integrations.
***
***
# Choosing the Right Scraper API Plan (/docs/scraper-api/additional-resources/choosing_scraper_api_plan)
import Link from "next/link";
Geonode Scraper API offers two pricing models:
* **Request-Based** pricing, where successful page extractions use requests from your available balance.
* **Unlimited** plans, where you pay a flat monthly price and your plan determines the number of concurrent threads.
This guide helps you decide which model is better for your workload.
## Quick Comparison
| | Request-Based | Unlimited |
| ------------------- | ------------------------------------ | ------------------------- |
| Pricing | Based on request volume | Flat monthly price |
| Monthly requests | Plan-dependent | Unlimited |
| Main pricing factor | Successful extractions | Concurrent threads |
| Cost | Based on your request plan and usage | Predictable monthly price |
| Best for | Variable or smaller workloads | High-volume workloads |
| Batch extraction | Supported | Supported |
| Crawl | Supported | Supported |
If you already know which pricing model you need, you can go directly to the detailed guide:
Request-Based
Pay based on successful page extractions and choose the request capacity that fits your workload.
View Request-Based Pricing
Unlimited
Get unlimited requests with a fixed monthly price based on concurrent threads.
View Unlimited Pricing
## Choose Request-Based Pricing If
Request-Based pricing is a good fit when your extraction volume changes from month to month or you do not need a large amount of concurrent processing.
Consider Request-Based pricing if:
* Your scraping volume is relatively small.
* Your workload is unpredictable.
* You only run extraction jobs occasionally.
* You want your pricing to be based on request usage.
* You are testing or starting a new scraping workflow.
* You do not need a high number of concurrent threads.
Request-Based pricing also includes a free tier with **1,500 requests per month**, which makes it suitable for getting started without immediately choosing a paid plan.
Explore Request-Based Pricing
## Choose Unlimited If
Unlimited plans are designed for workloads that require a larger amount of extraction capacity and predictable monthly pricing.
Consider an Unlimited plan if:
* You regularly process a large number of pages.
* You need to run many extractions in parallel.
* Your application can take advantage of higher concurrency.
* You want a predictable monthly cost.
* You regularly run large Batch or Crawl jobs.
* You do not want to manage a monthly request balance.
Unlimited plans are priced according to concurrent threads.
For example, the available plans provide different concurrency levels:
| Plan | Concurrent threads |
| -------- | -----------------: |
| Starter | 2 |
| Growth | 10 |
| Scale | 25 |
| Pro | 50 |
| Business | 100 |
When all available threads are busy, additional work waits until a thread becomes available.
Explore Unlimited Pricing
## Compare Common Workloads
Use your expected workload to determine which pricing model is more appropriate.
| Workload | Recommended |
| ----------------------------------- | ----------------- |
| Testing the Scraper API | Request-Based |
| Small or occasional extraction jobs | Request-Based |
| Variable monthly scraping volume | Request-Based |
| Regular production scraping | Depends on volume |
| Large Batch jobs | Unlimited |
| Large Crawl jobs | Unlimited |
| High-volume scraping | Unlimited |
| Many parallel extraction tasks | Unlimited |
| Predictable monthly scraping costs | Unlimited |
These recommendations are based on the difference between the two pricing models. Your actual choice should depend on both your request volume and the amount of parallel processing your application requires.
## Consider Concurrency
Concurrency is particularly important when comparing plans.
A higher concurrency level allows more extraction tasks to run at the same time.
For example:
| Concurrent threads | Parallel extraction capacity |
| -----------------: | ---------------------------: |
| 2 | Up to 2 extractions |
| 10 | Up to 10 extractions |
| 25 | Up to 25 extractions |
| 50 | Up to 50 extractions |
| 100 | Up to 100 extractions |
If your application processes URLs sequentially, increasing concurrency may not provide much benefit.
If your application can submit many extraction tasks at the same time, a higher-concurrency Unlimited plan can provide greater throughput.
## Consider Batch and Crawl Jobs
Batch and Crawl jobs have their own job-size limits.
These limits apply to an individual job and are separate from the monthly request allowance on request-based plans or the unlimited request volume on Unlimited plans.
For Unlimited plans, the job limits are:
| Plan | Max URLs per batch | Max pages per crawl |
| -------- | -----------------: | ------------------: |
| Starter | 50 | 50 |
| Growth | 500 | 500 |
| Scale | 1,000 | 1,000 |
| Pro | 2,000 | 2,000 |
| Business | 4,000 | 4,000 |
If a workload is larger than the maximum size of a single job, split it into multiple jobs.
## Example Scenarios
### You are testing the API
If you are testing a new extraction workflow or only need a small number of requests, start with Request-Based pricing.
The free tier provides **1,500 requests per month**.
### You scrape occasionally
If you only run extraction jobs when you need specific data and your monthly volume varies, Request-Based pricing can be a better fit.
You pay according to the request capacity you need instead of committing to a fixed Unlimited concurrency level.
### You process thousands of pages regularly
If your application regularly processes thousands of pages and can run extraction tasks in parallel, an Unlimited plan may be a better fit.
The main consideration becomes how many concurrent threads your workload needs.
### You run large Batch or Crawl jobs
If your application regularly runs large Batch or Crawl jobs, consider an Unlimited plan with enough concurrency and job capacity for your workload.
You can split larger workloads across multiple jobs when required.
## A Simple Decision
Use this as a quick starting point:
### Choose Request-Based if
* Your volume is small or unpredictable.
* You want to pay based on request usage.
* You do not need high concurrency.
* You are testing or starting a project.
### Choose Unlimited if
* Your volume is consistently high.
* You need many extractions to run in parallel.
* You want predictable monthly pricing.
* You regularly process large scraping workloads.
## You Can Change Plans as Your Workload Grows
Your initial choice does not need to be permanent.
You can start with Request-Based pricing while evaluating your workload. If your extraction volume grows or you need more parallel processing, you can move to an Unlimited plan.
Similarly, if your workload does not require high concurrency, a request-based plan may be more appropriate.
The best plan is the one that matches your actual workload rather than simply choosing the plan with the highest limits.
## Related Guides
Request-Based Pricing
Unlimited Scraper API Pricing
Extract a Single URL
Extract Multiple URLs
# FAQs (/docs/scraper-api/additional-resources/faq)
The Scraper API is a hosted extraction API. You send it a URL, and it returns clean Markdown or HTML from the target page. It also handles proxy routing, geo-targeting, JavaScript rendering, and anti-bot handling.
One credit means one Scraper API request. In these docs, you'll usually see the word request instead of credit because the dashboard and pricing model are request-based.
One successful page extraction uses one request. Batch jobs count each successfully extracted URL as one request, and crawl jobs count each successfully extracted page as one request.
There are no extra request multipliers for JavaScript rendering, proxy type, geo-targeting, or requesting both Markdown and HTML.
When you use all free requests for the billing month, new extraction requests stop working until you add paid requests, upgrade to a subscription, or wait for the next monthly renewal.
The API can return `402` when your account does not have enough available requests.
No. Free tier requests renew every billing month and expire at the end of that month. Unused free requests do not roll over.
Yes. Unused subscription requests roll over to the next month with no rollover cap while your subscription remains active.
No. Pay as you Go requests do not expire. They are a good fit when your scraping workload is occasional or bursty.
No. JavaScript rendering does not cost extra requests. A successful extraction with `render_js: true` still uses one request.
JavaScript rendering can take longer, so it is best to start with `render_js: false` and enable it only when the returned content is incomplete.
With proxies directly, you still have to write the scraper, manage retries, choose when to run a browser, parse HTML, handle noisy pages, and normalize output.
The Scraper API sits above that work. You send a URL and get extracted Markdown or HTML back. It is a better fit when you want page content quickly and do not want to maintain browser or proxy orchestration yourself.
Direct proxies are still useful when you need full control over the browser, request flow, cookies, sessions, or a custom scraper pipeline.
You can use the Scraper API with public webpages that your account is allowed to access. It works well with static pages, documentation pages, articles, product pages, and many JavaScript-rendered pages.
Complex sites can still return noisy output. Some sites include large navigation menus, tracking links, ads, images, recommendations, cookie banners, or challenge pages in the extracted content. In those cases, inspect the returned Markdown or HTML and add post-processing if your workflow needs cleaner fields.
Always follow the target site's terms, applicable laws, and your own compliance requirements.
The most common reason is that the page loads content with JavaScript after the initial HTML response. Try the same request with `render_js: true`.
```json
{
"url": "https://quotes.toscrape.com/js/",
"formats": ["markdown"],
"render_js": true
}
```
If the result is still incomplete, the site may require a user interaction, a longer wait condition, a login, or a site-specific scraping strategy.
The API extracts raw page content. Many modern sites include navigation, filters, image data, tracking links, cookie banners, recommendations, and footer content in the page itself. The API may return some of that content because it exists in the source page.
For LLM, search, or analytics workflows, it is a good idea to test a few pages from the same target site and add your own cleanup step if needed.
Yes. Pass `proxy.country` with a two-letter ISO country code.
```json
{
"proxy": {
"country": "US",
"type": "residential"
}
}
```
If you omit `proxy`, the API applies default residential proxy routing and tries to infer a useful country from the target URL when possible.
The extraction endpoint supports `residential`, `datacenter`, and `mix`.
```json
{
"proxy": {
"country": "US",
"type": "residential"
}
}
```
Yes. Use the `headers` object for headers that should be included in the target extraction request.
```json
{
"url": "https://example.com",
"formats": ["markdown"],
"headers": {
"Accept-Language": "en-US,en;q=0.9"
}
}
```
Do not send your Geonode API key inside this object. Your API key belongs in the `X-Api-Key` header sent to the Scraper API.
The public extraction schema supports custom headers, but it does not provide a full browser session management interface for logging in, clicking through flows, or maintaining user state across many pages. You can use the `headers` field to send authentication cookies or tokens with each extraction request.
If your use case requires authenticated scraping beyond what custom headers can provide, contact support so the team can recommend the right setup.
Use `sync` mode for quick pages when you want the result in the same HTTP response.
Use `async` mode for slower pages, JavaScript-heavy pages, and workflows where you want to start the extraction and poll for the result later.
The map endpoint discovers URLs under a base URL by reading sitemaps and links from the seed page.
```json
{
"url": "https://quotes.toscrape.com/",
"search": "author",
"include_subdomains": false,
"ignore_query_parameters": true
}
```
The `search` field filters URLs the API already discovered. It does not run a Google search or query an external search engine.
Open the Geonode dashboard and check your Scraper API request balance. Use the dashboard UI as the source of truth for your current balance and plan details.
Read the full Scraper API reference here:
[Open the Scraper API Reference](/docs/scraper-api/v1/reference)
# Pricing and Requests (/docs/scraper-api/additional-resources/pricing-and-requests)
Scraper API supports both **request-based pricing** and **Unlimited plans**.
Request-based plans charge per successful page extraction, while Unlimited plans use concurrency (threads) instead of request balances to determine throughput.
## Request Counting
The request model is intentionally simple for page extraction work:
* One page extraction equals one request.
* A batch job counts each successfully extracted URL as one request.
* A crawl job counts each successfully extracted page as one request.
* JavaScript rendering does not cost extra requests.
* Requesting Markdown, HTML, or both does not cost extra requests.
* Geo-targeting does not cost extra requests.
* Failed extraction requests are not charged.
* Authentication errors, validation errors, and requests rejected before extraction are not charged.
Some dashboard, API, or billing surfaces may still use older token or credit naming, such as `tokens_charged`, `estimated_tokens`, `tokens_charged_total`, or `tokens_reserved`. For Scraper API billing, read those values as charged, estimated, or reserved requests.
***
## Endpoint Usage
Not every API endpoint extracts a page. Use this table to understand how each endpoint category relates to request balance.
| Endpoint category | Request balance behavior |
| --------------------------------------------------- | ----------------------------------------------------------------------------------------------------------- |
| Single-page extraction | A successful `POST /v1/extract` extraction uses one request. |
| Async extraction | The completed extraction uses one request when the page is successfully extracted. |
| Batch extraction | Each successfully extracted URL in the batch uses one request. |
| Crawl jobs | Each successfully extracted crawled page uses one request. |
| Map | Discovers URLs only. It does not extract page content. |
| Statistics, webhooks, job lookup, and health checks | These endpoints do not extract page content. They are used to manage, inspect, or monitor Scraper API work. |
***
## Free Tier
The free tier includes **1,500 requests every month**. You do not need a credit card to start using the free tier.
Free tier requests:
* Renew every billing month.
* Expire at the end of the billing month.
* Do not roll over to the next month.
* Stop working when the monthly free allowance runs out unless you upgrade or add paid requests.
***
## Subscription Plans
Subscription plans include a monthly request allowance. Unused subscription requests roll over to the next month with no rollover cap while your subscription remains active.
If your workload requires unlimited requests, Geonode also offers **Unlimited plans**, which use concurrency-based pricing instead of request balances.
Use the Geonode dashboard to see the currently available subscription tiers, included monthly requests, and current prices. The dashboard is the source of truth for plan availability and billing details.
Subscription requests:
* Renew every billing month.
* Roll over when unused.
* Stay available while the subscription remains active.
* Can be combined with Pay as you Go requests if your account has both balances.
***
## Unlimited Plans
Unlimited plans charge a flat monthly price with no request counting. Instead of a request balance, each plan sets the number of concurrent threads—how many extractions can run in parallel at any moment.
* No request allowance and no top-ups — total monthly volume is not capped.
* No multipliers — JavaScript rendering, geo-targeting, and output format never change the price.
* Throughput scales with your plan's thread count — more threads means more pages processed in parallel.
Use the Geonode dashboard to view the available Unlimited plans and current pricing.
***
## Concurrency and Job Limits
**Concurrency (threads)** determines how many extractions your plan can run in parallel. When all threads are busy, additional work waits in a queue until a thread becomes available. Requests are not rejected because all threads are in use.
**Job-size limits** define the maximum number of URLs or pages that can be processed in a single Batch or Crawl job. These limits apply per job, not per month, so you can run as many jobs as needed.
| Plan | Concurrency | Max URLs per batch job | Max pages per crawl job |
| ---------------------- | ----------: | ---------------------: | ----------------------: |
| Starter (2 threads) | 2 | 50 | 50 |
| Growth (10 threads) | 10 | 500 | 500 |
| Scale (25 threads) | 25 | 1,000 | 1,000 |
| Pro (50 threads) | 50 | 2,000 | 2,000 |
| Business (100 threads) | 100 | 4,000 | 4,000 |
To process more URLs than a single job allows, split the work into multiple jobs.
***
## Pay as You Go
Pay as You Go lets you buy request top-ups without committing to a subscription. Pay as You Go requests are prepaid and do not expire.
Use the Geonode dashboard to see the available top-up sizes and current prices for your account.
Pay as You Go requests:
* Are prepaid.
* Do not expire.
* Can be used for bursty workloads.
* Remain available even if you do not use the API every month.
***
## What Happens When You Run Out
For request-based plans, if you run out of available requests, the API returns `402`. Add more requests from the dashboard, buy a Pay as You Go top-up, or upgrade to a subscription plan.
```json
{
"code": "PAYMENT_REQUIRED",
"message": "Insufficient request balance.",
"correlation_id": "req_...",
"retryable": false
}
```
The exact message may vary, but the response tells you that the request could not be processed because the account does not have enough available requests.
***
## Checking Usage
You can monitor your usage in the Geonode dashboard.
Depending on your plan, the dashboard shows:
* Remaining free requests for the current billing month.
* Current subscription allowance and rollover balance.
* Pay as You Go request balance.
* Thread usage and concurrency information for Unlimited plans.
***
## Billing Examples
If you extract one static page as Markdown, it uses one request.
```json
{
"url": "https://example.com",
"formats": ["markdown"],
"render_js": false
}
```
If you extract the same page as Markdown and HTML, it still uses one request.
```json
{
"url": "https://example.com",
"formats": ["markdown", "html"],
"render_js": false
}
```
If you render a JavaScript-heavy page before extracting it, it still uses one request.
```json
{
"url": "https://quotes.toscrape.com/js/",
"formats": ["markdown"],
"render_js": true
}
```
If the extraction fails, the failed extraction is not charged.
For Batch and Crawl jobs, check `token_summary.tokens_charged_total` on the job status response to see how many requests have been charged so far. `token_summary.tokens_reserved` shows the amount still reserved for queued or processing items.
# Request-Based Pricing (/docs/scraper-api/additional-resources/request_based_pricing)
The Geonode Scraper API supports two pricing models:
* Request-based pricing, where successful page extractions consume requests.
* Unlimited plans, where throughput is based on concurrent threads instead of a request balance.
This guide explains the request-based model, including the free tier, subscription request allowances, and Pay As You Go usage.
If your workload requires unlimited requests and predictable monthly pricing, see [Unlimited Scraper API Pricing](/docs/scraper-api/additional-resources/unlimited_scraper_api_pricing).
## Request-Based Plans
Request-based plans charge based on the number of successful page extractions.
| Plan | Requests | Price | Cost per 1K requests | Best for |
| ---- | ------------: | ------------: | -------------------: | ------------------------------------------ |
| Free | 1,500 / month | $0 | — | Trying the API without a paid subscription |
| 10K | 10,000 | $3.50 / month | $0.35 | Smaller recurring workloads |
| 50K | 50,000 | $13 / month | $0.26 | Regular scraping workloads |
| 250K | 250,000 | $43 / month | $0.17 | Larger scraping workloads |
| 1M | 1,000,000 | $126 / month | $0.13 | High-volume request-based workloads |
| 3M+ | 3,000,000+ | Custom | Custom | Production-scale workloads |
The Free plan renews every billing month and does not require a credit card.
For 3M+ plans, [contact sales](https://geonode.com/contact).
> Pricing and plan availability can change. Check the [Geonode Scraper API pricing page](https://geonode.com/products/scraper-api) for the latest pricing.
## How requests are counted
The request model is based on successful page extraction.
| Operation | Request usage |
| ---------------------- | ------------------------------------------------- |
| Single-page extraction | 1 request per successfully extracted page |
| Async extraction | 1 request when the page is successfully extracted |
| Batch extraction | 1 request for each successfully extracted URL |
| Crawl | 1 request for each successfully extracted page |
| Map | Does not extract page content |
| Job lookup | Does not extract page content |
| Statistics | Does not extract page content |
| Webhooks | Does not extract page content |
| Health checks | Does not extract page content |
## What does not cost additional requests?
The following options do not create additional request charges for a successful extraction:
* JavaScript Rendering
* Geo-targeting
* Markdown output
* HTML output
* Requesting multiple supported output formats
For example, extracting a page as both Markdown and HTML still uses one request.
## Failed requests
Failed extraction requests are not charged.
Requests rejected before extraction are also not charged, including:
* Authentication errors
* Validation errors
* Requests rejected before extraction starts
If an extraction does not successfully produce a page result, it does not consume a request from your request balance.
## Free Tier
The free tier includes **1,500 requests every month**.
You can start using the Scraper API without a credit card.
Free requests:
* Renew every billing month.
* Do not roll over to the next month.
* Expire at the end of the billing period.
* Stop being available when the monthly allowance is exhausted.
Once your free allowance is used, you can upgrade to a paid request-based plan or choose an Unlimited plan.
## Pay As You Go
Pay As You Go is part of the request-based pricing model.
It is useful when you need additional request capacity without moving to an Unlimited concurrency plan.
Pay As You Go requests are prepaid and can be used for additional scraping workloads.
The available top-up sizes and prices can vary. Use the Geonode dashboard to view the options currently available for your account.
Pay As You Go requests:
* Are prepaid.
* Do not expire.
* Can be used for bursty workloads.
* Can be used alongside applicable request-based balances.
## Subscription requests
Paid request-based subscriptions provide a monthly request allowance.
Unused subscription requests roll over to the next month while your subscription remains active, according to the applicable subscription terms.
Subscription requests:
* Renew every billing month.
* Can roll over when unused.
* Remain available while the subscription is active.
* Can be combined with applicable Pay As You Go balances.
Use the Geonode dashboard to view the current subscription options and available request balances.
## What happens when you run out?
If you have no available request balance, the API returns a `402` response.
```json
{
"code": "PAYMENT_REQUIRED",
"message": "Insufficient request balance.",
"correlation_id": "req_...",
"retryable": false
}
```
# Unlimited Scraper API Pricing (/docs/scraper-api/additional-resources/unlimited_scraper_api_pricing)
Geonode's Unlimited Scraper API lets you run unlimited extraction requests on a flat monthly plan.
Instead of using a monthly request balance, Unlimited plans are based on **concurrent threads**. Your thread count determines how many extractions can run at the same time.
## Unlimited Plans
Choose an Unlimited plan based on the amount of work you need to process in parallel.
| Plan | Concurrent threads | Price | Requests | Best for |
| ---------- | -----------------: | -------------: | ---------------------- | ----------------------------------- |
| Free Trial | 5 | $0 | Unlimited for 48 hours | Trying Unlimited before a paid plan |
| Starter | 2 | $47 / month | Unlimited | Smaller production workloads |
| Growth | 10 | $199 / month | Unlimited | Regular production workloads |
| Scale | 25 | $399 / month | Unlimited | Large scraping workloads |
| Pro | 50 | $699 / month | Unlimited | High-volume scraping workloads |
| Business | 100 | $1,799 / month | Unlimited | Very high-volume scraping workloads |
For Pro and Business plans, [contact sales](https://geonode.com/contact).
> Pricing and plan availability can change. Check the [Geonode Scraper API pricing page](https://geonode.com/products/scraper-api) for the latest pricing and available plans.
## How Unlimited Pricing Works
Unlimited plans do not use a monthly request allowance.
Instead, each plan provides a specific number of **concurrent threads**. This determines how many extractions can run at the same time.
For example, a plan with 10 concurrent threads can process up to 10 extraction tasks concurrently.
If all available threads are busy, additional work waits in the queue until a thread becomes available.
This means that a higher thread count increases your potential throughput when your workload can be processed in parallel.
## Understanding Concurrency
Concurrency is the number of extraction tasks that can run at the same time.
For example:
| Concurrent threads | Maximum parallel extractions |
| -----------------: | ---------------------------: |
| 2 | 2 |
| 10 | 10 |
| 25 | 25 |
| 50 | 50 |
| 100 | 100 |
When all threads are occupied, new work is queued rather than rejected because of concurrency.
## Job Limits
Unlimited monthly requests do not mean that an individual Batch or Crawl job can contain an unlimited number of URLs or pages.
Job-size limits apply separately to each Batch and Crawl job.
| Plan | Concurrency | Max URLs per batch job | Max pages per crawl job |
| -------- | ----------: | ---------------------: | ----------------------: |
| Starter | 2 | 50 | 50 |
| Growth | 10 | 500 | 500 |
| Scale | 25 | 1,000 | 1,000 |
| Pro | 50 | 2,000 | 2,000 |
| Business | 100 | 4,000 | 4,000 |
If your workload is larger than the maximum size of a single job, split it across multiple jobs.
## Choosing the Right Plan
Choose your plan based on how much work you need to process **at the same time**.
| Workload | Recommended plan |
| ---------------------------- | ---------------- |
| Testing and evaluation | Free Trial |
| Small production workloads | Starter |
| Regular production workloads | Growth |
| Large scraping workloads | Scale |
| High-volume scraping | Pro |
| Very high-volume scraping | Business |
If you are unsure which plan you need, start with the free trial and measure your workload before moving to a paid plan.
## What Is Included?
Unlimited plans provide access to the Scraper API features used for web extraction.
You can use features such as:
* HTML and Markdown output
* JavaScript rendering
* Proxy selection
* Geo-targeting
* Batch extraction
* Crawl jobs
These features do not create additional request charges on an Unlimited plan.
## Unlimited Requests
Unlimited means that your plan does not have a monthly request allowance.
You can continue submitting extraction requests throughout the billing period without tracking a remaining monthly request balance.
Your plan still determines how many extractions can run concurrently.
For example, with 10 concurrent threads:
1. Up to 10 extractions can run at the same time.
2. Additional work waits when all 10 threads are occupied.
3. A waiting extraction starts when a thread becomes available.
## Unlimited vs Request-Based Pricing
Unlimited and request-based pricing use different billing models.
| | Unlimited | Request-Based |
| ------------------------- | --------------------- | -------------------------------- |
| Pricing model | Flat monthly price | Based on request volume |
| Monthly request allowance | Unlimited | Plan-dependent |
| Main pricing factor | Concurrent threads | Successful extractions |
| Best for | High-volume workloads | Variable workloads |
| Monthly cost | Predictable | Based on selected plan and usage |
If you want to understand how request-based pricing works, including request counting, free requests, subscriptions, and Pay As You Go, see the [Request-Based Pricing](/docs/scraper-api/additional-resources/request_based_pricing) guide.
## Related Guides
* [Request-Based Pricing](/docs/scraper-api/additional-resources/request_based_pricing)
* [Extraction](/docs/scraper-api/guides/extraction/01_understanding_extraction)
* [Batch](/docs/scraper-api/guides/batch/00_understanding_batch)
* [Dashboard Overview](/docs/scraper-api/dashboard-guides/overview)
## Get Started
Ready to use the Unlimited Scraper API? Open the [Geonode dashboard](https://app.geonode.com/login) to choose a plan.
# Dashboard Overview (/docs/scraper-api/dashboard-guides/overview)
The GeoNode Dashboard provides a graphical interface for using the **Scraper** and **Map** APIs without writing code. From the dashboard, you can submit extraction and mapping requests, monitor usage, review previous jobs, and manage your account from a single place.
After signing in, you'll see the main dashboard with navigation on the left and your selected tool in the main workspace.
## Available Requests
The **Available Requests** card provides a quick overview of your current usage, including:
* Remaining requests available in your account.
* Your free request allowance.
* When your request quota renews.
* The number of requests used during the last 24 hours.
This allows you to monitor your usage without leaving the dashboard.
## Upgrade Your Plan
The **Upgrade Your Plan** section lets you review available pricing plans and upgrade your subscription whenever you need additional requests or higher usage limits.
## Dashboard Navigation
Use the navigation menu on the left to switch between the available tools.
* **Extraction** - Extract structured content from one or more web pages.
* **Mapping** - Discover URLs from websites before extracting their content.
* **Search** - Search the web for specific content.
* **Crawl** - Crawl a website and extract its content.
## Settings
Both the Scraper and Map dashboards include a **Settings** panel where you can configure your requests before starting a job.
Detailed explanations of each setting are available in the corresponding API Guides and are not repeated in the Dashboard documentation.
## Recent Jobs
The lower section of each dashboard displays your recent activity.
Depending on the selected tool, you can:
* View recent extraction or mapping jobs.
* Check the current job status.
* Search previous jobs.
* Open completed results.
* Review execution times.
Dedicated guides explain how to work with jobs in more detail.
## Statistics
The **Statistics** tab provides an overview of your API usage and request activity, helping you monitor your overall usage over time.
## Next Steps
Now that you're familiar with the dashboard layout, continue with one of the following guides:
* **Scraper Overview** – Learn how to extract content from web pages using the Dashboard.
* **Map Overview** – Learn how to discover URLs from websites using the Dashboard.
* **Crawl Overview** – Learn how to crawl a website and download results using the Dashboard.
* **Search Overview** – Learn how to search the web using the Dashboard.
# CrewAI (/docs/scraper-api/developer-guides/crewai)
{/* DRAFT: do not publish / do not add to meta.json until approved */}
`geonode-scraper-crewai` builds CrewAI tools you can call with `tool.run(...)` or attach to an `Agent`. Each tool returns a JSON-friendly dict.
Responses below are from live runs against `https://scraper.geonode.io`.
## Setup
CrewAI requires Python 3.10 through 3.13.
```bash
pip install geonode-scraper-crewai python-dotenv
```
Create a `.env` file:
```bash
GEONODE_SCRAPER_API_KEY=your_geonode_key
SCRAPER_API_BASE_URL=https://scraper.geonode.io
```
```python
import os
from dotenv import load_dotenv
from geonode_scraper_crewai import build_crewai_tools
from geonode_scraper_tools_core import ScraperToolSettings
load_dotenv()
settings = ScraperToolSettings(
host=os.environ.get("SCRAPER_API_BASE_URL", "https://scraper.geonode.io"),
api_key=os.environ["GEONODE_SCRAPER_API_KEY"],
)
tools = build_crewai_tools(settings=settings)
by_name = {tool.name: tool for tool in tools}
```
Use `https://scraper.geonode.io` for production.
### Response shape
Every tool returns a dict like:
```python
{
"ok": True,
"operation": "extract",
"attempts": 1,
"result": { ... },
}
```
Read the payload from `response["result"]`.
## Tools overview
| Group | Tool names |
| -------------- | ------------------------------------------------------------------------------------------------------- |
| **Extraction** | `scraper_extract_content`, `scraper_get_job_result`, `scraper_wait_for_job`, `scraper_list_jobs` |
| **Batch** | `scraper_create_batch`, `scraper_get_batch_status`, `scraper_wait_for_batch`, `scraper_list_batch_jobs` |
| **Crawl** | `scraper_create_crawl`, `scraper_get_crawl_status`, `scraper_wait_for_crawl`, `scraper_list_crawl_jobs` |
| **Map** | `scraper_map_urls`, `scraper_list_map_jobs`, `scraper_get_map_job` |
| **Search** | `scraper_search`, `scraper_list_search_jobs`, `scraper_get_search_job` |
| **Account** | `scraper_get_statistics`, `scraper_get_concurrency_usage`, `scraper_check_health` |
***
## Extraction
### scraper\_extract\_content (sync)
```python
response = by_name["scraper_extract_content"].run(
url="https://docs.geonode.com/docs/scraper-api/quick-start",
formats=["markdown"],
processing_mode="sync",
)
result = response["result"]
markdown = (result.get("data") or {}).get("markdown") or ""
print("ok:", response["ok"])
print("tokens:", result.get("tokens_charged"))
print("markdown_len:", len(markdown))
print("preview:", markdown[:200])
```
Response (live run):
```text
ok: True
tokens: 1
markdown_len: 15871
preview: ---
canonical: https://docs.geonode.com/docs/scraper-api/quick-start
meta-description: Get your API key, authenticate requests, and choose the right API for your use case.
...
```
### scraper\_extract\_content (async)
```python
response = by_name["scraper_extract_content"].run(
url="https://docs.geonode.com/docs/scraper-api/quick-start",
formats=["markdown"],
processing_mode="async",
)
result = response["result"]
print("job_id:", result["job_id"])
print("status:", result["status"])
```
Response (live run):
```text
job_id: e35c4d0e-9df3-43b9-9446-33c517b1cfc4
status: queued
```
### scraper\_wait\_for\_job
```python
response = by_name["scraper_wait_for_job"].run(
job_id="e35c4d0e-9df3-43b9-9446-33c517b1cfc4",
timeout_seconds=120,
)
result = response["result"]
markdown = (result.get("data") or {}).get("markdown") or ""
print("status:", result["status"])
print("markdown_len:", len(markdown))
print("poll_attempts:", response.get("poll_attempts"))
```
Response (live run):
```text
status: completed
markdown_len: 15871
poll_attempts: 4
```
### scraper\_get\_job\_result
```python
response = by_name["scraper_get_job_result"].run(
job_id="e35c4d0e-9df3-43b9-9446-33c517b1cfc4",
)
result = response["result"]
print("status:", result["status"])
print("tokens:", result.get("tokens_charged"))
```
Response (live run):
```text
status: completed
tokens: 1
```
### scraper\_list\_jobs
```python
response = by_name["scraper_list_jobs"].run(page=1, page_size=3)
result = response["result"]
print("page:", result["page"], "page_size:", result["page_size"])
for job in result.get("jobs") or []:
print(job["job_id"], job["status"], job.get("url"))
```
Response (live run):
```text
page: 1 page_size: 3
752b8599-5915-441c-bc1f-b9fb0d938f72 completed ...
0bcd424e-b5ab-4b85-9cbd-5b44f7ed433c completed http://example.com/
0dbb126c-4331-4b4b-9d55-0c380ba69ae7 completed http://example.com/
```
***
## Batch
### scraper\_create\_batch
```python
response = by_name["scraper_create_batch"].run(
urls=[
"https://docs.geonode.com/docs/scraper-api/quick-start",
"https://docs.geonode.com/docs/scraper-api",
],
formats=["markdown"],
)
result = response["result"]
print("job_id:", result["job_id"])
print("accepted_urls:", result["accepted_urls"])
print("status:", result["status"])
```
Response (live run):
```text
job_id: 037be6d5-06e8-480f-9e22-4eac6cfb2966
accepted_urls: 2
status: queued
```
### scraper\_get\_batch\_status
```python
response = by_name["scraper_get_batch_status"].run(
job_id="037be6d5-06e8-480f-9e22-4eac6cfb2966",
page=1,
page_size=10,
)
result = response["result"]
print(result["status"], result["completed_urls"], "/", result["total_urls"])
```
Response (live run, mid-job):
```text
processing 0 / 2
```
### scraper\_wait\_for\_batch
```python
response = by_name["scraper_wait_for_batch"].run(
job_id="037be6d5-06e8-480f-9e22-4eac6cfb2966",
timeout_seconds=120,
)
result = response["result"]
print(result["status"], result["completed_urls"], "/", result["total_urls"])
print("poll_attempts:", response.get("poll_attempts"))
```
Response (live run):
```text
completed 2 / 2
poll_attempts: 3
```
### scraper\_list\_batch\_jobs
```python
response = by_name["scraper_list_batch_jobs"].run(page=1, page_size=3)
for job in (response["result"].get("jobs") or []):
print(
job["job_id"],
job["status"],
job["completed_urls"],
"/",
job["accepted_urls"],
)
```
Response (live run):
```text
037be6d5-06e8-480f-9e22-4eac6cfb2966 completed 2 / 2
600bb35a-d6fc-4e05-be49-816b8f3ad5d7 completed 2 / 2
55b11790-5c7f-4945-bc13-bd9a365a1835 completed 2 / 2
```
***
## Crawl
### scraper\_create\_crawl
```python
response = by_name["scraper_create_crawl"].run(
url="https://docs.geonode.com/docs/scraper-api",
depth=2,
limit=3,
formats=["markdown"],
same_domain_only=True,
)
result = response["result"]
print("job_id:", result["job_id"])
print("estimated_pages:", result["estimated_pages"])
print("status:", result["status"])
```
Response (live run):
```text
job_id: dbc0fa1a-1fea-496c-abe4-a33c2b91efb7
estimated_pages: 3
status: queued
```
### scraper\_get\_crawl\_status
```python
response = by_name["scraper_get_crawl_status"].run(
job_id="dbc0fa1a-1fea-496c-abe4-a33c2b91efb7",
page=1,
page_size=10,
)
result = response["result"]
print(
result["status"],
result.get("completed_pages"),
"/",
result.get("total_pages"),
)
```
Response (live run, early poll):
```text
processing 0 / 1
```
### scraper\_wait\_for\_crawl
```python
response = by_name["scraper_wait_for_crawl"].run(
job_id="dbc0fa1a-1fea-496c-abe4-a33c2b91efb7",
timeout_seconds=180,
)
result = response["result"]
print(result["status"], result["completed_pages"], "/", result["total_pages"])
print("poll_attempts:", response.get("poll_attempts"))
```
Response (live run):
```text
completed 3 / 3
poll_attempts: 4
```
### scraper\_list\_crawl\_jobs
```python
response = by_name["scraper_list_crawl_jobs"].run(page=1, page_size=2)
for job in (response["result"].get("jobs") or []):
print(
job["job_id"],
job["status"],
job["completed_pages"],
"/",
job["total_pages"],
)
```
Response (live run):
```text
dbc0fa1a-1fea-496c-abe4-a33c2b91efb7 completed 3 / 3
ab83f113-5086-46d3-a045-3a42263b5c87 completed 3 / 3
```
***
## Map
### scraper\_map\_urls
```python
response = by_name["scraper_map_urls"].run(
url="https://docs.geonode.com/docs/scraper-api",
)
result = response["result"]
links = result.get("links") or []
print("link_count:", result.get("links_count") or len(links))
for link in links[:5]:
print(link.get("source"), link.get("url"))
```
Response (live run):
```text
link_count: 112
sitemap https://docs.geonode.com/docs/scraper-api
sitemap https://docs.geonode.com/docs/scraper-api/quick-start
sitemap https://docs.geonode.com/docs/scraper-api/additional-resources/choosing_scraper_api_plan
sitemap https://docs.geonode.com/docs/scraper-api/additional-resources/faq
sitemap https://docs.geonode.com/docs/scraper-api/additional-resources/pricing-and-requests
```
### scraper\_list\_map\_jobs
```python
response = by_name["scraper_list_map_jobs"].run(page=1, page_size=2)
for job in (response["result"].get("jobs") or []):
print(job["job_id"], job["status"], job.get("links_count"), job.get("url"))
```
Response (live run):
```text
88b13fab-8d1b-421d-aa0e-256dd705b3aa completed 112 https://docs.geonode.com/docs/scraper-api
3abfc672-2aff-4daf-9f68-f033978ddfbd completed 112 https://docs.geonode.com/docs/scraper-api
```
### scraper\_get\_map\_job
```python
response = by_name["scraper_get_map_job"].run(
job_id="88b13fab-8d1b-421d-aa0e-256dd705b3aa",
)
result = response["result"]
print("status:", result["status"])
print("link_count:", result.get("links_count"))
for link in (result.get("links") or [])[:3]:
print(link.get("source"), link.get("url"))
```
Response (live run):
```text
status: completed
link_count: 112
sitemap https://docs.geonode.com/docs/scraper-api
sitemap https://docs.geonode.com/docs/scraper-api/quick-start
sitemap https://docs.geonode.com/docs/scraper-api/additional-resources/choosing_scraper_api_plan
```
***
## Search
### scraper\_search
```python
response = by_name["scraper_search"].run(query="geonode scraper api")
result = response["result"]
print("job_id:", result["job_id"])
print("hit_count:", result.get("results_count") or len(result.get("results") or []))
for hit in (result.get("results") or [])[:5]:
print(hit["position"], hit["title"], hit["url"])
```
Response (live run):
```text
job_id: 7eeb9599-c39a-4eba-8e0d-2eaed5a8deab
hit_count: 15
1 Wholesale proxies & web data infrastructure | Geonode https://geonode.com/
2 Geonode - PyPI https://pypi.org/user/Geonode/
3 pavel.s - PyPI https://pypi.org/user/pavel.s/
4 Geonode Documentation | Geonode https://docs.geonode.com/
5 GeoNode https://geonode.org/
```
### scraper\_list\_search\_jobs
```python
response = by_name["scraper_list_search_jobs"].run(page=1, page_size=2)
for job in (response["result"].get("jobs") or []):
print(job["job_id"], job.get("query"), job["status"], job.get("results_count"))
```
Response (live run):
```text
7eeb9599-c39a-4eba-8e0d-2eaed5a8deab geonode scraper api completed 15
f80269e1-1911-4fd4-b1a1-fc0f7bbae7f9 geonode scraper api completed 15
```
### scraper\_get\_search\_job
```python
response = by_name["scraper_get_search_job"].run(
job_id="7eeb9599-c39a-4eba-8e0d-2eaed5a8deab",
)
result = response["result"]
print("status:", result["status"])
print("hit_count:", result.get("results_count"))
for hit in (result.get("results") or [])[:3]:
print(hit["position"], hit["title"], hit["url"])
```
Response (live run):
```text
status: completed
hit_count: 15
1 Wholesale proxies & web data infrastructure | Geonode https://geonode.com/
2 Geonode - PyPI https://pypi.org/user/Geonode/
3 pavel.s - PyPI https://pypi.org/user/pavel.s/
```
***
## Statistics, usage, and health
### scraper\_get\_statistics
```python
response = by_name["scraper_get_statistics"].run()
result = response["result"]
print("extraction_count:", result.get("extraction_count"))
print("success_rate:", result.get("success_rate"))
```
Response (live run):
```text
extraction_count: 4401
success_rate: 0.95137
```
### scraper\_get\_concurrency\_usage
```python
response = by_name["scraper_get_concurrency_usage"].run()
result = response["result"]
print(
result["work_concurrency_in_use"],
"/",
result["work_concurrency_limit"],
)
```
Response (live run):
```text
0 / 50
```
### scraper\_check\_health
```python
response = by_name["scraper_check_health"].run()
result = response["result"]
print(result.get("service"), result.get("status"), result.get("version"))
```
Response (live run):
```text
Scraper API ok 0.1.0
```
***
## Selecting a subset of tools
```python
tools = build_crewai_tools(
settings=settings,
operations=["extract", "map_urls", "create_crawl", "wait_for_crawl"],
)
print([tool.name for tool in tools])
```
Response (live run):
```text
['scraper_extract_content', 'scraper_map_urls', 'scraper_create_crawl', 'scraper_wait_for_crawl']
```
***
## Use tools with an Agent
Attach the tools to a CrewAI agent:
```python
from crewai import Agent
from geonode_scraper_crewai import build_crewai_tools
from geonode_scraper_tools_core import ScraperToolSettings
settings = ScraperToolSettings(
host="https://scraper.geonode.io",
api_key=os.environ["GEONODE_SCRAPER_API_KEY"],
)
agent = Agent(
role="Web Researcher",
goal="Extract and inspect web content.",
backstory="Focused on pulling structured data from URLs.",
tools=build_crewai_tools(settings=settings),
)
```
The agent can call the same tools shown above. Direct checks still use `tool.run(...)`.
***
## Toolkit helper
```python
from geonode_scraper_crewai import ScraperCrewAIToolkit
from geonode_scraper_tools_core import ScraperToolSettings
settings = ScraperToolSettings(
host="https://scraper.geonode.io",
api_key=os.environ["GEONODE_SCRAPER_API_KEY"],
)
toolkit = ScraperCrewAIToolkit.from_settings(settings)
tools = toolkit.get_tools()
```
For package versions and changelog, see [geonode-scraper-crewai on PyPI](https://pypi.org/project/geonode-scraper-crewai/).
# LangChain (/docs/scraper-api/developer-guides/langchain)
{/* DRAFT: do not publish / do not add to meta.json until approved */}
`geonode-scraper-langchain` builds LangChain `StructuredTool` objects you can call with `tool.invoke(...)` or pass into an agent. Each tool returns a JSON-friendly dict.
Responses below are from live runs against `https://scraper.geonode.io`.
## Setup
```bash
pip install geonode-scraper-langchain python-dotenv
```
Create a `.env` file:
```bash
GEONODE_SCRAPER_API_KEY=your_geonode_key
SCRAPER_API_BASE_URL=https://scraper.geonode.io
```
```python
import os
from dotenv import load_dotenv
from geonode_scraper_langchain import build_langchain_tools
from geonode_scraper_tools_core import ScraperToolSettings
load_dotenv()
settings = ScraperToolSettings(
host=os.environ.get("SCRAPER_API_BASE_URL", "https://scraper.geonode.io"),
api_key=os.environ["GEONODE_SCRAPER_API_KEY"],
)
tools = build_langchain_tools(settings=settings)
by_name = {tool.name: tool for tool in tools}
```
Use `https://scraper.geonode.io` for production.
### Response shape
Every tool returns a dict like:
```python
{
"ok": True,
"operation": "extract",
"attempts": 1,
"result": { ... },
}
```
Read the payload from `response["result"]`.
## Tools overview
| Group | Tool names |
| -------------- | ------------------------------------------------------------------------------------------------------- |
| **Extraction** | `scraper_extract_content`, `scraper_get_job_result`, `scraper_wait_for_job`, `scraper_list_jobs` |
| **Batch** | `scraper_create_batch`, `scraper_get_batch_status`, `scraper_wait_for_batch`, `scraper_list_batch_jobs` |
| **Crawl** | `scraper_create_crawl`, `scraper_get_crawl_status`, `scraper_wait_for_crawl`, `scraper_list_crawl_jobs` |
| **Map** | `scraper_map_urls`, `scraper_list_map_jobs`, `scraper_get_map_job` |
| **Search** | `scraper_search`, `scraper_list_search_jobs`, `scraper_get_search_job` |
| **Account** | `scraper_get_statistics`, `scraper_get_concurrency_usage`, `scraper_check_health` |
***
## Extraction
### scraper\_extract\_content (sync)
```python
response = by_name["scraper_extract_content"].invoke(
{
"url": "https://docs.geonode.com/docs/scraper-api/quick-start",
"formats": ["markdown"],
"processing_mode": "sync",
}
)
result = response["result"]
markdown = (result.get("data") or {}).get("markdown") or ""
print("ok:", response["ok"])
print("tokens:", result.get("tokens_charged"))
print("markdown_len:", len(markdown))
print("preview:", markdown[:200])
```
Response (live run):
```text
ok: True
tokens: 1
markdown_len: 15871
preview: ---
canonical: https://docs.geonode.com/docs/scraper-api/quick-start
meta-description: Get your API key, authenticate requests, and choose the right API for your use case.
...
```
### scraper\_extract\_content (async)
```python
response = by_name["scraper_extract_content"].invoke(
{
"url": "https://docs.geonode.com/docs/scraper-api/quick-start",
"formats": ["markdown"],
"processing_mode": "async",
}
)
result = response["result"]
print("job_id:", result["job_id"])
print("status:", result["status"])
```
Response (live run):
```text
job_id: 2ac66637-905c-4892-a6b9-5ace58b2ffe9
status: queued
```
### scraper\_wait\_for\_job
```python
response = by_name["scraper_wait_for_job"].invoke(
{
"job_id": "2ac66637-905c-4892-a6b9-5ace58b2ffe9",
"timeout_seconds": 120,
}
)
result = response["result"]
markdown = (result.get("data") or {}).get("markdown") or ""
print("status:", result["status"])
print("markdown_len:", len(markdown))
print("poll_attempts:", response.get("poll_attempts"))
```
Response (live run):
```text
status: completed
markdown_len: 15871
poll_attempts: 3
```
### scraper\_get\_job\_result
```python
response = by_name["scraper_get_job_result"].invoke(
{"job_id": "2ac66637-905c-4892-a6b9-5ace58b2ffe9"}
)
result = response["result"]
print("status:", result["status"])
print("tokens:", result.get("tokens_charged"))
```
Response (live run):
```text
status: completed
tokens: 1
```
### scraper\_list\_jobs
```python
response = by_name["scraper_list_jobs"].invoke({"page": 1, "page_size": 3})
result = response["result"]
print("page:", result["page"], "page_size:", result["page_size"])
for job in result.get("jobs") or []:
print(job["job_id"], job["status"], job.get("url"))
```
Response (live run):
```text
page: 1 page_size: 3
752b8599-5915-441c-bc1f-b9fb0d938f72 completed ...
0bcd424e-b5ab-4b85-9cbd-5b44f7ed433c completed http://example.com/
0dbb126c-4331-4b4b-9d55-0c380ba69ae7 completed http://example.com/
```
***
## Batch
### scraper\_create\_batch
```python
response = by_name["scraper_create_batch"].invoke(
{
"urls": [
"https://docs.geonode.com/docs/scraper-api/quick-start",
"https://docs.geonode.com/docs/scraper-api",
],
"formats": ["markdown"],
}
)
result = response["result"]
print("job_id:", result["job_id"])
print("accepted_urls:", result["accepted_urls"])
print("status:", result["status"])
```
Response (live run):
```text
job_id: 600bb35a-d6fc-4e05-be49-816b8f3ad5d7
accepted_urls: 2
status: queued
```
### scraper\_get\_batch\_status
```python
response = by_name["scraper_get_batch_status"].invoke(
{
"job_id": "600bb35a-d6fc-4e05-be49-816b8f3ad5d7",
"page": 1,
"page_size": 10,
}
)
result = response["result"]
print(result["status"], result["completed_urls"], "/", result["total_urls"])
```
Response (live run, mid-job):
```text
processing 0 / 2
```
### scraper\_wait\_for\_batch
```python
response = by_name["scraper_wait_for_batch"].invoke(
{
"job_id": "600bb35a-d6fc-4e05-be49-816b8f3ad5d7",
"timeout_seconds": 120,
}
)
result = response["result"]
print(result["status"], result["completed_urls"], "/", result["total_urls"])
print("poll_attempts:", response.get("poll_attempts"))
```
Response (live run):
```text
completed 2 / 2
poll_attempts: 3
```
### scraper\_list\_batch\_jobs
```python
response = by_name["scraper_list_batch_jobs"].invoke({"page": 1, "page_size": 3})
for job in (response["result"].get("jobs") or []):
print(
job["job_id"],
job["status"],
job["completed_urls"],
"/",
job["accepted_urls"],
)
```
Response (live run):
```text
600bb35a-d6fc-4e05-be49-816b8f3ad5d7 completed 2 / 2
55b11790-5c7f-4945-bc13-bd9a365a1835 completed 2 / 2
d8e92d3a-c939-4ec6-b133-3e04c6de71d1 completed 2 / 2
```
***
## Crawl
### scraper\_create\_crawl
```python
response = by_name["scraper_create_crawl"].invoke(
{
"url": "https://docs.geonode.com/docs/scraper-api",
"depth": 2,
"limit": 3,
"formats": ["markdown"],
"same_domain_only": True,
}
)
result = response["result"]
print("job_id:", result["job_id"])
print("estimated_pages:", result["estimated_pages"])
print("status:", result["status"])
```
Response (live run):
```text
job_id: ab83f113-5086-46d3-a045-3a42263b5c87
estimated_pages: 3
status: queued
```
### scraper\_get\_crawl\_status
```python
response = by_name["scraper_get_crawl_status"].invoke(
{
"job_id": "ab83f113-5086-46d3-a045-3a42263b5c87",
"page": 1,
"page_size": 10,
}
)
result = response["result"]
print(
result["status"],
result.get("completed_pages"),
"/",
result.get("total_pages"),
)
```
Response (live run, early poll):
```text
processing 0 / 1
```
### scraper\_wait\_for\_crawl
```python
response = by_name["scraper_wait_for_crawl"].invoke(
{
"job_id": "ab83f113-5086-46d3-a045-3a42263b5c87",
"timeout_seconds": 180,
}
)
result = response["result"]
print(result["status"], result["completed_pages"], "/", result["total_pages"])
```
Response (live run):
```text
completed 3 / 3
```
### scraper\_list\_crawl\_jobs
```python
response = by_name["scraper_list_crawl_jobs"].invoke({"page": 1, "page_size": 2})
for job in (response["result"].get("jobs") or []):
print(
job["job_id"],
job["status"],
job["completed_pages"],
"/",
job["total_pages"],
)
```
Response (live run):
```text
ab83f113-5086-46d3-a045-3a42263b5c87 completed 3 / 3
9f447ced-4aa3-45a3-b9af-421f751b8cec completed 3 / 3
```
***
## Map
### scraper\_map\_urls
```python
response = by_name["scraper_map_urls"].invoke(
{"url": "https://docs.geonode.com/docs/scraper-api"}
)
result = response["result"]
links = result.get("links") or []
print("link_count:", result.get("links_count") or len(links))
for link in links[:5]:
print(link.get("source"), link.get("url"))
```
Response (live run):
```text
link_count: 112
sitemap https://docs.geonode.com/docs/scraper-api
sitemap https://docs.geonode.com/docs/scraper-api/quick-start
sitemap https://docs.geonode.com/docs/scraper-api/additional-resources/choosing_scraper_api_plan
sitemap https://docs.geonode.com/docs/scraper-api/additional-resources/faq
sitemap https://docs.geonode.com/docs/scraper-api/additional-resources/pricing-and-requests
```
### scraper\_list\_map\_jobs
```python
response = by_name["scraper_list_map_jobs"].invoke({"page": 1, "page_size": 2})
for job in (response["result"].get("jobs") or []):
print(job["job_id"], job["status"], job.get("links_count"), job.get("url"))
```
Response (live run):
```text
3abfc672-2aff-4daf-9f68-f033978ddfbd completed 112 https://docs.geonode.com/docs/scraper-api
c7cb6e40-ddfa-416a-b255-00dfdfd717bc completed 112 https://docs.geonode.com/docs/scraper-api
```
### scraper\_get\_map\_job
```python
response = by_name["scraper_get_map_job"].invoke(
{"job_id": "3abfc672-2aff-4daf-9f68-f033978ddfbd"}
)
result = response["result"]
print("status:", result["status"])
print("link_count:", result.get("links_count"))
for link in (result.get("links") or [])[:3]:
print(link.get("source"), link.get("url"))
```
Response (live run):
```text
status: completed
link_count: 112
sitemap https://docs.geonode.com/docs/scraper-api
sitemap https://docs.geonode.com/docs/scraper-api/quick-start
sitemap https://docs.geonode.com/docs/scraper-api/additional-resources/choosing_scraper_api_plan
```
***
## Search
### scraper\_search
```python
response = by_name["scraper_search"].invoke({"query": "geonode scraper api"})
result = response["result"]
print("job_id:", result["job_id"])
print("hit_count:", result.get("results_count") or len(result.get("results") or []))
for hit in (result.get("results") or [])[:5]:
print(hit["position"], hit["title"], hit["url"])
```
Response (live run):
```text
job_id: f80269e1-1911-4fd4-b1a1-fc0f7bbae7f9
hit_count: 15
1 Wholesale proxies & web data infrastructure | Geonode https://geonode.com/
2 Geonode - PyPI https://pypi.org/user/Geonode/
3 pavel.s - PyPI https://pypi.org/user/pavel.s/
4 Geonode Documentation | Geonode https://docs.geonode.com/
5 GeoNode https://geonode.org/
```
### scraper\_list\_search\_jobs
```python
response = by_name["scraper_list_search_jobs"].invoke({"page": 1, "page_size": 2})
for job in (response["result"].get("jobs") or []):
print(job["job_id"], job.get("query"), job["status"], job.get("results_count"))
```
Response (live run):
```text
f80269e1-1911-4fd4-b1a1-fc0f7bbae7f9 geonode scraper api completed 15
29809a61-49b6-4bf6-925b-aa32dd91e441 geonode scraper api completed 15
```
### scraper\_get\_search\_job
```python
response = by_name["scraper_get_search_job"].invoke(
{"job_id": "f80269e1-1911-4fd4-b1a1-fc0f7bbae7f9"}
)
result = response["result"]
print("status:", result["status"])
print("hit_count:", result.get("results_count"))
for hit in (result.get("results") or [])[:3]:
print(hit["position"], hit["title"], hit["url"])
```
Response (live run):
```text
status: completed
hit_count: 15
1 Wholesale proxies & web data infrastructure | Geonode https://geonode.com/
2 Geonode - PyPI https://pypi.org/user/Geonode/
3 pavel.s - PyPI https://pypi.org/user/pavel.s/
```
***
## Statistics, usage, and health
### scraper\_get\_statistics
```python
response = by_name["scraper_get_statistics"].invoke({})
result = response["result"]
print("extraction_count:", result.get("extraction_count"))
print("success_rate:", result.get("success_rate"))
```
Response (live run):
```text
extraction_count: 4392
success_rate: 0.95128
```
### scraper\_get\_concurrency\_usage
```python
response = by_name["scraper_get_concurrency_usage"].invoke({})
result = response["result"]
print(
result["work_concurrency_in_use"],
"/",
result["work_concurrency_limit"],
)
```
Response (live run):
```text
0 / 50
```
### scraper\_check\_health
```python
response = by_name["scraper_check_health"].invoke({})
result = response["result"]
print(result.get("service"), result.get("status"), result.get("version"))
```
Response (live run):
```text
Scraper API ok 0.1.0
```
***
## Selecting a subset of tools
```python
tools = build_langchain_tools(
settings=settings,
operations=["extract", "map_urls", "create_batch", "wait_for_batch"],
)
print([tool.name for tool in tools])
```
Response (live run):
```text
['scraper_extract_content', 'scraper_map_urls', 'scraper_create_batch', 'scraper_wait_for_batch']
```
***
## Use tools with an agent
Pass the tool list into your LangChain agent or bind them to a chat model. Example with `bind_tools`:
```python
from langchain_openai import ChatOpenAI
llm = ChatOpenAI(model="gpt-4o-mini")
llm_with_tools = llm.bind_tools(tools)
# The model can request tool calls; your agent loop then runs tool.invoke(...)
```
Or build an agent that owns the tools (exact API depends on your LangChain version):
```python
# Pseudocode: wire `tools` into your preferred LangChain agent helper
# agent = create_agent(model=llm, tools=tools)
# agent.invoke({"messages": [("user", "Extract markdown from https://example.com")]})
```
The tool calls themselves match the `tool.invoke({...})` examples above.
***
## Toolkit helper
You can also use the toolkit class:
```python
from geonode_scraper_langchain import ScraperLangChainToolkit
from geonode_scraper_tools_core import ScraperToolSettings
settings = ScraperToolSettings(
host="https://scraper.geonode.io",
api_key=os.environ["GEONODE_SCRAPER_API_KEY"],
)
toolkit = ScraperLangChainToolkit.from_settings(settings)
tools = toolkit.get_tools()
```
For package versions and changelog, see [geonode-scraper-langchain on PyPI](https://pypi.org/project/geonode-scraper-langchain/).
# Python SDK (/docs/scraper-api/developer-guides/python-sdk)
{/* DRAFT: do not publish / do not add to meta.json until approved */}
The Geonode Scraper Python SDK wraps every public Scraper API endpoint. This guide:
1. **API overview:** what each API group does and when to use it
2. **Shared request options:** enums used across APIs (`formats`, `processing_mode`, proxy, wait)
3. **Each API section:** that API’s methods, request fields, then code + live response per method
Responses below are from live runs against `https://scraper.geonode.io`.
## Setup
```bash
pip install geonode-scraper-sdk python-dotenv
```
Create a `.env` file:
```bash
GEONODE_SCRAPER_API_KEY=your_geonode_key
SCRAPER_API_BASE_URL=https://scraper.geonode.io
```
```python
import os
from dotenv import load_dotenv
from geonode_scraper_sdk import Configuration, ApiClient
load_dotenv()
configuration = Configuration(
host=os.environ.get("SCRAPER_API_BASE_URL", "https://scraper.geonode.io"),
api_key={"ApiKeyAuth": os.environ["GEONODE_SCRAPER_API_KEY"]},
)
```
Use `https://scraper.geonode.io` for production. If `host` is omitted, the client defaults to `http://localhost`.
## API overview
| API | SDK class | Endpoint | When to use |
| -------------- | --------------- | ---------------- | ---------------------------------------- |
| **Extract** | `ExtractionApi` | `/v1/extract` | One URL → HTML/Markdown. Sync or async. |
| **Batch** | `BatchApi` | `/v1/batch` | Many known URLs in one job. |
| **Crawl** | `CrawlApi` | `/v1/crawl` | Seed URL + follow links (depth/limit). |
| **Map** | `MapApi` | `/v1/map` | Discover URLs without scraping content. |
| **Search** | `SearchApi` | `/v1/search` | Web search → ranked URLs. |
| **Usage** | `UsageApi` | `/v1/usage` | Live concurrency vs plan limit. |
| **Statistics** | `StatisticsApi` | `/v1/statistics` | Historical counts, tokens, success rate. |
| **System** | `SystemApi` | `/health` | Service health check. |
| **Webhooks** | `WebhooksApi` | `/v1/webhooks` | Callbacks when async jobs finish. |
**Typical flows:** single page → Extract sync · known URL list → Batch · whole site → Map then Crawl/Batch · discovery → Search then Extract · production async → ASYNC + Webhooks.
Each API section below lists that API’s methods and request fields, then walks every method with code and a live response.
***
## Shared request options
Enums and helpers reused by Extract, Batch, and Crawl. Per-API field tables live in each API section.
### Output formats
```python
from geonode_scraper_sdk import OutputFormat
# Available values:
OutputFormat.HTML # "html": raw page HTML
OutputFormat.MARKDOWN # "markdown": cleaned Markdown
```
Pass as a list: `formats=[OutputFormat.MARKDOWN]` or `[OutputFormat.HTML, OutputFormat.MARKDOWN]`. Extract defaults to `[HTML]`; Batch/Crawl default to server-side defaults if omitted.
### Processing mode (Extract only)
```python
from geonode_scraper_sdk import ProcessingMode
ProcessingMode.SYNC # "sync": block until content is ready (default)
ProcessingMode.ASYNC # "async": return job_id immediately; poll get_job_result
```
Batch, Crawl, Map, and Search are always async jobs (create → poll status / get job).
### Proxy settings
```python
from geonode_scraper_sdk import ProxySettings, ProxyType
ProxySettings(
country="US", # ISO 3166-1 alpha-2 (optional)
type=ProxyType.RESIDENTIAL, # residential | datacenter | mix
)
```
### Wait config (JS rendering)
Used with `render_js=True` to control headless browser timing:
```python
from geonode_scraper_sdk import WaitConfig, WaitUntil
WaitConfig(
wait_until=WaitUntil.NETWORKIDLE, # commit | domcontentloaded | load | networkidle
wait_for="#content", # CSS selector (optional)
wait_timeout=10000, # ms, 0-30000
)
```
### Custom headers
Pass a dict on Extract/Batch requests: `headers={"User-Agent": "my-bot/1.0"}`.
### Job status values
Async jobs move through: `queued` → `processing` → `completed` | `failed` | `cancelled`.
***
## Extraction API
`ExtractionApi` (`/v1/extract`)
**Methods:**
* `extract_v1_extract_post(extract_request)`: sync or async extract
* `get_job_result_v1_extract_job_id_get(job_id)`: fetch async job result
* `list_jobs_v1_extract_jobs_get(...)`: paginate past extract jobs
**`ExtractRequest` fields:**
| Field | Type | Notes |
| ----------------- | ---------------- | ---------------------------------- |
| `url` | `str` | Required. Target URL. |
| `formats` | `[OutputFormat]` | Default `[HTML]`. |
| `processing_mode` | `ProcessingMode` | `SYNC` (default) or `ASYNC`. |
| `render_js` | `bool` | Headless browser. Default `False`. |
| `proxy` | `ProxySettings` | Optional. |
| `headers` | `dict[str, str]` | Optional request headers. |
| `wait_config` | `WaitConfig` | Optional browser wait policy. |
### extract\_v1\_extract\_post (sync)
Scrape one URL and return content in the same response.
```python
from geonode_scraper_sdk import (
ApiClient,
ExtractRequest,
ExtractionApi,
OutputFormat,
ProcessingMode,
)
with ApiClient(configuration) as api_client:
api = ExtractionApi(api_client)
response = api.extract_v1_extract_post(
ExtractRequest(
url="https://docs.geonode.com/docs/scraper-api/quick-start",
formats=[OutputFormat.MARKDOWN],
processing_mode=ProcessingMode.SYNC,
)
)
markdown = response.data.markdown if response.data else ""
print("Scraped content length:", len(markdown or ""))
print("Tokens charged:", response.tokens_charged)
print("Preview:", (markdown or "")[:300])
```
Response (live run):
```text
Scraped content length: 15871
Tokens charged: 1
Preview: ---
canonical: https://docs.geonode.com/docs/scraper-api/quick-start
meta-description: Get your API key, authenticate requests, and choose the right API for your use case.
...
```
### extract\_v1\_extract\_post (async)
Submit a job and receive a `job_id` immediately. Poll with `get_job_result_v1_extract_job_id_get`.
```python
import time
from geonode_scraper_sdk import (
ApiClient,
ExtractRequest,
ExtractionApi,
OutputFormat,
ProcessingMode,
)
with ApiClient(configuration) as api_client:
api = ExtractionApi(api_client)
submit = api.extract_v1_extract_post(
ExtractRequest(
url="https://docs.geonode.com/docs/scraper-api/quick-start",
formats=[OutputFormat.MARKDOWN],
processing_mode=ProcessingMode.ASYNC,
)
)
print("Job ID:", submit.job_id)
while True:
job = api.get_job_result_v1_extract_job_id_get(str(submit.job_id))
print("Status:", job.status)
if str(job.status).lower().endswith("completed"):
if job.data and job.data.markdown:
print("Scraped content length:", len(job.data.markdown))
break
time.sleep(2)
```
Response (live run):
```text
Job ID: 6d92b4d5-c9c2-4b66-aa0a-98c07c3a31da
Status: queued
Status: completed
Scraped content length: 15871
```
### get\_job\_result\_v1\_extract\_job\_id\_get
Fetch status and content for a single extract job (used after async submit or to re-read a past job).
```python
from geonode_scraper_sdk import ApiClient, ExtractionApi
JOB_ID = "6d92b4d5-c9c2-4b66-aa0a-98c07c3a31da"
with ApiClient(configuration) as api_client:
api = ExtractionApi(api_client)
job = api.get_job_result_v1_extract_job_id_get(JOB_ID)
md_len = len(job.data.markdown) if job.data and job.data.markdown else 0
print("Job ID:", JOB_ID)
print("Status:", job.status)
print("Markdown length:", md_len)
```
Response (live run):
```text
Job ID: 6d92b4d5-c9c2-4b66-aa0a-98c07c3a31da
Status: JobStatus.COMPLETED
Markdown length: 15871
```
### list\_jobs\_v1\_extract\_jobs\_get
Paginate past extract jobs. Optional filters: `status`, `start_date`, `end_date`, `page`, `page_size`.
```python
from geonode_scraper_sdk import ApiClient, ExtractionApi
with ApiClient(configuration) as api_client:
api = ExtractionApi(api_client)
page = api.list_jobs_v1_extract_jobs_get(page=1, page_size=3)
jobs = getattr(page, "items", None) or getattr(page, "jobs", None) or []
for job in jobs:
print(job.job_id, job.status)
```
Response (live run):
```text
752b8599-5915-441c-bc1f-b9fb0d938f72 JobStatus.COMPLETED
0bcd424e-b5ab-4b85-9cbd-5b44f7ed433c JobStatus.COMPLETED
0706f42e-d9fd-4076-bfae-3cb365f4b634 JobStatus.COMPLETED
```
### extract\_v1\_extract\_post (JS + proxy)
Same method with `render_js`, residential proxy, and wait config for dynamic pages.
```python
from geonode_scraper_sdk import (
ApiClient,
ExtractRequest,
ExtractionApi,
OutputFormat,
ProcessingMode,
ProxySettings,
ProxyType,
WaitConfig,
WaitUntil,
)
with ApiClient(configuration) as api_client:
api = ExtractionApi(api_client)
response = api.extract_v1_extract_post(
ExtractRequest(
url="https://example.com",
formats=[OutputFormat.MARKDOWN],
processing_mode=ProcessingMode.SYNC,
render_js=True,
proxy=ProxySettings(country="US", type=ProxyType.RESIDENTIAL),
wait_config=WaitConfig(
wait_until=WaitUntil.NETWORKIDLE,
wait_timeout=10000,
),
)
)
markdown = response.data.markdown if response.data else ""
print("Scraped content length:", len(markdown or ""))
print("Tokens charged:", response.tokens_charged)
print("Preview:", (markdown or "")[:250])
```
Response (live run):
```text
Scraped content length: 250
Tokens charged: 1
Preview: ---
meta-viewport: width=device-width, initial-scale=1
title: Example Domain
---
# Example Domain
This domain is for use in documentation examples without needing permission. Avoid use in operations.
[Learn more](https://iana.org/domains/example)
```
***
## Batch API
`BatchApi` (`/v1/batch`)
**Methods:**
* `create_batch_v1_batch_post(batch_request)`: submit URLs as one job
* `get_batch_status_v1_batch_job_id_get(job_id, page, page_size)`: poll status / results
* `list_batch_jobs_v1_batch_jobs_get(...)`: list past batch jobs
* `cancel_batch_v1_batch_job_id_delete(job_id)`: cancel a running batch
**`BatchRequest` fields:**
| Field | Type | Notes |
| --------------------- | ---------------- | ------------------------------------------------- |
| `urls` | `[str]` | Required. 1-1000 URLs. |
| `formats` | `[OutputFormat]` | Optional. |
| `render_js` | `bool` | Apply to every URL. |
| `proxy` | `ProxySettings` | Optional. |
| `headers` | `dict[str, str]` | Optional. |
| `wait_config` | `WaitConfig` | Optional. |
| `ignore_invalid_urls` | `bool` | Default `True`. Skip bad URLs instead of failing. |
### create\_batch\_v1\_batch\_post
Submit many URLs as one batch job.
```python
from geonode_scraper_sdk import ApiClient, BatchApi, BatchRequest, OutputFormat
with ApiClient(configuration) as api_client:
api = BatchApi(api_client)
accepted = api.create_batch_v1_batch_post(
BatchRequest(
urls=[
"https://docs.geonode.com/docs/scraper-api/quick-start",
"https://docs.geonode.com/docs/scraper-api",
],
formats=[OutputFormat.MARKDOWN],
)
)
print("Batch job:", accepted.job_id)
print("Accepted URLs:", accepted.accepted_urls)
```
Response (live run):
```text
Batch job: 863e95fa-24c0-4638-aa09-a70235fdade1
Accepted URLs: 2
```
### get\_batch\_status\_v1\_batch\_job\_id\_get
Poll progress and read per-URL results (paginated with `page`, `page_size`).
```python
import time
from geonode_scraper_sdk import ApiClient, BatchApi, BatchRequest, OutputFormat
with ApiClient(configuration) as api_client:
api = BatchApi(api_client)
accepted = api.create_batch_v1_batch_post(
BatchRequest(
urls=["https://docs.geonode.com/docs/scraper-api/quick-start"],
formats=[OutputFormat.MARKDOWN],
)
)
while True:
status = api.get_batch_status_v1_batch_job_id_get(
job_id=accepted.job_id, page=1, page_size=10
)
print(
status.status,
status.completed_urls,
"/",
status.total_urls,
)
if str(status.status).lower().endswith("completed"):
break
time.sleep(3)
```
Response (live run):
```text
queued 0 / 1
processing 0 / 1
completed 1 / 1
```
### list\_batch\_jobs\_v1\_batch\_jobs\_get
List batch jobs with optional `status`, `start_date`, `end_date`, pagination.
```python
from geonode_scraper_sdk import ApiClient, BatchApi
with ApiClient(configuration) as api_client:
api = BatchApi(api_client)
page = api.list_batch_jobs_v1_batch_jobs_get(page=1, page_size=3)
for job in (page.jobs or [])[:3]:
print(job.job_id, job.status, job.completed_urls, "/", job.accepted_urls)
```
Response (live run):
```text
1255853d-eef7-45c4-9abe-7e3863535584 JobStatus.COMPLETED 1 / 1
863e95fa-24c0-4638-aa09-a70235fdade1 JobStatus.COMPLETED 2 / 2
d403ad87-66ce-4a48-823b-21e8ebbfb899 JobStatus.COMPLETED 2 / 2
```
### cancel\_batch\_v1\_batch\_job\_id\_delete
Stop scheduling new batch items. In-flight extractions drain.
```python
from geonode_scraper_sdk import ApiClient, BatchApi, BatchRequest, OutputFormat
with ApiClient(configuration) as api_client:
api = BatchApi(api_client)
accepted = api.create_batch_v1_batch_post(
BatchRequest(
urls=["https://docs.geonode.com/docs/scraper-api/quick-start"],
formats=[OutputFormat.MARKDOWN],
)
)
cancelled = api.cancel_batch_v1_batch_job_id_delete(accepted.job_id)
print("Cancelled:", cancelled.job_id, cancelled.status)
```
Response (live run):
```text
Cancelled: 883bd4ad-cb9c-49bc-99f5-d007fce5a217 JobStatus.CANCELLED
```
***
## Crawl API
`CrawlApi` (`/v1/crawl`)
**Methods:**
* `create_crawl_v1_crawl_post(crawl_request)`: start crawl from seed URL
* `get_crawl_status_v1_crawl_job_id_get(job_id, page, page_size)`: poll status / pages
* `list_crawl_jobs_v1_crawl_jobs_get(...)`: list past crawl jobs
* `cancel_crawl_v1_crawl_job_id_delete(job_id)`: cancel a running crawl
**`CrawlRequest` fields:**
| Field | Type | Notes |
| -------------------- | ---------------- | --------------------------------- |
| `url` | `str` | Required seed URL. |
| `depth` | `int` | BFS depth, 1-10. Default `2`. |
| `limit` | `int` | Max pages, 1-10000. Default `50`. |
| `same_domain_only` | `bool` | Default `True`. |
| `include_subdomains` | `bool` | Default `False`. |
| `formats` | `[OutputFormat]` | Optional per-page formats. |
| `render_js` | `bool` | Optional. |
| `proxy` | `ProxySettings` | Optional. |
| `wait_config` | `WaitConfig` | Optional. |
### create\_crawl\_v1\_crawl\_post
Start a crawl from a seed URL.
```python
from geonode_scraper_sdk import ApiClient, CrawlApi, CrawlRequest, OutputFormat
with ApiClient(configuration) as api_client:
api = CrawlApi(api_client)
accepted = api.create_crawl_v1_crawl_post(
CrawlRequest(
url="https://docs.geonode.com/docs/scraper-api",
depth=2,
limit=5,
formats=[OutputFormat.MARKDOWN],
same_domain_only=True,
)
)
print("Crawl job:", accepted.job_id)
print("Estimated pages:", accepted.estimated_pages)
```
Response (live run):
```text
Crawl job: efd86d85-f843-4b71-ad58-d5fec23a0ec3
Estimated pages: 5
```
### get\_crawl\_status\_v1\_crawl\_job\_id\_get
Poll crawl progress and read scraped pages (paginated).
```python
import time
from geonode_scraper_sdk import ApiClient, CrawlApi, CrawlRequest, OutputFormat
with ApiClient(configuration) as api_client:
api = CrawlApi(api_client)
accepted = api.create_crawl_v1_crawl_post(
CrawlRequest(
url="https://docs.geonode.com/docs/scraper-api",
depth=2,
limit=5,
formats=[OutputFormat.MARKDOWN],
same_domain_only=True,
)
)
while True:
status = api.get_crawl_status_v1_crawl_job_id_get(
job_id=accepted.job_id, page=1, page_size=10
)
print(
status.status,
status.completed_pages,
"/",
status.total_pages,
)
if str(status.status).lower().endswith("completed"):
break
time.sleep(4)
```
Response (live run):
```text
queued 0 / 5
processing 0 / 5
...
completed 5 / 5
```
### list\_crawl\_jobs\_v1\_crawl\_jobs\_get
List crawl jobs. Optional filters: `url`, `status`, `start_date`, `end_date`.
```python
from geonode_scraper_sdk import ApiClient, CrawlApi
with ApiClient(configuration) as api_client:
api = CrawlApi(api_client)
page = api.list_crawl_jobs_v1_crawl_jobs_get(page=1, page_size=2)
for job in (page.jobs or [])[:2]:
print(job.job_id, job.status, job.completed_pages, "/", job.total_pages)
```
Response (live run):
```text
34ebec40-5353-4933-ac11-6be09b85c448 JobStatus.COMPLETED 1 / 1
efd86d85-f843-4b71-ad58-d5fec23a0ec3 JobStatus.COMPLETED 5 / 5
```
### cancel\_crawl\_v1\_crawl\_job\_id\_delete
Cancel a queued or processing crawl.
```python
from geonode_scraper_sdk import ApiClient, CrawlApi, CrawlRequest, OutputFormat
with ApiClient(configuration) as api_client:
api = CrawlApi(api_client)
accepted = api.create_crawl_v1_crawl_post(
CrawlRequest(
url="https://docs.geonode.com/docs/scraper-api",
depth=1,
limit=50,
formats=[OutputFormat.MARKDOWN],
)
)
cancelled = api.cancel_crawl_v1_crawl_job_id_delete(accepted.job_id)
print("Cancelled:", cancelled.job_id, cancelled.status)
```
Response (live run):
```text
Cancelled: d8e5dec2-7450-4023-8fba-858243ac5339 JobStatus.CANCELLED
```
***
## Map API
`MapApi` (`/v1/map`)
**Methods:**
* `map_urls_v1_map_post(map_request)`: discover URLs under a base URL
* `list_map_jobs_v1_map_jobs_get(...)`: list past map jobs
* `get_map_job_v1_map_job_id_get(job_id)`: fetch a completed map job
**`MapRequest` fields:**
| Field | Type | Notes |
| ------------------------- | ------ | --------------------------------------- |
| `url` | `str` | Required base URL. |
| `include_subdomains` | `bool` | Widen discovery scope. Default `False`. |
| `ignore_query_parameters` | `bool` | Normalize URLs. Default `True`. |
| `search` | `str` | Optional path/url filter. |
### map\_urls\_v1\_map\_post
Discover URLs synchronously (returns links inline). Large sites may also create a persisted map job you can fetch later.
```python
from geonode_scraper_sdk import ApiClient, MapApi, MapRequest
with ApiClient(configuration) as api_client:
api = MapApi(api_client)
result = api.map_urls_v1_map_post(
MapRequest(url="https://docs.geonode.com/docs/scraper-api")
)
print("Link count:", len(result.links or []))
for link in (result.links or [])[:5]:
print(link.source, link.url)
```
Response (live run):
```text
Link count: 112
sitemap https://docs.geonode.com/docs/scraper-api
sitemap https://docs.geonode.com/docs/scraper-api/quick-start
sitemap https://docs.geonode.com/docs/scraper-api/additional-resources/choosing_scraper_api_plan
sitemap https://docs.geonode.com/docs/scraper-api/additional-resources/faq
sitemap https://docs.geonode.com/docs/scraper-api/additional-resources/pricing-and-requests
```
### list\_map\_jobs\_v1\_map\_jobs\_get
List past map jobs. Optional filters: `url`, `status`, `start_date`, `end_date`.
```python
from geonode_scraper_sdk import ApiClient, MapApi
with ApiClient(configuration) as api_client:
api = MapApi(api_client)
page = api.list_map_jobs_v1_map_jobs_get(page=1, page_size=2)
for job in (page.jobs or [])[:2]:
print(job.job_id, job.status, job.url)
```
Response (live run):
```text
b38bdc1c-2b53-4ccb-8947-6673cf5b53e0 JobStatus.COMPLETED https://docs.geonode.com/docs/scraper-api
bc074160-7d94-4a74-bb77-1e4f96f9c5cb JobStatus.COMPLETED https://docs.geonode.com/docs/scraper-api
```
### get\_map\_job\_v1\_map\_job\_id\_get
Retrieve full link list for a completed map job.
```python
from geonode_scraper_sdk import ApiClient, MapApi
JOB_ID = "b38bdc1c-2b53-4ccb-8947-6673cf5b53e0"
with ApiClient(configuration) as api_client:
api = MapApi(api_client)
detail = api.get_map_job_v1_map_job_id_get(JOB_ID)
links = detail.links or []
print("Job ID:", JOB_ID)
print("Link count:", len(links))
for link in links[:3]:
print(link.source, link.url)
```
Response (live run):
```text
Job ID: b38bdc1c-2b53-4ccb-8947-6673cf5b53e0
Link count: 112
sitemap https://docs.geonode.com/docs/scraper-api
sitemap https://docs.geonode.com/docs/scraper-api/quick-start
sitemap https://docs.geonode.com/docs/scraper-api/additional-resources/choosing_scraper_api_plan
```
***
## Search API
`SearchApi` (`/v1/search`)
**Methods:**
* `search_v1_search_post(search_request)`: run a search query
* `list_search_jobs_v1_search_jobs_get(...)`: list past search jobs
* `get_search_job_v1_search_job_id_get(job_id)`: fetch a completed search job
**`SearchRequest` fields:**
| Field | Type | Notes |
| ------------ | ----- | ----------------------------------------------- |
| `query` | `str` | Required search string. |
| `page` | `int` | Result page 1-20. Default `1`. |
| `safe` | `str` | `off` \| `moderate` \| `strict`. Default `off`. |
| `time_range` | `str` | Optional: `day`, `week`, `month`, `year`. |
| `locale` | `str` | Optional locale hint. |
### search\_v1\_search\_post
Run a search query. Returns results inline and a `job_id` for later lookup.
```python
from geonode_scraper_sdk import ApiClient, SearchApi, SearchRequest
with ApiClient(configuration) as api_client:
api = SearchApi(api_client)
result = api.search_v1_search_post(
SearchRequest(query="geonode scraper api")
)
print("Job ID:", result.job_id)
print("Hit count:", len(result.results or []))
for hit in (result.results or [])[:5]:
print(hit.position, hit.title, hit.url)
```
Response (live run):
```text
Job ID: 6aef3e93-f714-4eb5-ab7a-c0b00fee5dea
Hit count: 15
1 Wholesale proxies & web data infrastructure | Geonode https://geonode.com/
2 Geonode - PyPI https://pypi.org/user/Geonode/
3 pavel.s - PyPI https://pypi.org/user/pavel.s/
4 Geonode Documentation | Geonode https://docs.geonode.com/
5 GeoNode https://geonode.org/
```
### list\_search\_jobs\_v1\_search\_jobs\_get
List past search jobs. Optional filters: `query`, `status`, `start_date`, `end_date`.
```python
from geonode_scraper_sdk import ApiClient, SearchApi
with ApiClient(configuration) as api_client:
api = SearchApi(api_client)
page = api.list_search_jobs_v1_search_jobs_get(page=1, page_size=2)
for job in (page.jobs or [])[:2]:
print(job.job_id, job.query, job.status)
```
Response (live run):
```text
6aef3e93-f714-4eb5-ab7a-c0b00fee5dea geonode scraper api JobStatus.COMPLETED
0a0375f1-a36a-4c51-929d-5fa4b2425877 geonode scraper api JobStatus.COMPLETED
```
### get\_search\_job\_v1\_search\_job\_id\_get
Re-fetch full results for a search job.
```python
from geonode_scraper_sdk import ApiClient, SearchApi
JOB_ID = "6aef3e93-f714-4eb5-ab7a-c0b00fee5dea"
with ApiClient(configuration) as api_client:
api = SearchApi(api_client)
detail = api.get_search_job_v1_search_job_id_get(JOB_ID)
hits = detail.results or []
print("Job ID:", JOB_ID)
print("Hit count:", len(hits))
for hit in hits[:3]:
print(hit.position, hit.title, hit.url)
```
Response (live run):
```text
Job ID: 6aef3e93-f714-4eb5-ab7a-c0b00fee5dea
Hit count: 15
1 Wholesale proxies & web data infrastructure | Geonode https://geonode.com/
2 Geonode - PyPI https://pypi.org/user/Geonode/
3 pavel.s - PyPI https://pypi.org/user/pavel.s/
```
***
## Usage API
`UsageApi` (`/v1/usage`)
**Methods:**
* `get_concurrency_usage_v1_usage_concurrency_get()`: live concurrency in use vs plan limit
### get\_concurrency\_usage\_v1\_usage\_concurrency\_get
Check live work-concurrency slots against your plan limit.
```python
from geonode_scraper_sdk import ApiClient, UsageApi
with ApiClient(configuration) as api_client:
api = UsageApi(api_client)
usage = api.get_concurrency_usage_v1_usage_concurrency_get()
print(usage.work_concurrency_in_use, usage.work_concurrency_limit)
```
Response (live run):
```text
0 50
```
***
## Statistics API
`StatisticsApi` (`/v1/statistics`)
**Methods:**
* `get_statistics_v1_statistics_get(...)`: extraction counts, tokens, success rate
### get\_statistics\_v1\_statistics\_get
Historical extraction counts, token usage, success rate.
```python
from geonode_scraper_sdk import ApiClient, StatisticsApi
with ApiClient(configuration) as api_client:
api = StatisticsApi(api_client)
stats = api.get_statistics_v1_statistics_get()
print("Extraction count:", stats.extraction_count)
print("Success rate:", stats.success_rate)
print("Recent token days:", len(stats.tokens_used or []))
```
Response (live run):
```text
Extraction count: 4378
Success rate: 0.95112
Recent token days: 7
```
***
## System API
`SystemApi` (`/health`)
**Methods:**
* `health_check_health_get()`: service health check
### health\_check\_health\_get
Confirm the Scraper API is up.
```python
from geonode_scraper_sdk import ApiClient, SystemApi
with ApiClient(configuration) as api_client:
api = SystemApi(api_client)
health = api.health_check_health_get()
print(health.service, health.status, health.version)
```
Response (live run):
```text
Scraper API HealthStatus.OK 0.1.0
```
***
## Webhooks API
`WebhooksApi` (`/v1/webhooks`)
Event types: `extract_completed`, `batch_completed`, `crawl_completed`.
**Methods:**
* `create_webhook_v1_webhooks_post(webhook_create)`: register a webhook
* `list_webhooks_v1_webhooks_get(...)`: list webhooks
* `get_webhook_v1_webhooks_webhook_id_get(webhook_id)`: get one webhook
* `update_webhook_v1_webhooks_webhook_id_patch(webhook_id, webhook_update)`: update
* `delete_webhook_v1_webhooks_webhook_id_delete(webhook_id)`: delete
* `list_deliveries_v1_webhooks_webhook_id_deliveries_get(...)`: delivery history
* `rotate_secret_v1_webhooks_webhook_id_rotate_secret_post(webhook_id)`: rotate signing secret
**`WebhookCreate` fields:**
| Field | Type | Notes |
| ------------- | ------------------ | ------------------------------------------------------------- |
| `url` | `str` | Required callback URL. |
| `event_type` | `WebhookEventType` | `extract_completed`, `batch_completed`, or `crawl_completed`. |
| `description` | `str` | Optional label. |
### create\_webhook\_v1\_webhooks\_post
```python
from geonode_scraper_sdk import (
ApiClient,
WebhookCreate,
WebhookEventType,
WebhooksApi,
)
with ApiClient(configuration) as api_client:
api = WebhooksApi(api_client)
created = api.create_webhook_v1_webhooks_post(
WebhookCreate(
url="https://example.com/webhook",
event_type=WebhookEventType.EXTRACT_COMPLETED,
description="sdk guide demo",
)
)
print("Created:", created.id, created.url, created.event_type)
```
Response (live run):
```text
Created: b9ea9375-90a0-47cb-bb37-48c4c0804925 https://example.com/webhook WebhookEventType.EXTRACT_COMPLETED
```
### list\_webhooks\_v1\_webhooks\_get
```python
from geonode_scraper_sdk import ApiClient, WebhooksApi
with ApiClient(configuration) as api_client:
api = WebhooksApi(api_client)
page = api.list_webhooks_v1_webhooks_get(page=1, page_size=5)
items = getattr(page, "items", None) or []
print("Webhook count:", len(items))
for wh in items:
print(wh.id, wh.url, wh.event_type)
```
Response (live run):
```text
Webhook count: 0
```
### get\_webhook\_v1\_webhooks\_webhook\_id\_get
```python
from geonode_scraper_sdk import ApiClient, WebhooksApi
WEBHOOK_ID = "b9ea9375-90a0-47cb-bb37-48c4c0804925"
with ApiClient(configuration) as api_client:
api = WebhooksApi(api_client)
wh = api.get_webhook_v1_webhooks_webhook_id_get(WEBHOOK_ID)
print(wh.url, wh.event_type, wh.is_active)
```
Response (live run):
```text
https://example.com/webhook WebhookEventType.EXTRACT_COMPLETED True
```
### update\_webhook\_v1\_webhooks\_webhook\_id\_patch
```python
from geonode_scraper_sdk import ApiClient, WebhookUpdate, WebhooksApi
WEBHOOK_ID = "b9ea9375-90a0-47cb-bb37-48c4c0804925"
with ApiClient(configuration) as api_client:
api = WebhooksApi(api_client)
updated = api.update_webhook_v1_webhooks_webhook_id_patch(
WEBHOOK_ID,
WebhookUpdate(description="updated by sdk guide"),
)
print("Description:", updated.description)
```
Response (live run):
```text
Description: updated by sdk guide
```
### rotate\_secret\_v1\_webhooks\_webhook\_id\_rotate\_secret\_post
```python
from geonode_scraper_sdk import ApiClient, WebhooksApi
WEBHOOK_ID = "b9ea9375-90a0-47cb-bb37-48c4c0804925"
with ApiClient(configuration) as api_client:
api = WebhooksApi(api_client)
rotated = api.rotate_secret_v1_webhooks_webhook_id_rotate_secret_post(WEBHOOK_ID)
print("New secret length:", len(rotated.secret or ""))
```
Response (live run):
```text
New secret length: 64
```
### list\_deliveries\_v1\_webhooks\_webhook\_id\_deliveries\_get
```python
from geonode_scraper_sdk import ApiClient, WebhooksApi
WEBHOOK_ID = "b9ea9375-90a0-47cb-bb37-48c4c0804925"
with ApiClient(configuration) as api_client:
api = WebhooksApi(api_client)
page = api.list_deliveries_v1_webhooks_webhook_id_deliveries_get(
WEBHOOK_ID, page=1, page_size=5
)
items = getattr(page, "items", None) or []
print("Delivery count:", len(items))
```
Response (live run):
```text
Delivery count: 0
```
### delete\_webhook\_v1\_webhooks\_webhook\_id\_delete
```python
from geonode_scraper_sdk import ApiClient, WebhooksApi
WEBHOOK_ID = "b9ea9375-90a0-47cb-bb37-48c4c0804925"
with ApiClient(configuration) as api_client:
api = WebhooksApi(api_client)
api.delete_webhook_v1_webhooks_webhook_id_delete(WEBHOOK_ID)
print("Deleted:", WEBHOOK_ID)
```
Response (live run):
```text
Deleted: b9ea9375-90a0-47cb-bb37-48c4c0804925
```
***
## Error handling
Non-2xx HTTP responses raise `ApiException` with `status`, `body`, and parsed `data`. Invalid models can fail before the HTTP call (Pydantic `ValidationError`).
**Validation error** (empty URL):
```python
from geonode_scraper_sdk import ApiClient, ExtractRequest, ExtractionApi
with ApiClient(configuration) as api_client:
api = ExtractionApi(api_client)
try:
api.extract_v1_extract_post(ExtractRequest(url=""))
except Exception as exc:
print(type(exc).__name__, exc)
```
Response (live run):
```text
ValidationError 1 validation error for ExtractRequest
url
String should have at least 1 character [type=string_too_short, input_value='', input_type=str]
```
**API error** (fake job ID):
```python
from geonode_scraper_sdk import ApiClient, ApiException, ExtractionApi
with ApiClient(configuration) as api_client:
api = ExtractionApi(api_client)
try:
api.get_job_result_v1_extract_job_id_get(
"00000000-0000-0000-0000-000000000000"
)
except ApiException as exc:
print(exc.status)
print(exc.body)
```
Response (live run):
```text
404
{"code":"NOT_FOUND","message":"Job 00000000-0000-0000-0000-000000000000 not found","correlation_id":"f430c05f-5aea-4e98-bcf2-34fabb743f8c","retryable":false,"details":null}
```
For package versions and `*_with_http_info()` variants, see [geonode-scraper-sdk on PyPI](https://pypi.org/project/geonode-scraper-sdk/).
# Tools Core (/docs/scraper-api/developer-guides/tools-core)
{/* DRAFT: do not publish / do not add to meta.json until approved */}
`geonode-scraper-tools-core` gives you named scraper operations and JSON-friendly dictionaries. Call them via `ScraperToolService`, or select a subset with `get_operations()`. LangChain and CrewAI packages use this layer under the hood.
Responses below are from live runs against `https://scraper.geonode.io`.
## Setup
```bash
pip install geonode-scraper-tools-core python-dotenv
```
Create a `.env` file:
```bash
GEONODE_SCRAPER_API_KEY=your_geonode_key
SCRAPER_API_BASE_URL=https://scraper.geonode.io
```
```python
import os
from dotenv import load_dotenv
from geonode_scraper_tools_core import ScraperToolSettings, ScraperToolService
load_dotenv()
settings = ScraperToolSettings(
host=os.environ.get("SCRAPER_API_BASE_URL", "https://scraper.geonode.io"),
api_key=os.environ["GEONODE_SCRAPER_API_KEY"],
)
service = ScraperToolService(settings)
```
Use `https://scraper.geonode.io` for production.
### ScraperToolSettings fields
| Field | Type | Notes |
| ----------------------- | ------- | -------------------------------------- |
| `host` | `str` | Required. API base URL. |
| `api_key` | `str` | Required. Geonode Scraper API key. |
| `verify_ssl` | `bool` | Default `True`. |
| `request_timeout` | timeout | Optional HTTP timeout. |
| `max_retries` | `int` | Default `0`. |
| `retry_backoff_seconds` | `float` | Default `1.0`. |
| `poll_interval_seconds` | `float` | Default `3.0` (used by wait helpers). |
| `poll_timeout_seconds` | `float` | Default `60.0` (used by wait helpers). |
### Response shape
Every service method returns a dict like:
```python
{
"ok": True,
"operation": "extract",
"attempts": 1,
"result": { ... }, # real payload
}
```
Wait helpers also include `poll_attempts`. Read the payload from `response["result"]`.
## Operations overview
| Group | Operations |
| -------------- | ----------------------------------------------------------------------- |
| **Extraction** | `extract`, `get_job_result`, `wait_for_job`, `list_jobs` |
| **Batch** | `create_batch`, `get_batch_status`, `wait_for_batch`, `list_batch_jobs` |
| **Crawl** | `create_crawl`, `get_crawl_status`, `wait_for_crawl`, `list_crawl_jobs` |
| **Map** | `map_urls`, `list_map_jobs`, `get_map_job` |
| **Search** | `search`, `list_search_jobs`, `get_search_job` |
| **Account** | `get_statistics`, `get_concurrency_usage`, `health_check` |
`wait_for_job`, `wait_for_batch`, and `wait_for_crawl` poll until a job finishes, so you do not write poll loops yourself.
***
## Extraction
**Operations:** `extract`, `get_job_result`, `wait_for_job`, `list_jobs`
### extract (sync)
Scrape one URL and wait for content in the same call.
**Main inputs:** `url`, `formats` (`markdown` / `html`), `processing_mode` (`sync` / `async`), `render_js`, `proxy_country`, `proxy_type`, `headers`, `wait_for`, `wait_timeout`, `wait_until`
```python
response = service.extract(
url="https://docs.geonode.com/docs/scraper-api/quick-start",
formats=["markdown"],
processing_mode="sync",
)
result = response["result"]
markdown = (result.get("data") or {}).get("markdown") or ""
print("ok:", response["ok"])
print("tokens:", result.get("tokens_charged"))
print("markdown_len:", len(markdown))
print("preview:", markdown[:200])
```
Response (live run):
```text
ok: True
tokens: 1
markdown_len: 15871
preview: ---
canonical: https://docs.geonode.com/docs/scraper-api/quick-start
meta-description: Get your API key, authenticate requests, and choose the right API for your use case.
...
```
### extract (async)
Same method with `processing_mode="async"`. Returns a `job_id` immediately.
```python
response = service.extract(
url="https://docs.geonode.com/docs/scraper-api/quick-start",
formats=["markdown"],
processing_mode="async",
)
result = response["result"]
print("job_id:", result["job_id"])
print("status:", result["status"])
```
Response (live run):
```text
job_id: 8a2bbbd5-7034-412b-acdf-e30226ae32b6
status: queued
```
### wait\_for\_job
Poll an async extract job until it finishes (or times out).
**Inputs:** `job_id`, optional `timeout_seconds`, `poll_interval_seconds`
```python
response = service.wait_for_job(
job_id="8a2bbbd5-7034-412b-acdf-e30226ae32b6",
timeout_seconds=120,
)
result = response["result"]
markdown = (result.get("data") or {}).get("markdown") or ""
print("status:", result["status"])
print("markdown_len:", len(markdown))
print("poll_attempts:", response.get("poll_attempts"))
```
Response (live run):
```text
status: completed
markdown_len: 15871
poll_attempts: 2
```
### get\_job\_result
Fetch the current state or final result for one extract job (no waiting).
**Inputs:** `job_id`
```python
response = service.get_job_result(
job_id="8a2bbbd5-7034-412b-acdf-e30226ae32b6",
)
result = response["result"]
print("status:", result["status"])
print("tokens:", result.get("tokens_charged"))
```
Response (live run):
```text
status: completed
tokens: 1
```
### list\_jobs
List past extract jobs.
**Inputs:** optional `job_id`, `url`, `status`, `output`, `start_date`, `end_date`, `page`, `page_size`
```python
response = service.list_jobs(page=1, page_size=3)
result = response["result"]
print("page:", result["page"], "page_size:", result["page_size"])
for job in result.get("jobs") or []:
print(job["job_id"], job["status"], job.get("url"))
```
Response (live run):
```text
page: 1 page_size: 3
752b8599-5915-441c-bc1f-b9fb0d938f72 completed ...
0bcd424e-b5ab-4b85-9cbd-5b44f7ed433c completed http://example.com/
0dbb126c-4331-4b4b-9d55-0c380ba69ae7 completed http://example.com/
```
***
## Batch
**Operations:** `create_batch`, `get_batch_status`, `wait_for_batch`, `list_batch_jobs`
### create\_batch
Submit many URLs as one job.
**Inputs:** `urls`, `formats`, optional `render_js`, `proxy_country`, `proxy_type`, `headers`
```python
response = service.create_batch(
urls=[
"https://docs.geonode.com/docs/scraper-api/quick-start",
"https://docs.geonode.com/docs/scraper-api",
],
formats=["markdown"],
)
result = response["result"]
print("job_id:", result["job_id"])
print("accepted_urls:", result["accepted_urls"])
print("status:", result["status"])
```
Response (live run):
```text
job_id: 55b11790-5c7f-4945-bc13-bd9a365a1835
accepted_urls: 2
status: queued
```
### get\_batch\_status
Poll progress and partial results (paginated).
**Inputs:** `job_id`, `page`, `page_size`
```python
response = service.get_batch_status(
job_id="55b11790-5c7f-4945-bc13-bd9a365a1835",
page=1,
page_size=10,
)
result = response["result"]
print(
result["status"],
result["completed_urls"],
"/",
result["total_urls"],
)
```
Response (live run, mid-job):
```text
processing 0 / 2
```
### wait\_for\_batch
Poll until the batch finishes.
**Inputs:** `job_id`, optional `timeout_seconds`, `poll_interval_seconds`
```python
response = service.wait_for_batch(
job_id="55b11790-5c7f-4945-bc13-bd9a365a1835",
timeout_seconds=120,
)
result = response["result"]
print(
result["status"],
result["completed_urls"],
"/",
result["total_urls"],
)
print("poll_attempts:", response.get("poll_attempts"))
```
Response (live run):
```text
completed 2 / 2
poll_attempts: 3
```
### list\_batch\_jobs
**Inputs:** optional `status`, `start_date`, `end_date`, `page`, `page_size`
```python
response = service.list_batch_jobs(page=1, page_size=3)
for job in (response["result"].get("jobs") or []):
print(
job["job_id"],
job["status"],
job["completed_urls"],
"/",
job["accepted_urls"],
)
```
Response (live run):
```text
55b11790-5c7f-4945-bc13-bd9a365a1835 completed 2 / 2
d8e92d3a-c939-4ec6-b133-3e04c6de71d1 completed 2 / 2
883bd4ad-cb9c-49bc-99f5-d007fce5a217 completed 0 / 1
```
***
## Crawl
**Operations:** `create_crawl`, `get_crawl_status`, `wait_for_crawl`, `list_crawl_jobs`
### create\_crawl
Start from a seed URL.
**Inputs:** `url`, `depth`, `limit`, `formats`, `same_domain_only`, `include_subdomains`, optional `render_js`, `proxy_country`, `proxy_type`
```python
response = service.create_crawl(
url="https://docs.geonode.com/docs/scraper-api",
depth=2,
limit=3,
formats=["markdown"],
same_domain_only=True,
)
result = response["result"]
print("job_id:", result["job_id"])
print("estimated_pages:", result["estimated_pages"])
print("status:", result["status"])
```
Response (live run):
```text
job_id: 9f447ced-4aa3-45a3-b9af-421f751b8cec
estimated_pages: 3
status: queued
```
### get\_crawl\_status
**Inputs:** `job_id`, `page`, `page_size`
```python
response = service.get_crawl_status(
job_id="9f447ced-4aa3-45a3-b9af-421f751b8cec",
page=1,
page_size=10,
)
result = response["result"]
print(
result["status"],
result.get("completed_pages"),
"/",
result.get("total_pages"),
)
```
Response (live run, early poll):
```text
processing 0 / 1
```
### wait\_for\_crawl
**Inputs:** `job_id`, optional `timeout_seconds`, `poll_interval_seconds`
```python
response = service.wait_for_crawl(
job_id="9f447ced-4aa3-45a3-b9af-421f751b8cec",
timeout_seconds=180,
)
result = response["result"]
print(
result["status"],
result["completed_pages"],
"/",
result["total_pages"],
)
print("poll_attempts:", response.get("poll_attempts"))
```
Response (live run):
```text
completed 3 / 3
poll_attempts: 19
```
### list\_crawl\_jobs
**Inputs:** optional `url`, `status`, `start_date`, `end_date`, `page`, `page_size`
```python
response = service.list_crawl_jobs(page=1, page_size=2)
for job in (response["result"].get("jobs") or []):
print(
job["job_id"],
job["status"],
job["completed_pages"],
"/",
job["total_pages"],
)
```
Response (live run):
```text
9f447ced-4aa3-45a3-b9af-421f751b8cec completed 3 / 3
d8e5dec2-7450-4023-8fba-858243ac5339 completed 0 / 1
```
***
## Map
**Operations:** `map_urls`, `list_map_jobs`, `get_map_job`
### map\_urls
Discover URLs under a base URL (sitemap + HTML links). Does not scrape page content.
**Inputs:** `url`, optional `search`, `include_subdomains`, `ignore_query_parameters`
```python
response = service.map_urls(
url="https://docs.geonode.com/docs/scraper-api",
)
result = response["result"]
links = result.get("links") or []
print("link_count:", result.get("links_count") or len(links))
for link in links[:5]:
print(link.get("source"), link.get("url"))
```
Response (live run):
```text
link_count: 112
sitemap https://docs.geonode.com/docs/scraper-api
sitemap https://docs.geonode.com/docs/scraper-api/quick-start
sitemap https://docs.geonode.com/docs/scraper-api/additional-resources/choosing_scraper_api_plan
sitemap https://docs.geonode.com/docs/scraper-api/additional-resources/faq
sitemap https://docs.geonode.com/docs/scraper-api/additional-resources/pricing-and-requests
```
### list\_map\_jobs
**Inputs:** optional `url`, `status`, `start_date`, `end_date`, `page`, `page_size`
```python
response = service.list_map_jobs(page=1, page_size=2)
for job in (response["result"].get("jobs") or []):
print(job["job_id"], job["status"], job.get("links_count"), job.get("url"))
```
Response (live run):
```text
c7cb6e40-ddfa-416a-b255-00dfdfd717bc completed 112 https://docs.geonode.com/docs/scraper-api
b38bdc1c-2b53-4ccb-8947-6673cf5b53e0 completed 112 https://docs.geonode.com/docs/scraper-api
```
### get\_map\_job
**Inputs:** `job_id`
```python
response = service.get_map_job(
job_id="c7cb6e40-ddfa-416a-b255-00dfdfd717bc",
)
result = response["result"]
print("status:", result["status"])
print("link_count:", result.get("links_count"))
for link in (result.get("links") or [])[:3]:
print(link.get("source"), link.get("url"))
```
Response (live run):
```text
status: completed
link_count: 112
sitemap https://docs.geonode.com/docs/scraper-api
sitemap https://docs.geonode.com/docs/scraper-api/quick-start
sitemap https://docs.geonode.com/docs/scraper-api/additional-resources/choosing_scraper_api_plan
```
***
## Search
**Operations:** `search`, `list_search_jobs`, `get_search_job`
### search
Run a web search and get ranked hits (also returns a `job_id`).
**Inputs:** `query`, optional `locale`, `page`, `safe` (`off` / `moderate` / `strict`), `time_range` (`day` / `week` / `month` / `year`)
```python
response = service.search(query="geonode scraper api")
result = response["result"]
print("job_id:", result["job_id"])
print("hit_count:", result.get("results_count") or len(result.get("results") or []))
for hit in (result.get("results") or [])[:5]:
print(hit["position"], hit["title"], hit["url"])
```
Response (live run):
```text
job_id: 29809a61-49b6-4bf6-925b-aa32dd91e441
hit_count: 15
1 Wholesale proxies & web data infrastructure | Geonode https://geonode.com/
2 Geonode - PyPI https://pypi.org/user/Geonode/
3 pavel.s - PyPI https://pypi.org/user/pavel.s/
4 Geonode Documentation | Geonode https://docs.geonode.com/
5 GeoNode https://geonode.org/
```
### list\_search\_jobs
**Inputs:** optional `query`, `status`, `start_date`, `end_date`, `page`, `page_size`
```python
response = service.list_search_jobs(page=1, page_size=2)
for job in (response["result"].get("jobs") or []):
print(job["job_id"], job.get("query"), job["status"], job.get("results_count"))
```
Response (live run):
```text
29809a61-49b6-4bf6-925b-aa32dd91e441 geonode scraper api completed 15
6aef3e93-f714-4eb5-ab7a-c0b00fee5dea geonode scraper api completed 15
```
### get\_search\_job
**Inputs:** `job_id`
```python
response = service.get_search_job(
job_id="29809a61-49b6-4bf6-925b-aa32dd91e441",
)
result = response["result"]
print("status:", result["status"])
print("hit_count:", result.get("results_count"))
for hit in (result.get("results") or [])[:3]:
print(hit["position"], hit["title"], hit["url"])
```
Response (live run):
```text
status: completed
hit_count: 15
1 Wholesale proxies & web data infrastructure | Geonode https://geonode.com/
2 Geonode - PyPI https://pypi.org/user/Geonode/
3 pavel.s - PyPI https://pypi.org/user/pavel.s/
```
***
## Statistics, usage, and health
**Operations:** `get_statistics`, `get_concurrency_usage`, `health_check`
### get\_statistics
**Inputs:** optional `start_date`, `end_date`
```python
response = service.get_statistics()
result = response["result"]
print("extraction_count:", result.get("extraction_count"))
print("success_rate:", result.get("success_rate"))
```
Response (live run):
```text
extraction_count: 4383
success_rate: 0.95117
```
### get\_concurrency\_usage
No inputs. Checks live concurrency slots vs your plan limit.
```python
response = service.get_concurrency_usage()
result = response["result"]
print(
result["work_concurrency_in_use"],
"/",
result["work_concurrency_limit"],
)
```
Response (live run):
```text
0 / 50
```
### health\_check
No inputs. Confirms the Scraper API is up.
```python
response = service.health_check()
result = response["result"]
print(result.get("service"), result.get("status"), result.get("version"))
```
Response (live run):
```text
Scraper API ok 0.1.0
```
***
## Selecting a subset of operations
Use `get_operations()` when you only want some tools (for example before wiring into an agent framework).
```python
from geonode_scraper_tools_core import get_operations
ops = get_operations(["extract", "map_urls", "create_batch", "wait_for_batch"])
print([op.key for op in ops])
```
Response:
```text
['extract', 'map_urls', 'create_batch', 'wait_for_batch']
```
`OPERATIONS` is the full registry (all 21). Each entry has `key`, `tool_name`, `description`, `args_schema`, and `service_method`.
***
## Used by LangChain and CrewAI
Most agent users install:
* `geonode-scraper-langchain`
* `geonode-scraper-crewai`
Those packages wrap this core service. Use tools-core directly when you want named operations and plain dicts in your own Python code, without an agent framework.
For package versions and changelog, see [geonode-scraper-tools-core on PyPI](https://pypi.org/project/geonode-scraper-tools-core/).
# Before You Start (/docs/scraper-api/getting-started/00_before_you_start)
import { Step, Steps } from "fumadocs-ui/components/steps";
This guide covers the basic requirements needed to follow the Extraction guides and ensure your account is set up and ready to make requests.
## Prerequisites
#### Create a Geonode Account
Sign up for a Geonode account if you do not already have one.
#### Get Access to the Scraper API
Make sure your account has access to the Scraper API.
#### Generate an API Key
1. Go to `https://app.geonode.com/scraper-api`.
2. Click `Get Code`.
3. The `API & Integrations` modal will open.
4. Click `Copy API Key` to copy your API key.
## Authentication
Include your API key in the `X-Api-Key` request header.
```bash
X-Api-Key: YOUR_API_KEY
```
Requests without a valid API key will be rejected.
## Base URL
All Extraction API endpoints are available through the following base URL:
```bash
https://scraper.geonode.io
```
## Next Steps
Now that your account is set up, continue to **Understanding Extraction** to learn how the Extraction API works and when to use its different features.
# Build an Amazon Product Snapshot from Search Results (/docs/scraper-api/real-world/build_amazon_product_snapshot)
This guide walks you through a real marketplace workflow with the Geonode Scraper API: **as if we are building it together**.
You want a **small competitive product snapshot** on Amazon.com: what shows up for a keyword, at what price, with ratings. You **already know the site**. You do **not** need the Geonode Search API to find Amazon.
By the end, you will understand **when to use Extract vs Batch** on a hard, JavaScript-heavy marketplace, why residential geo targeting matters, and how to keep a demo bounded when pages block or fail.
Companion example code lives in `geonode-scraper-examples/amazon-product-snapshot/`.
This is a **small teaching demo** (one search page, capped ASINs). Amazon frequently blocks or challenges automated traffic. Follow Amazon’s terms and applicable law. This guide does **not** teach bypassing captchas, logging into accounts, or large-scale scraping. For production hardness, contact support.
## Use case
Imagine you are on a pricing, marketplace, or assortment team. You need a short list of products for one keyword:
* ASIN and product URL
* Title
* Price (USD when present)
* Star rating and review count when present
Starting point: a **known Amazon.com search URL**, for example:
```text
https://www.amazon.com/s?k=logitech+mx+master
```
## What we are going to build
A master JSON snapshot. Each product looks roughly like this (shortened):
```json
{
"asin": "B0FB21526X",
"url": "https://www.amazon.com/dp/B0FB21526X",
"title": "Logitech MX Master 3S Bluetooth Wireless Mouse...",
"price": 89.99,
"currency": "USD",
"rating": 4.6,
"reviews_count": 1200
}
```
| Metric | Demo target |
| ---------- | -------------: |
| SERP pages | 1 |
| Max ASINs | 10 |
| Proxy | US residential |
| JS render | on |
## The plan (what you should expect)
```text
1. Extract → Amazon search (SERP) Markdown/HTML
2. Parse → ASINs (local; fallback list if SERP empty)
3. Batch → /dp/{ASIN} product pages
4. Parse → title, price, rating (local)
5. Merge → amazon_products.json
```
POST /v1/extract"] --> B["2. Parse ASINs local"]
end
subgraph products["Product pages"]
direction LR
C["3. Batch PDPs POST /v1/batch"] --> D["4. Parse local"]
D --> E["5. Merge"]
E --> F["amazon_products.json"]
end
discovery --> C
classDef api fill:#e8f5f3,stroke:#054D4D,color:#033333
classDef local fill:#f5f5f5,stroke:#666,color:#222
classDef result fill:#fff8e6,stroke:#b8860b,color:#333
class A,C api
class B,D,E local
class F result
`}
/>
| API | Why we use it here |
| -------------------------------------------------------------------------- | ------------------------------------------- |
| [Extract](/docs/scraper-api/guides/extraction/01_understanding_extraction) | One known SERP URL; JS + residential + wait |
| [Batch](/docs/scraper-api/guides/batch/00_understanding_batch) | Many `/dp/` URLs with the same settings |
| Local parse | ASINs and product fields |
**Geonode Search is not used.** You already know `amazon.com`.
**Map is not used.** You are not inventorying a sitemap; you start from one search URL.
**Crawl is not used.** You do not want an unbounded walk of Amazon.
Compare with other real-world guides:
| Guide | Difference |
| ------------------------------------------------------------------------------------------------- | ---------------------------------------------- |
| [Retail category price catalog](/docs/scraper-api/real-world/build_retail_category_price_catalog) | Sitemap + PLP pagination on a lighter retailer |
| [Docs knowledge corpus](/docs/scraper-api/real-world/build_docs_knowledge_corpus) | Crawl public docs |
| [Basic Auth headers](/docs/scraper-api/real-world/build_authenticated_basic_auth) | Authenticated Extract, not marketplace SERP |
***
## Before you start
You need:
* A Geonode API key
* Python 3.9+, `requests`, `python-dotenv`
```bash
export GEONODE_SCRAPER_API_KEY="YOUR_API_KEY"
```
Optional:
```bash
export AMAZON_SEARCH_URL="https://www.amazon.com/s?k=logitech+mx+master"
export AMAZON_MAX_ASINS="10"
```
***
## Step 1: Extract the search results page
### What we need
HTML/Markdown of one Amazon SERP so we can collect ASINs.
### What we will do
`POST /v1/extract` with:
* `render_js: true`
* `proxy: { "country": "US", "type": "residential" }`
* `wait_config` with `wait_until: "networkidle"` (Amazon is a heavy SPA)
### What you should expect
A large Markdown/HTML blob with `/dp/{ASIN}` links, or a soft block / empty shell. Save raw output either way so you can debug.
```python
import os
import requests
api_key = os.environ["GEONODE_SCRAPER_API_KEY"]
url = os.environ.get(
"AMAZON_SEARCH_URL",
"https://www.amazon.com/s?k=logitech+mx+master",
)
response = requests.post(
"https://scraper.geonode.io/v1/extract",
headers={
"X-Api-Key": api_key,
"Content-Type": "application/json",
},
json={
"url": url,
"formats": ["markdown", "html"],
"render_js": True,
"processing_mode": "sync",
"proxy": {"country": "US", "type": "residential"},
"wait_config": {
"wait_until": "networkidle",
"wait_timeout": 20000,
},
},
)
# Save markdown / html for ASIN parsing
```
### Takeaway
**Hard marketplaces need JS + residential geo + patience.** Retry on 429/503/504. If the SERP is blocked, do not pretend Batch will invent ASINs.
***
## Step 2: Parse ASINs locally
### What we need
A capped list of product URLs: `https://www.amazon.com/dp/{ASIN}`.
### What we will do
Regex over SERP Markdown/HTML for `/dp/`, `/gp/product/`, and similar. Dedupe. Cap at `AMAZON_MAX_ASINS` (demo default **10**).
If parsing finds nothing, the companion repo falls back to a checked-in `fallback_asins.json` so later Batch/parse steps still teach the pipeline.
### What you should expect
`asins.json` + `product_urls.json`.
### Takeaway
**Discovery is local once Extract returns page content.** Keep the ASIN cap small for demos and cost control.
***
## Step 3: Batch product detail pages
### What we need
Title/price/rating material from each `/dp/` URL.
### What we will do
`POST /v1/batch` with the same `render_js`, US residential proxy, and wait settings. Poll until completed. Save per-ASIN Markdown and HTML.
### What you should expect
Some URLs succeed, some fail, Amazon demos often show partial completion (for example 6/10). That is normal; parse what you have.
```python
# urls = ["https://www.amazon.com/dp/B0…", ...]
response = requests.post(
"https://scraper.geonode.io/v1/batch",
headers={
"X-Api-Key": api_key,
"Content-Type": "application/json",
},
json={
"urls": urls,
"ignore_invalid_urls": True,
"formats": ["markdown", "html"],
"render_js": True,
"proxy": {"country": "US", "type": "residential"},
"wait_config": {
"wait_until": "networkidle",
"wait_timeout": 20000,
},
},
)
job_id = response.json()["job_id"]
# Poll GET /v1/batch/{job_id} until completed
```
### Takeaway
**Batch reuses Extract settings across a URL list.** It does not guarantee every Amazon PDP returns price HTML.
***
## Step 4–5: Parse and merge
### What we need
The business artifact: `amazon_products.json`.
### What we will do
Locally parse title (including `og:title`), price patterns / `a-offscreen` price HTML, rating, and review counts. Merge into a slim master file with counts and optional SERP meta.
### What you should expect
A spreadsheet-ready list. Some products may have `price: null` when the PDP shell loaded without a clear price node.
### Takeaway
Same pattern as other real-world guides: **Geonode fetches; your merge step is the deliverable.**
***
## Why this API mix worked
| Tool | Role |
| ----------- | ----------------------------------------- |
| Extract | One SERP you already know |
| Batch | Many PDPs, same browser/proxy config |
| Map / Crawl | Wrong shape for a single keyword snapshot |
### When not to use this pattern
* You only have known ASINs → skip SERP; start at Batch
* You need unbounded site coverage → still not Crawl-on-Amazon for demos
* You need account pages → not this guide (and often not feasible via simple headers)
***
## Cost and request usage
On [request-based pricing](/docs/scraper-api/additional-resources/request_based_pricing):
| Stage | Rough volume | Notes |
| ------------- | -----------------------: | ----------------------------------- |
| SERP Extract | 1 | JS + residential |
| Batch PDPs | up to `AMAZON_MAX_ASINS` | Partial failures still consume work |
| Parse / merge | 0 | Local |
Keep the ASIN cap small while teaching.
***
## Limitations and good practice
* Amazon may return captchas, empty shells, or timeouts, expect flaky runs.
* Prefer first-party / permitted use cases; respect robots and terms.
* Do not store or publish customer account data.
* Multi-page SERP pagination (`page=`) is on you, same idea as retail PLP pagination in the [BBB catalog guide](/docs/scraper-api/real-world/build_retail_category_price_catalog).
* Production monitoring usually needs retries, alerting, and support for harder setups.
***
## Recap
1. **Extract** one Amazon search URL with JS + US residential.
2. **Parse** ASINs (cap the list; keep a fallback for demos).
3. **Batch** `/dp/` pages with the same settings.
4. **Parse / merge** into `amazon_products.json`.
That is the Amazon teaching story next to the lighter retail catalog: **same Extract → Batch shape, harder site, smaller demo, honest limits.**
# Extract a Basic Auth Page with Authorization Headers (/docs/scraper-api/real-world/build_authenticated_basic_auth)
This guide walks you through an **authenticated** Extract workflow with the Geonode Scraper API: **as if we are building it together**.
Geonode does **not** open a login UI. You send site credentials in request `headers`. Here we use **HTTP Basic Auth** — the same pattern many staging sites and internal tools use.
By the end you will understand:
* `X-Api-Key` → authenticates you to **Geonode**
* `headers.Authorization: Basic …` → authenticates the request to the **target site**
Companion example code lives in `geonode-scraper-examples/basic-auth-protected-page/`.
## Use case
You need Markdown from a URL that returns **401 Unauthorized** without credentials (staging wiki, internal tool, password-protected doc).
For a **reproducible public demo**, this guide uses well-known Basic Auth practice URLs (not a production customer site):
```text
https://the-internet.herokuapp.com/basic_auth
```
Public demo credentials: username `admin`, password `admin`.
The **API pattern** is what you reuse on your own Basic Auth staging or internal pages.
## What we are going to build
A small auth corpus plus before/after meta:
```json
{
"business_use_case": "authenticated_extract_with_basic_auth_headers",
"auth": "headers.Authorization: Basic … — Geonode does not log in",
"page_count": 2,
"pages": [
{
"url": "https://the-internet.herokuapp.com/basic_auth",
"authenticated_ok": true,
"excerpt": "Congratulations! You must have the proper credentials."
}
]
}
```
| Metric | Demo target |
| ------- | -----------------------------: |
| Auth | HTTP Basic (`admin` / `admin`) |
| APIs | Extract + Batch |
| Outcome | `auth_corpus.json` |
## The plan
```text
1. Probe → Extract WITHOUT Authorization
2. Extract → same URL WITH Authorization: Basic
3. Batch → more protected URLs, same header
4. Parse → title + excerpt (local)
5. Merge → auth_corpus.json
```
no Authorization"] --> B["2. Extract Basic Auth"]
B --> C["3. Batch same header"]
end
subgraph output["Output"]
direction LR
D["4. Parse"] --> E["5. Merge"]
E --> F["auth_corpus.json"]
end
auth_flow --> D
classDef api fill:#e8f5f3,stroke:#054D4D,color:#033333
classDef local fill:#f5f5f5,stroke:#666,color:#222
classDef result fill:#fff8e6,stroke:#b8860b,color:#333
class A,B,C api
class D,E local
class F result
`}
/>
| API | Why we use it here |
| ---------------------------------------------------------------------------------- | --------------------------------------- |
| [Extract](/docs/scraper-api/guides/extraction/01_understanding_extraction) | Clear before/after on one URL |
| [Batch](/docs/scraper-api/guides/batch/00_understanding_batch) | Same Basic header on several URLs |
| [Custom headers](/docs/scraper-api/guides/making-requests/07_using_custom_headers) | `Authorization` is a target-site header |
Compare with other real-world guides (all **public**, no auth):
| Guide | Auth |
| ------------------------------------------------------------------------------------------------- | -------------------------- |
| [B2B agency lead list](/docs/scraper-api/real-world/build_b2b_agency_lead_list) | None |
| [Retail category price catalog](/docs/scraper-api/real-world/build_retail_category_price_catalog) | None |
| [Docs knowledge corpus](/docs/scraper-api/real-world/build_docs_knowledge_corpus) | None |
| **This guide** | **`Authorization: Basic`** |
***
## Before you start
* A Geonode API key
* Python 3.9+, `requests`
```bash
export GEONODE_SCRAPER_API_KEY="YOUR_API_KEY"
```
Heavy SaaS apps that rely on fragile browser cookies often fight proxies. Basic Auth on a simple page is the clearest way to teach authenticated Extract. Apply the same `headers` pattern to **your** staging or internal Basic Auth URLs.
***
## Step 1: Probe without Authorization
### What we need
Proof the page is protected.
### What we will do
`POST /v1/extract` with **no** `headers.Authorization`.
### What you should expect
No “Congratulations” success body (API may surface an error or empty/locked content).
```python
import os
import requests
api_key = os.environ["GEONODE_SCRAPER_API_KEY"]
url = "https://the-internet.herokuapp.com/basic_auth"
response = requests.post(
"https://scraper.geonode.io/v1/extract",
headers={
"X-Api-Key": api_key,
"Content-Type": "application/json",
},
json={
"url": url,
"formats": ["markdown"],
"render_js": False,
"processing_mode": "sync",
},
)
# Should NOT contain the protected success copy
```
### Takeaway
**If anonymous Extract already shows the secret page, you are not testing auth.**
***
## Step 2: Extract with Basic Auth
### What we need
The same URL, authorized.
### What we will do
Encode `username:password` as Base64 and send:
```text
Authorization: Basic YWRtaW46YWRtaW4=
```
(`admin:admin` → `YWRtaW46YWRtaW4=`)
Do **not** put the Geonode API key inside `headers`.
```python
import base64
import os
import requests
api_key = os.environ["GEONODE_SCRAPER_API_KEY"]
url = "https://the-internet.herokuapp.com/basic_auth"
token = base64.b64encode(b"admin:admin").decode("ascii")
response = requests.post(
"https://scraper.geonode.io/v1/extract",
headers={
"X-Api-Key": api_key,
"Content-Type": "application/json",
},
json={
"url": url,
"formats": ["markdown"],
"render_js": False,
"processing_mode": "sync",
"headers": {
"Authorization": f"Basic {token}",
},
},
)
# Expect Markdown containing "Congratulations"
```
### Takeaway
**`headers.Authorization` is target-site auth.** Same field works for `Bearer` tokens when a site expects that instead of Basic.
***
## Step 3: Batch with the same header
### What we need
Several protected URLs without rewriting Extract.
### What we will do
`POST /v1/batch` with `urls` and the same `headers.Authorization`. Demo list:
* `https://the-internet.herokuapp.com/basic_auth`
* `https://httpbin.org/basic-auth/admin/admin` (same `admin` / `admin`)
### What you should expect
One job, one Basic header, multiple Markdown results.
### Takeaway
**Batch reuses headers; it does not invent a login.**
***
## Step 4–5: Parse and merge
Locally turn Markdown into `auth_corpus.json`, including optional before/after meta from probe vs authenticated Extract.
***
## Why this API mix worked
| Tool | Role |
| ------------------ | -------------------------------------------------------------------------------------------------------------- |
| Extract | Prove Basic Auth before/after |
| Batch | Same header, many URLs |
| Cookie / SSO flows | Out of scope here — see FAQ on [custom headers vs full session UI](/docs/scraper-api/additional-resources/faq) |
Use only URLs and credentials you are allowed to access. Practice-site credentials are public by design; production secrets stay in `.env`.
***
## Cost and request usage
| Stage | Rough volume |
| -------------- | -----------: |
| Probe Extract | 1 |
| Authed Extract | 1 |
| Batch | N URLs |
| Parse / merge | 0 (local) |
`render_js` is off for these simple pages.
***
## Recap
1. **Probe** without `Authorization`.
2. **Extract** with `Authorization: Basic …`.
3. **Batch** more protected URLs with the same header.
4. **Parse / merge** into `auth_corpus.json`.
That is the authenticated Extract story: **Geonode key for the API, Basic (or Bearer) for the site.**
# Build a B2B Agency Lead List from a Directory (/docs/scraper-api/real-world/build_b2b_agency_lead_list)
This guide walks you through a real B2B workflow with the Geonode Scraper API: **as if we are building it together**.
You want a list of companies you can use for sales, partnerships, or research. You do **not** already have their websites. You only know the market (Berlin tech / creative agencies).
By the end, you will understand **which API to reach for at each stage**, what a good result looks like, and why we skip Crawl (and avoid mapping every company website).
## Use case
Imagine you are on a partnerships or outbound team. You need a spreadsheet-ready lead list with:
* Company name and website
* Location, size, founding year, hourly rate
* Services / industries
* Public email and phone from the company site
Starting point: **market intent only**: for example, “Berlin software / agency companies.”
Directory used in this walkthrough:
```text
https://techbehemoths.com/companies/software-development/berlin
```
The same pattern works on other public directories.
## What we are going to build
A master JSON lead file. Each company looks roughly like this (shortened):
```json
{
"slug": "andberlin",
"name": "&Berlin Creative Agency",
"website": "https://www.andberlin.co",
"profile_url": "https://techbehemoths.com/company/andberlin",
"hourly_rate": "$70-150/h",
"founded": 2022,
"employees": 10,
"verified": true,
"locations": ["Berlin"],
"services": ["Branding", "Web Design", "Web Development"],
"industries": [{ "name": "Business services", "percent": 10 }],
"emails": ["example@example.com"],
"phones": [],
"has_contact": true
}
```
In a full run against the Berlin software-development directory slice used for this example:
| Metric | Approx. result |
| -------------------------------- | -------------: |
| Companies from directory pages | \~73 |
| Homepages successfully extracted | \~69 |
| With at least one email | \~60 |
| With at least one phone | \~46 |
Exact counts depend on pagination depth and site availability.
## The plan (what you should expect)
We will move through eight stages. At each stage we pick one Geonode product on purpose:
```text
1. Search → find a directory listing URL
2. Map → try link discovery (expect failure on JS listings: teaching moment)
3. Extract → render listing pages + pagination → profile URLs
4. Batch → extract all profile pages → HTML
5. Parse → firmographics + website (local)
6. Batch → extract all company homepages → HTML
7. Parse → emails / phones (local)
8. Merge → master lead file
```
/v1/search"] --> B["Listing URL"]
B --> C["2. Map /v1/map"]
C -->|"JS grids → 0 links"| D["3. Extract /v1/extract + JS"]
end
subgraph enrich["Enrich"]
direction LR
subgraph profiles["Profiles"]
direction TB
E["4. Batch /v1/batch"] --> F["5. Parse local"]
end
subgraph contacts["Contacts"]
direction TB
G["6. Batch sites /v1/batch"] --> H["7. Parse emails / phones"]
end
end
subgraph output["Output"]
direction LR
I["8. Merge"] --> J["companies_master.json"]
end
discover --> enrich
F --> G
F --> I
H --> I
classDef api fill:#e8f5f3,stroke:#054D4D,color:#033333
classDef local fill:#f5f5f5,stroke:#666,color:#222
classDef result fill:#fff8e6,stroke:#b8860b,color:#333
class A,C,D,E,G api
class F,H,I local
class B,J result
`}
/>
| API | Why we use it here |
| -------------------------------------------------------------------------- | ------------------------------------------------- |
| [Search](/docs/scraper-api/guides/search/01_search_overview) | You do not know the best listing URL yet |
| [Map](/docs/scraper-api/guides/map/00_understanding_map) | Fast URL inventory check: often fails on JS grids |
| [Extract](/docs/scraper-api/guides/extraction/01_understanding_extraction) | Render JS listing pages and collect profile links |
| [Batch](/docs/scraper-api/guides/batch/00_understanding_batch) | Many known URLs (profiles, then homepages) |
| Local parsing | Turn HTML into structured fields, then merge |
**Crawl is not used.** After Extract you already know the profile URLs. After profile parse you already know the websites. You do not need a deep multi-page walk of one domain.
***
## Before you start
You need:
* A Geonode API key
* Python 3.9 or later
* The `requests` package
```bash
export GEONODE_SCRAPER_API_KEY="YOUR_API_KEY"
pip install requests
```
On Windows PowerShell:
```powershell
$env:GEONODE_SCRAPER_API_KEY="YOUR_API_KEY"
pip install requests
```
Never put your Geonode API key directly in your source code or commit it to your repository. Store it in an environment variable or `.env` file instead.
***
## Step 1: Find a directory with Search
### What we need
A **listing URL**: a page that shows many company cards. We do not hardcode one on day one; we discover candidates with Search.
### What we will do
Send a natural-language query that matches the market, for example `100 berlin startups`.
### What you should expect
Search returns several sources (`url`, `title`, sometimes a snippet). Your job is to **pick one strong directory**. For this demo we keep a **mini list** from a single directory so the walkthrough stays clear. You can add more directories later with the same pipeline.
```python
import os
import requests
api_key = os.environ["GEONODE_SCRAPER_API_KEY"]
response = requests.post(
"https://scraper.geonode.io/v1/search",
headers={
"X-Api-Key": api_key,
"Content-Type": "application/json",
},
json={
"query": "100 berlin startups",
"page": 1,
"locale": "en",
},
)
response.raise_for_status()
results = response.json()
print(results)
```
From the results, choose one directory listing. In this walkthrough we use TechBehemoths Berlin software-development companies:
```text
https://techbehemoths.com/companies/software-development/berlin
```
Open that URL in a browser so you know what “success” looks like: company cards, profile links, pagination.
### Takeaway
**Search answers “where do I start?”**: not “give me every company yet.” After this step you should have **one listing URL** saved and ready for the next APIs.
***
## Step 2: Try Map on the listing (expect JS limits)
### What we need
A list of company **profile URLs** from the directory (for example `/company/andberlin`).
### What we will do
Run Map on the listing. Map reads **sitemaps and static HTML links**. It does **not** execute JavaScript.
We try Map **on purpose**. On modern directories the company grid is often rendered client-side. Seeing zero profile links teaches you when to switch tools.
### What you should expect
On TechBehemoths-style JS listings, Map often returns **zero** `/company/...` links even though the browser shows dozens of companies. That is normal: not a broken API key.
```python
listing_url = "https://techbehemoths.com/companies/software-development/berlin"
response = requests.post(
"https://scraper.geonode.io/v1/map",
headers={
"X-Api-Key": api_key,
"Content-Type": "application/json",
},
json={
"url": listing_url,
"search": "company",
},
)
response.raise_for_status()
print(response.json())
```
Use Map when you need a fast URL inventory for a site with a sitemap or static navigation: for example, finding `/kontakt`, `/impressum`, or `/about` on a single company domain later. Do not rely on Map alone for JS-heavy listing grids.
### Takeaway
**Map failed here as a discovery path: and that is the lesson.** For JS directory grids, move to **Extract with `render_js`**.
***
## Step 3: Extract listing pages with JavaScript rendering
### What we need
The profile URLs that Map could not see.
### What we will do
Extract the listing with `render_js: true`, wait for the page to settle, then parse `/company/{slug}` links from the Markdown. Repeat for `?page=2`, `?page=3`, … until you have enough companies for your demo.
### What you should expect
Each extracted page should contain many profile links in the Markdown. After a few pages you have a clean list of profile URLs ready for Batch.
```python
response = requests.post(
"https://scraper.geonode.io/v1/extract",
headers={
"X-Api-Key": api_key,
"Content-Type": "application/json",
},
json={
"url": listing_url,
"formats": ["markdown"],
"render_js": True,
"processing_mode": "sync",
"proxy": {"country": "DE", "type": "residential"},
"wait_config": {
"wait_until": "networkidle",
"wait_timeout": 5000,
},
},
)
response.raise_for_status()
markdown = response.json()["data"]["markdown"]
# Parse /company/{slug} links from markdown, then repeat for page=2, page=3, ...
```
Example profile URLs:
```text
https://techbehemoths.com/company/andberlin
https://techbehemoths.com/company/why-studio
...
```
### Takeaway
**Extract is how you discover URLs on JS listings.** Pagination is usually the same Extract call with `?page=N`: you do not need Crawl for this directory pattern.
***
## Step 4: Batch extract all profile pages
### What we need
HTML for every company profile so we can read firmographics and the **website** field.
### What we will do
Submit **one Batch job** with all known profile URLs instead of calling Extract in a loop.
### What you should expect
You get a `job_id`. Poll `GET /v1/batch/{job_id}` until the job completes. Then save each result’s HTML. Concurrency follows your plan’s thread limits automatically: you do not manage worker threads in your app.
```python
profile_urls = [
"https://techbehemoths.com/company/andberlin",
"https://techbehemoths.com/company/why-studio",
# ... more profile URLs
]
response = requests.post(
"https://scraper.geonode.io/v1/batch",
headers={
"X-Api-Key": api_key,
"Content-Type": "application/json",
},
json={
"urls": profile_urls,
"ignore_invalid_urls": True,
"formats": ["html"],
"render_js": True,
"wait_config": {"wait_until": "domcontentloaded"},
},
)
response.raise_for_status()
job = response.json()
job_id = job["job_id"]
# Poll GET /v1/batch/{job_id} until status is completed, then save each result HTML
```
### Takeaway
**Many known URLs → Batch.** Extract was for discovery on listing pages. Batch is for volume on URLs you already have.
***
## Step 5: Parse profiles locally
### What we need
Structured company records: name, website, size, services, and so on.
### What we will do
Parse the saved profile HTML locally (no Geonode call). Directory profiles usually expose firmographics and a website: but rarely a public email.
### What you should expect
JSON objects with fields such as:
* `name`, `website`, `hourly_rate`, `founded`, `employees`
* `locations`, `services`, `industries`
* `profile_url`
You should see a **website** and still see **no email** on most rows. That is why the next steps exist.
### Takeaway
**Directory pages give firmographics. Contact data usually lives on the company site.** Keep the website field: it is the bridge to enrichment.
***
## Step 6: Batch extract company homepages
### What we need
Emails and phones. Homepages (and footers) are the fastest place to look first.
### What we will do
Batch-extract each company’s **homepage** from the website field. We intentionally **do not Map every company domain**.
### Why we skip Map-on-every-website here
* Map returns **URLs only**. You still need Extract/Batch afterward.
* Agency sites often expose huge sitemaps (hundreds of URLs). Mapping each domain is slow and noisy.
* Public emails/phones are often already on the homepage (`mailto:`, `tel:`, footer text).
### What you should expect
A second Batch job over homepage URLs. Some sites fail or block; that is fine: those companies stay in the master file with empty contacts.
```python
homepage_urls = [
"https://www.andberlin.co",
"https://www.why.de",
# ... websites from the profile parse step
]
response = requests.post(
"https://scraper.geonode.io/v1/batch",
headers={
"X-Api-Key": api_key,
"Content-Type": "application/json",
},
json={
"urls": homepage_urls,
"ignore_invalid_urls": True,
"formats": ["html"],
"render_js": True,
"wait_config": {"wait_until": "domcontentloaded"},
},
)
response.raise_for_status()
# Poll the batch job, then save homepage HTML per company
```
### Takeaway
**Homepage-first enrichment is cheaper and faster than Map fan-out.** If a company still has no contact after this, *then* consider Map on that one domain for `/contact` / `/impressum`.
***
## Step 7: Parse contacts and merge the master file
### What we need
The final lead list: firmographics + emails/phones in one record per company.
### What we will do
1. From each homepage HTML, extract `mailto:` / email-like strings and `tel:` / phone-like text.
2. Join profile fields + contacts on a stable key (slug or normalized website).
3. Light-clean placeholders (`example.com`, theme demos), dedupe phones, keep companies even when homepage Batch failed.
### What you should expect
A contacts file with `emails` / `phones` arrays, then a master file where many rows have `has_contact: true`.
### Takeaway
**Scraper API gets you the pages. Local parse + merge makes the CRM-ready lead list.** Keep parsing simple and filter junk before outreach.
***
## Why this API mix worked
### Search was enough to start
You only needed market intent. Search returns candidate sources; you pick one listing and continue.
### Map was useful as a check, not as the main discovery path
| Situation | Better tool |
| ------------------------------------------ | ---------------------------- |
| JS directory grid / infinite scroll cards | **Extract** with `render_js` |
| Static site or sitemap-heavy single domain | **Map** |
| Deep walk of one large website | **Crawl** |
| Many known URLs (profiles or homepages) | **Batch** |
### Extract + Batch covered the whole lead pipeline
1. Extract discovers profile URLs from rendered listing pages (+ pagination).
2. Batch pulls all profile HTML.
3. Batch pulls all homepage HTML.
4. Local code turns HTML into CRM-ready fields.
That is usually enough for **directory → firmographics → homepage contacts**.
### When to add Map or Crawl later
* **Map a company site** after the homepage has no email: discover `/contact`, `/kontakt`, `/impressum`, then Batch only those few URLs.
* **Crawl** a single large corporate domain when you need broad content discovery beyond a short contact URL list.
Do not start with Crawl across dozens of unrelated company domains for this use case.
***
## Cost and request usage
On [request-based pricing](/docs/scraper-api/additional-resources/request_based_pricing), successful **page extractions** consume requests. Map and job-status polling do not count as page extractions the same way content extraction does.
Approximate request shape for a run like this example:
| Stage | Rough volume | Notes |
| -------------------------- | -----------: | --------------------------------------------- |
| Listing Extract | \~3 pages | page 1–3 with JS rendering |
| Profile Batch | \~73 URLs | 1 request per successfully extracted profile |
| Homepage Batch | \~69 URLs | 1 request per successfully extracted homepage |
| **Total page extractions** | **\~145** | Order-of-magnitude for this demo depth |
### Cost contrast: Map every company website first
If you Map \~70 company domains, you may discover hundreds of URLs per site. Batching those multiplies spend and runtime. For homepage-first contact enrichment, **Batch homepages first**, then Map selectively only for companies still missing contact data.
Pricing and plan details can change: check:
* [Request-Based Pricing](/docs/scraper-api/additional-resources/request_based_pricing)
* [Unlimited Scraper API Pricing](/docs/scraper-api/additional-resources/unlimited_scraper_api_pricing)
* [Choosing a Scraper API Plan](/docs/scraper-api/additional-resources/choosing_scraper_api_plan)
Unlimited plans bill by **concurrency (threads)**, not a monthly request balance: useful when you run large Batches often.
***
## Limitations and good practice
* Use publicly available pages and respect site terms / robots rules for your jurisdiction and use case.
* Directory data can be incomplete or outdated; prefer the company website as the contact source of truth.
* Homepage parsing will miss contacts that exist only behind forms, images, or Impressum-only pages: that is when selective Map + Batch helps.
* Filter obvious placeholder emails before exporting to a CRM or outreach tool.
* This guide produces a **lead list**, not a compliance or consent platform. Apply your own outreach and privacy policies.
***
## Recap
1. **Search** finds the directory.
2. **Map** may fail on JS listings: switch to **Extract**.
3. **Extract** + pagination collects profile URLs.
4. **Batch** extracts all profiles, then all homepages.
5. Local parse + merge produces the master B2B lead file.
# Build a Docs Knowledge Corpus with Crawl (/docs/scraper-api/real-world/build_docs_knowledge_corpus)
This guide walks you through a real documentation workflow with the Geonode Scraper API: **as if we are building it together**.
You want a **knowledge corpus**: many docs pages as clean Markdown for RAG, internal search, or a support bot. You **already know the docs site**. You do **not** have a ready-made list of every guide URL.
By the end, you will understand **when to use Crawl** instead of Map, Extract, or Batch, and how a single async job discovers linked pages and returns their content.
Companion example code lives in `geonode-scraper-examples/geonode-docs-corpus/`.
## Use case
Imagine you are on a docs, support, or AI team. You need:
* Many documentation pages as Markdown
* Stable URLs and titles for indexing
* A bounded crawl (not the whole internet)
Starting point: **a known docs root**, for example Geonode’s public docs (the same seed used in the Crawl API guides):
```text
https://docs.geonode.com/docs/scraper-api
```
Seeding the Scraper API section (instead of only the docs homepage) walks nested guides under Extraction, Batch, Crawl, Map, and Search.
## What we are going to build
A master JSON corpus. Each page looks roughly like this (shortened):
```json
{
"url": "https://docs.geonode.com/docs/scraper-api/guides/crawl/01_first-crawl",
"title": "Your First Crawl",
"section": "scraper-api",
"path": "/docs/scraper-api/guides/crawl/01_first-crawl",
"depth": 2,
"excerpt": "In this guide, you'll create your first crawl job...",
"markdown_len": 4200
}
```
| Metric | Demo target |
| ----------------------- | -----------------------------------: |
| Seed | Scraper API docs section |
| Crawl `limit` / `depth` | 30 / 3 |
| Outcome | `docs_corpus.json` with page records |
{/*  */}
## The plan (what you should expect)
Fewer stages than a directory lead list or retail catalog, because **Crawl discovers and extracts in one job**:
```text
1. Crawl → seed docs → linked pages + markdown
2. Parse → title, section, excerpt (local)
3. Merge → docs_corpus.json
```
POST /v1/crawl"] --> B["Poll GET /v1/crawl/job_id"]
B --> C["Page markdown files"]
end
subgraph output["Output"]
direction LR
D["2. Parse local"] --> E["3. Merge"]
E --> F["docs_corpus.json"]
end
crawl_job --> D
classDef api fill:#e8f5f3,stroke:#054D4D,color:#033333
classDef local fill:#f5f5f5,stroke:#666,color:#222
classDef result fill:#fff8e6,stroke:#b8860b,color:#333
class A,B api
class D,E local
class C,F result
`}
/>
| API | Why we use it here |
| ------------------------------------------------------ | ----------------------------------------------------------- |
| [Crawl](/docs/scraper-api/guides/crawl/01_first-crawl) | One seed, unknown linked tree, content + discovery together |
| Local parsing | Turn Markdown into indexable fields |
**Search is not used.** You already know the docs URL.
**Map is not required.** Map would only return URLs; you would still Batch Extract. Crawl returns content.
**Batch is not required.** You do not have a URL list until Crawl finishes, and by then pages are already extracted.
Compare with other real-world guides:
| Guide | Why Crawl was skipped |
| ------------------------------------------------------------------------------------------------- | ---------------------------------------------------- |
| [B2B agency lead list](/docs/scraper-api/real-world/build_b2b_agency_lead_list) | Directory + pagination gave profile URLs; then Batch |
| [Retail category price catalog](/docs/scraper-api/real-world/build_retail_category_price_catalog) | Sitemap + PLP Extract gave product URLs; then Batch |
***
## Before you start
You need:
* A Geonode API key
* Python 3.9 or later
* The `requests` package
```bash
export GEONODE_SCRAPER_API_KEY="YOUR_API_KEY"
pip install requests
```
On Windows PowerShell:
```powershell
$env:GEONODE_SCRAPER_API_KEY="YOUR_API_KEY"
pip install requests
```
Never put your Geonode API key directly in your source code or commit it to your repository. Store it in an environment variable or `.env` file instead.
***
## Step 1: Crawl the docs section
### What we need
Many docs pages with Markdown, without hand-picking every guide URL.
### What we will do
Create one Crawl job from the Scraper API docs seed. Cap the walk with `limit` and `depth`. Keep `same_domain_only: true`. Docs are mostly static, so start with `render_js: false`.
### What you should expect
A `202` response with `job_id`. Poll `GET /v1/crawl/{job_id}` until `status` is `completed`. Save each completed page’s Markdown.
```python
import os
import time
import requests
api_key = os.environ["GEONODE_SCRAPER_API_KEY"]
base = "https://scraper.geonode.io"
response = requests.post(
f"{base}/v1/crawl",
headers={
"X-Api-Key": api_key,
"Content-Type": "application/json",
},
json={
"url": "https://docs.geonode.com/docs/scraper-api",
"formats": ["markdown"],
"limit": 30,
"depth": 3,
"same_domain_only": True,
"include_subdomains": False,
"render_js": False,
},
)
response.raise_for_status()
job_id = response.json()["job_id"]
while True:
job = requests.get(
f"{base}/v1/crawl/{job_id}",
headers={"X-Api-Key": api_key},
).json()
print(job["status"], job.get("completed_pages"), "/", job.get("total_pages"))
if job["status"] in {"completed", "failed", "cancelled"}:
break
time.sleep(5)
# job["results"] → save data.markdown per page
```
Crawling only `https://docs.geonode.com/` with shallow depth may stop at a handful of top-level hubs (Getting Started, Proxies, Scraper API). Seeding **`/docs/scraper-api`** walks nested Extraction, Batch, Crawl, Map, and Search guides, better for a RAG demo.
### Takeaway
**Crawl answers “give me linked pages with content from this seed.”** Bound the job with `limit` and `depth` so demos stay cheap and predictable.
***
## Step 2: Parse pages locally
### What we need
Structured records: URL, title, section, short excerpt.
### What we will do
Parse saved Markdown locally (no Geonode call). Prefer the first `#` heading or front-matter `title:` as the page title. Derive `section` from the URL path (`/docs/scraper-api/...` → `scraper-api`).
### What you should expect
A `pages.json` with one object per crawled URL, plus section counts.
### Takeaway
**Scraper API gets you the pages. Local parse makes the corpus indexable.** Keep excerpts short for demos; store full Markdown files separately if you need RAG chunks later.
***
## Step 3: Merge the master corpus
### What we need
The demo deliverable: one slim file a search or AI pipeline can load.
### What we will do
Drop failed pages and internal file names. Keep `url`, `title`, `section`, `path`, `depth`, `excerpt`, `markdown_len`. Add `page_count` and section histogram.
### What you should expect
`docs_corpus.json`, the knowledge-corpus equivalent of a lead list or product catalog master file.
### Takeaway
Same pattern as other real-world guides: **Geonode fetches; your merge step is the business artifact.**
***
## Why this API mix worked
### Crawl was required
You had **one seed** and needed **many unknown linked pages with content**. That is Crawl’s job.
### Map would have been incomplete
| Tool | What you get |
| ------- | ------------------------------------------------ |
| Map | URLs only, still need Extract/Batch |
| Extract | One page per call, you write the link walker |
| Batch | Needs a URL list you already have |
| Crawl | Discover + extract, async, capped by limit/depth |
### When not to use Crawl
* You already have every URL → **Batch**
* You only need a link inventory → **Map**
* You need one page → **Extract**
* You do not know which site → **Search** first, then Crawl that site
***
## Cost and request usage
On [request-based pricing](/docs/scraper-api/additional-resources/request_based_pricing), successful page extractions in a crawl consume requests (roughly one per completed page).
| Stage | Rough volume | Notes |
| ------------- | ------------------: | ------------------ |
| Crawl | up to `limit` pages | Demo uses limit 30 |
| Parse / merge | 0 | Local only |
Keep `limit` small while teaching. Raise it when you need a fuller corpus.
***
## Limitations and good practice
* Prefer first-party or permitted docs sites for demos.
* Docs trees change; re-run Crawl when the corpus must stay fresh.
* Shallow depth on a marketing homepage may miss nested guides, seed the section you care about.
* Excerpts are not chunking strategy; production RAG usually splits Markdown further.
***
## Recap
1. **Crawl** the docs section with `limit` / `depth`.
2. **Parse** Markdown into titles, sections, excerpts.
3. **Merge** into `docs_corpus.json` for RAG or search.
That is the Crawl-first story: **one seed, unknown URLs, content included**, the gap left by the B2B lead list and retail catalog examples.
# Build a Retail Category Price Catalog from a Public Sitemap (/docs/scraper-api/real-world/build_retail_category_price_catalog)
This guide walks you through a real e-commerce workflow with the Geonode Scraper API: **as if we are building it together**.
You want a **competitive catalog snapshot**: what a retailer sells in one category, at what price, on promo or not. You **already know the website**. You do **not** need Search to find it.
By the end, you will understand **which API to reach for at each stage**, why Map on a homepage is not the same as the XML sitemap you open in a browser, and why Extract does not paginate for you.
Companion example code lives in `geonode-scraper-examples/bedbathandbeyond-catalog/`.
## Use case
Imagine you are on a category, marketplace, or pricing team. You need a spreadsheet-ready product list with:
* Product title, brand, SKU / item number
* Current price (USD) and sale flag
* Category breadcrumb
* Product URL and retailer product id
Starting point: **a known retailer**, not a search query. Example: Bed Bath & Beyond.
Sitemap and category used in this walkthrough:
```text
https://www.bedbathandbeyond.com/sitemap.xml
https://www.bedbathandbeyond.com/sitemap/ctaxonomy/ctaxonomy.xml
https://www.bedbathandbeyond.com/c/towels/bath-towels?t=18652
```
The same pattern works on other public retailers that publish a sitemap index and `/c/...` category pages (product listing pages, or **PLPs**). Each SKU lives on a product detail page (**PDP**), often ending in `/product.html`.
## What we are going to build
A master JSON catalog file. Each product looks roughly like this (shortened):
```json
{
"product_id": "33411469",
"sku": "37850619",
"url": "https://www.bedbathandbeyond.com/Bedding-Bath/.../33411469/product.html",
"title": "American Soft Linen 100% Cotton Turkish Bath Towels...",
"brand": "American Soft Linen",
"price": 55.49,
"list_price": null,
"on_sale": true,
"promo": "Labor Day Sale",
"currency": "USD",
"category": ["Bedding & Bath", "Bath Linens", "Towels", "Bath Towels"],
"availability": "unknown"
}
```
In a demo run against **two PLP pages** of Bath Towels:
| Metric | Approx. result |
| --------------------------------- | -------------: |
| Category PLPs in taxonomy sitemap | \~1430 |
| Product URLs from 2 listing pages | \~36 |
| PDP HTML files saved | \~34 |
| With a parsed price | \~30 |
| Price range (USD) | \~$33–$120 |
Exact counts depend on pagination depth, Batch failures, and site availability.
## The plan (what you should expect)
We move through six stages. At each stage we pick one Geonode product on purpose:
```text
1. Map → sitemap.xml (index of sub-sitemaps)
2. Map → ctaxonomy.xml (all category PLP URLs)
3. Extract → one PLP + pagination → product URLs
4. Batch → all PDPs → HTML
5. Parse → title, price, SKU, sale (local)
6. Merge → products_master.json
```
/v1/map"] --> B["Sub-sitemap URLs"]
B --> C["2. Map ctaxonomy.xml /v1/map"]
C --> D["Category PLPs /c/..."]
end
subgraph products["Collect products"]
direction LR
E["3. Extract PLP /v1/extract + JS + page=N"] --> F["Product URLs"]
F --> G["4. Batch PDPs /v1/batch"]
end
subgraph output["Output"]
direction LR
H["5. Parse HTML local"] --> I["6. Merge"]
I --> J["products_master.json"]
end
discover --> products
G --> H
classDef api fill:#e8f5f3,stroke:#054D4D,color:#033333
classDef local fill:#f5f5f5,stroke:#666,color:#222
classDef result fill:#fff8e6,stroke:#b8860b,color:#333
class A,C,E,G api
class H,I local
class B,D,F,J result
`}
/>
| API | Why we use it here |
| -------------------------------------------------------------------------- | --------------------------------------------------------------------- |
| [Map](/docs/scraper-api/guides/map/00_understanding_map) | Fast URL inventory from sitemap XML and taxonomy |
| [Extract](/docs/scraper-api/guides/extraction/01_understanding_extraction) | Render a JS category grid and parse product links; one call = one URL |
| [Batch](/docs/scraper-api/guides/batch/00_understanding_batch) | Many known PDP URLs in one job |
| Local parsing | Turn HTML into price/SKU fields, then merge |
**Search is not used.** You already know the retailer.
**Crawl is not required for this demo.** After Extract you already have product URLs. Crawl is the right tool later if you want a BFS walk of one category with a page `limit` instead of looping `?page=`.
***
## Before you start
You need:
* A Geonode API key
* Python 3.9 or later
* The `requests` package
```bash
export GEONODE_SCRAPER_API_KEY="YOUR_API_KEY"
pip install requests
```
On Windows PowerShell:
```powershell
$env:GEONODE_SCRAPER_API_KEY="YOUR_API_KEY"
pip install requests
```
Never put your Geonode API key directly in your source code or commit it to your repository. Store it in an environment variable or `.env` file instead.
***
## Step 1: Map the sitemap index (not the homepage)
### What we need
The retailer’s **sitemap index**: a list of sub-sitemap XML files (taxonomy, refinements, keyword pages, and so on).
### What we will do
Open the sitemap in a browser so you know what “success” looks like:
```text
https://www.bedbathandbeyond.com/sitemap.xml
```
You should see a `` with `` entries such as:
```text
https://www.bedbathandbeyond.com/sitemap/ctaxonomy/ctaxonomy.xml
https://www.bedbathandbeyond.com/sitemap/ctaxonomy/refinements.xml
https://www.bedbathandbeyond.com/sitemap/keyword-search-pages/keyword-search-pages.xml
```
Then call Map on **that same URL** (`sitemap.xml`), not on `https://www.bedbathandbeyond.com`.
### Why homepage Map looks different
Map means: **discover links from this seed**. It is not “pretty-print the XML file Chrome showed you.”
| Seed you Map | Typical result |
| ------------- | ------------------------------------------------------ |
| Homepage `/` | Mix of site links, often **PDPs** (`.../product.html`) |
| `sitemap.xml` | Inventory of URLs Map can see from that document |
If Map on `sitemap.xml` does not return the `.xml` children you see in the browser, **Extract the sitemap with `render_js: false`** and parse `` tags locally. That is the same lesson as JS directories: Map first, Extract when the document content is what you need.
```python
import os
import requests
api_key = os.environ["GEONODE_SCRAPER_API_KEY"]
response = requests.post(
"https://scraper.geonode.io/v1/map",
headers={
"X-Api-Key": api_key,
"Content-Type": "application/json",
},
json={
"url": "https://www.bedbathandbeyond.com/sitemap.xml",
"include_subdomains": True,
},
)
response.raise_for_status()
print(response.json())
```
### Takeaway
**Map answers “what URLs can we inventory from this seed?”** For a sitemap index, seed **`sitemap.xml`**. Save `ctaxonomy.xml` as the next seed.
***
## Step 2: Map the category taxonomy (all PLPs)
### What we need
Every **category listing URL** (`/c/...`) so you can pick one demo category (Bath Towels).
### What we will do
Map (or Extract-parse `` if Map times out) on:
```text
https://www.bedbathandbeyond.com/sitemap/ctaxonomy/ctaxonomy.xml
```
In the browser this file is a `` of category pages, for example:
```text
https://www.bedbathandbeyond.com/c/towels/bath-towels?t=18652
```
Taxonomy sitemaps can be huge. Map may return **HTTP 408** (Map request timed out during URL discovery). That is a documented Map error, not a bad API key. Retry, or Extract the XML and parse every `` that contains `/c/`.
```python
response = requests.post(
"https://scraper.geonode.io/v1/map",
headers={
"X-Api-Key": api_key,
"Content-Type": "application/json",
},
json={
"url": "https://www.bedbathandbeyond.com/sitemap/ctaxonomy/ctaxonomy.xml",
"include_subdomains": True,
"ignore_query_parameters": False,
},
)
response.raise_for_status()
links = response.json().get("links") or []
# Keep URLs whose path contains /c/
```
In this example run, taxonomy discovery produced **about 1430** category PLPs. We pick one:
```text
https://www.bedbathandbeyond.com/c/towels/bath-towels?t=18652
```
### Takeaway
**Second Map is the category tree**, not product SKUs. PDPs usually live on listing pages, not in this taxonomy file.
***
## Step 3: Extract the category PLP (pagination is your loop)
### What we need
**Product URLs** (`.../product.html`) from the Bath Towels grid.
### What we will do
1. Extract **page 1** with JavaScript rendering and `wait_config`.
2. Parse the highest `page=` in the Markdown (the UI can show the last page on page 1).
3. Call Extract again for `?page=2`, `?page=3`, … yourself.
### Extract does not auto-paginate
One Extract request = **one URL**. There is no “follow all pages” flag. That is the same pattern as directory listings in the [B2B lead list guide](/docs/scraper-api/real-world/build_b2b_agency_lead_list).
Bed Bath & Beyond uses a query parameter:
```text
https://www.bedbathandbeyond.com/c/towels/bath-towels?t=18652&page=33
```
### wait\_config (when the grid is JS)
From [Waiting for Dynamic Content](/docs/scraper-api/guides/making-requests/04_waiting_for_dynamic_content), wait order is:
```text
wait_until → wait_for → wait_timeout → extract
```
For a product grid:
```python
response = requests.post(
"https://scraper.geonode.io/v1/extract",
headers={
"X-Api-Key": api_key,
"Content-Type": "application/json",
},
json={
"url": "https://www.bedbathandbeyond.com/c/towels/bath-towels?t=18652",
"formats": ["markdown"],
"render_js": True,
"processing_mode": "sync",
"proxy": {"country": "US", "type": "residential"},
"wait_config": {
"wait_until": "networkidle",
"wait_for": 'a[href*="product.html"]',
"wait_timeout": 5000,
},
},
)
response.raise_for_status()
markdown = response.json()["data"]["markdown"]
# Parse product.html links; parse max page=N; repeat Extract for page=2...
```
`429` means **request throttled or work concurrency limit reached** ([Error Handling](/docs/scraper-api/troubleshooting/error-handling)). Slow down, honor `Retry-After` if present, and retry with backoff. Sync Extract with `render_js` uses workers.
This demo extracted **2 PLP pages** and collected **\~36** product URLs.
{/*  */}
### Takeaway
**Extract discovers PDPs on JS category pages.** You implement pagination. Crawl is optional if you prefer one async job with a `limit` instead of a page loop.
***
## Step 4: Batch extract all product pages
### What we need
HTML for every PDP so we can read price, SKU, and title.
### What we will do
Submit **one Batch job** with all known product URLs instead of calling Extract in a loop.
### What you should expect
You get a `job_id`. Poll `GET /v1/batch/{job_id}` until the job completes. Save each result’s HTML (use the numeric product id in the filename so files are not all named `product.html`).
Some URLs may fail; keep going. In this run: **36** submitted, **34** HTML files, **2** failed.
```python
product_urls = [
"https://www.bedbathandbeyond.com/Bedding-Bath/.../33411469/product.html",
# ... URLs from step 3
]
response = requests.post(
"https://scraper.geonode.io/v1/batch",
headers={
"X-Api-Key": api_key,
"Content-Type": "application/json",
},
json={
"urls": product_urls,
"ignore_invalid_urls": True,
"formats": ["html"],
"render_js": True,
"proxy": {"country": "US", "type": "residential"},
"wait_config": {
"wait_until": "networkidle",
"wait_timeout": 3000,
},
},
)
response.raise_for_status()
job_id = response.json()["job_id"]
# Poll GET /v1/batch/{job_id} until completed, then save each result HTML
```
### Takeaway
**Many known URLs → Batch.** Extract was for PLP discovery. Batch is for volume on PDPs you already have.
***
## Step 5: Parse product HTML locally
### What we need
Structured product rows: title, price, SKU, sale flag, category.
### What we will do
Parse saved HTML locally (no Geonode call). Retailer PDPs often embed a compact analytics object (in this example, `ensighten.items` with `price`, `sku`, `productId`, `productName`) plus breadcrumbs.
### What you should expect
JSON objects with fields such as:
* `product_id`, `sku`, `url`, `title`, `brand`
* `price`, `list_price`, `on_sale`, `currency`
* `category` (breadcrumb)
You may **not** get a reliable in-stock flag from HTML alone (add-to-cart vs “out of stock” copy is noisy). Treat `availability` as best-effort.
{/*  */}
### Takeaway
**Scraper API gets you the pages. Local parse makes the pricing spreadsheet.** Prefer structured blobs in the HTML over scraping the entire 3000-line document.
***
## Step 6: Merge the master catalog
### What we need
The final deliverable: one slim file a category or pricing person can use.
### What we will do
Drop parse internals (`source_file`). Keep business fields. Add stats (`with_price`, `on_sale_count`, min/max/avg USD). Optionally list Batch-failed URLs so you can retry later.
### What you should expect
`products_master.json` with a `products` array and summary metrics.
### Takeaway
**Same idea as a B2B master lead file:** Geonode fetches pages; your merge step is the business artifact.
***
## Why this API mix worked
### You already knew the site (no Search)
Search is for **market intent** (“Berlin agencies”). Here the seed is a **known domain + public sitemap**.
### Map was the right first tool (unlike the JS directory)
| Situation | Better tool |
| ------------------------------------- | ----------------------------------------------------------- |
| Sitemap index / taxonomy XML | **Map** (Extract `` if Map 408 or misses XML children) |
| JS category grid / pagination | **Extract** with `render_js` + `wait_config` |
| Many known PDP URLs | **Batch** |
| Deep walk of one site with a page cap | **Crawl** (optional; not used in this demo) |
### Extract + Batch covered listing → SKUs → HTML
1. Map builds the category list from taxonomy.
2. Extract discovers product URLs from rendered PLPs (+ your page loop).
3. Batch pulls all PDP HTML.
4. Local code turns HTML into a catalog JSON.
That is usually enough for **sitemap → category → products → prices**.
### When to add Crawl later
Use [Crawl](/docs/scraper-api/guides/crawl/01_first-crawl) from one PLP if you want async BFS with `depth` and `limit` instead of writing a `?page=` loop. Keep `limit` small for demos.
Do not Crawl the entire retailer from the homepage for this use case.
***
## Cost and request usage
On [request-based pricing](/docs/scraper-api/additional-resources/request_based_pricing), successful **page extractions** consume requests. Map and job-status polling do not count as page extractions the same way content extraction does.
Approximate request shape for a run like this example (2 PLP pages):
| Stage | Rough volume | Notes |
| -------------------------- | -----------: | ---------------------------------------------- |
| Map sitemap + taxonomy | 2 Map calls | Plus Extract on XML only if Map misses `` |
| PLP Extract | 2 pages | page 1–2 with JS rendering (`--max-pages 2`) |
| PDP Batch | \~36 URLs | 1 request per successfully extracted product |
| **Total page extractions** | **\~38+** | Order-of-magnitude for this demo depth |
If you Extract **all** Bath Towels pages (for example 30+), listing Extract grows linearly. Batch then grows with unique PDPs.
### Cost contrast: Crawl the whole site
A homepage Crawl with a high `limit` can fetch hundreds of mixed URLs (guides, search pages, PDPs). For one category, **Extract that PLP + Batch those PDPs** stays bounded.
Pricing and plan details can change: check:
* [Request-Based Pricing](/docs/scraper-api/additional-resources/request_based_pricing)
* [Unlimited Scraper API Pricing](/docs/scraper-api/additional-resources/unlimited_scraper_api_pricing)
* [Choosing a Scraper API Plan](/docs/scraper-api/additional-resources/choosing_scraper_api_plan)
Unlimited plans bill by **concurrency (threads)**, not a monthly request balance: useful when you run large Batches often.
***
## Limitations and good practice
* Use publicly available pages and respect site terms / robots rules for your jurisdiction and use case.
* Prices and promo flags change; this is a **point-in-time snapshot**, not a live feed unless you re-run Batch.
* `availability` parsed from HTML can be wrong; confirm in the UI if stock is a business-critical field.
* Pagination windows in the UI may not show the true last page on page 1; still parse `page=` links and cap `--max-pages` for demos.
* This guide produces a **catalog file**, not a commercial scraping product. Apply your own policies.
***
## Recap
1. **Map** `sitemap.xml` to find sub-sitemaps (Extract XML if Map misses ``).
2. **Map** `ctaxonomy.xml` to list category PLPs (handle Map **408** on huge files).
3. **Extract** one PLP with JS + `wait_config`; loop `?page=` yourself.
4. **Batch** all PDPs to HTML.
5. Local parse + merge produces the master price catalog.
Companion: B2B directories use **Search → Map-often-fails → Extract listings**. Retailers with a sitemap use **Map twice → Extract PLP → Batch PDPs**.
# Scraper API Request Parameters (/docs/scraper-api/snippets/requests-parameters)
import Link from "next/link";
Output Formats
JavaScript Rendering
Waiting for Dynamic Content
Processing Modes
Proxy and Geo-Targeting
Using Custom Headers
# Code Examples (/docs/scraper-api/troubleshooting/code-examples)
These examples call `POST /v1/extract` in synchronous mode and print the extracted Markdown. Before running them, set your Scraper API base URL and API key.
```bash
export SCRAPER_API_BASE_URL="https://scraper.geonode.io"
export GEONODE_SCRAPER_API_KEY="YOUR_API_KEY"
```
## cURL
Use this example when you want to test the API from a terminal before writing application code.
```bash
curl -X POST "$SCRAPER_API_BASE_URL/v1/extract" \
-H "X-Api-Key: $GEONODE_SCRAPER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "https://example.com",
"formats": ["markdown"],
"render_js": false,
"processing_mode": "sync"
}'
```
## Python
This example uses Python's standard library so you do not have to install any extra packages.
```python
import json
import os
from urllib import request, error
base_url = os.environ["SCRAPER_API_BASE_URL"]
api_key = os.environ["GEONODE_SCRAPER_API_KEY"]
payload = {
"url": "https://example.com",
"formats": ["markdown"],
"render_js": False,
"processing_mode": "sync",
}
req = request.Request(
f"{base_url}/v1/extract",
data=json.dumps(payload).encode("utf-8"),
headers={
"X-Api-Key": api_key,
"Content-Type": "application/json",
},
method="POST",
)
try:
with request.urlopen(req, timeout=60) as response:
result = json.loads(response.read().decode("utf-8"))
print(result["data"]["markdown"])
except error.HTTPError as exc:
body = exc.read().decode("utf-8")
print(f"Request failed with HTTP {exc.code}: {body}")
```
## Node.js
This example uses the built-in `fetch` API available in current Node.js versions.
```javascript
const baseUrl = process.env.SCRAPER_API_BASE_URL;
const apiKey = process.env.GEONODE_SCRAPER_API_KEY;
const response = await fetch(`${baseUrl}/v1/extract`, {
method: "POST",
headers: {
"X-Api-Key": apiKey,
"Content-Type": "application/json",
},
body: JSON.stringify({
url: "https://example.com",
formats: ["markdown"],
render_js: false,
processing_mode: "sync",
}),
});
const result = await response.json();
if (!response.ok) {
throw new Error(`Request failed with HTTP ${response.status}: ${JSON.stringify(result)}`);
}
console.log(result.data.markdown);
```
# Error Handling (/docs/scraper-api/troubleshooting/error-handling)
The Scraper API returns standard HTTP status codes and JSON error bodies. Handle both validation errors and extraction errors in your client, because a request can fail before extraction starts or after the API tries to process the target page.
## HTTP Status Codes
| HTTP status | Meaning | Returned by |
| ----------- | --------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------- |
| `200` | Synchronous extraction, map request, job lookup, statistics request, webhook lookup, or health check succeeded. | Extract (sync), map, get job, get batch/crawl status, list jobs, statistics, webhook get/list, health |
| `201` | Webhook subscription was created. | Create webhook |
| `202` | Async extraction, batch, or crawl job was accepted, or a running job cancellation was accepted. | Extract (async), create batch, create crawl, cancel batch, cancel crawl |
| `204` | Webhook subscription was deleted. | Delete webhook |
| `400` | Invalid request. | Extract, map |
| `401` | API key is missing or invalid. | All authenticated endpoints |
| `402` | Payment required or insufficient request balance. | Extract, map, crawl |
| `404` | Job, webhook, or other requested resource was not found. | Get job, get batch/crawl status, cancel batch/crawl, webhook get/update/delete/deliveries |
| `408` | Map request timed out during URL discovery. | Map |
| `409` | Batch or crawl job cannot be cancelled in its current state. | Cancel batch, cancel crawl |
| `422` | Validation error or extraction failed. | Extract (body validation and extraction errors), create batch, create crawl, create/update webhook |
| `429` | Request throttled or work concurrency limit reached. | Extract, create batch, create crawl, map |
| `500` | Internal server error. | Extract, map, webhook create/get/update/delete/list/deliveries |
| `502` | Billing service returned an upstream error. Retryable depends on the billing error. | Extract, map, create crawl |
| `503` | Service or billing service temporarily unavailable. | Extract, map, create batch, create crawl |
| `504` | Synchronous extraction timed out waiting for the browser worker. Retryable. | Extract (sync) |
## Validation Error
Validation errors can return a `detail` array. This usually means the request body does not match the expected schema, such as an invalid URL or unsupported value.
```json
{
"detail": [
{
"type": "value_error",
"loc": ["body", "url"],
"msg": "Value error, URL must contain a valid hostname.",
"input": "not-a-url"
}
]
}
```
## Extraction Error
Extraction failures return an `error` object and `tokens_charged`.
```json
{
"error": {
"code": "UNPROCESSABLE_CONTENT",
"message": "The target page could not be extracted.",
"retryable": false,
"details": null
},
"tokens_charged": 0
}
```
The response field is currently named `tokens_charged` in the API schema. In Scraper API docs and billing language, treat this value as the number of requests charged.
## Extraction Error Codes
Extraction error codes include:
* `RATE_LIMITED`
* `WORK_CONCURRENCY_LIMITED`
* `WORK_CONCURRENCY_LEASE_EXPIRED`
* `TEMPORARY_BLOCK`
* `NETWORK_ERROR`
* `CAPTCHA_CHALLENGE`
* `PERMANENT_BLOCK`
* `INVALID_URL`
* `UNPROCESSABLE_CONTENT`
* `AUTH_REQUIRED`
* `TIMEOUT`
* `PAYMENT_REQUIRED`
* `INTERNAL_ERROR`
* `PROXY_ERROR`
Use the `retryable` field to decide whether a retry may help. For retryable failures, use exponential backoff and avoid retrying in a tight loop.
The OpenAPI schema includes `429` responses, but no public numeric rate limit is specified in the current contract. If you receive `429`, slow down the client and retry after a delay.
# API Overview (/docs/scraper-api/v1)
The Geonode Scraper API helps you extract, discover, and process web content without managing browsers, proxies, or scraping infrastructure.
Send a URL and receive structured content as Markdown or HTML. The API also supports JavaScript rendering, geo-targeting, batch processing, website crawling, and webhook notifications.
## What You Can Do
With the Scraper API, you can:
* Extract content from webpages
* Process multiple URLs in batch jobs
* Crawl websites and discover pages
* Receive webhook notifications when jobs complete
* Use geo-targeted proxy routing
* Extract content from JavaScript-powered websites
* Retrieve links found on webpages
## Available APIs
Choose the API that best matches your use case.
| API | Use When |
| ---------- | -------------------------------------------------------- |
| Extraction | You already know the URL and want to extract its content |
| Batch | You have multiple URLs that need to be processed |
| Crawl | You want to discover and extract pages across a website |
| Webhooks | You want to receive notifications when jobs complete |
## When to Use the Scraper API
Use the Scraper API when you need:
* Clean Markdown or HTML output
* JavaScript rendering
* Geo-targeted extraction
* Managed proxy infrastructure
* Batch processing
* Website crawling
* Asynchronous processing
* Webhook notifications
If you need complete control over browser automation, request handling, or custom scraping logic, consider using the Geonode Proxy API instead.
## Getting Started
If you're new to the Scraper API, start with the Quick Start Guide.
The Quick Start Guide covers:
1. Creating an API key
2. Sending your first request
3. Understanding responses
4. Exploring available APIs
## Next Steps
Continue to the **Quick Start Guide** to make your first request and explore the Scraper API.
# Reference (/docs/scraper-api/v1/reference)
Use the API Reference when you already know what you want to build and need the exact endpoint for a specific operation.
If you're new to the Scraper API, start with the Quick Start Guide.
## Base URL
All API requests use the following base URL:
```text
https://scraper.geonode.io
```
You can store it as an environment variable:
```bash title="request.sh"
export SCRAPER_API_BASE_URL="https://scraper.geonode.io"
```
## Authentication
All Scraper API requests require an API key.
Send your API key using the `X-Api-Key` request header.
```http
X-Api-Key: YOUR_API_KEY
```
You can store the API key as an environment variable:
```bash title="request.sh"
export GEONODE_SCRAPER_API_KEY="YOUR_API_KEY"
```
Keep your API key private and never expose it in frontend applications, public repositories, screenshots, or logs.
## API Categories
The Scraper API is organized into the following categories.
| Category | Purpose |
| ---------- | ------------------------------------------- |
| Extraction | Extract content from webpages |
| Batch | Process multiple URLs in a single job |
| Crawl | Discover and extract pages across a website |
| Map | Discover URLs from a website |
| Statistics | Retrieve usage statistics |
| Webhooks | Receive notifications when jobs complete |
| System | Service health and status |
## Endpoints
### System
| Method | Endpoint | Description |
| ------ | --------- | ----------------------- |
| `GET` | `/health` | Check API health status |
### Extraction
| Method | Endpoint | Description |
| ------ | ---------------------- | ------------------------------ |
| `POST` | `/v1/extract` | Extract content from a webpage |
| `GET` | `/v1/extract/{job_id}` | Retrieve an extraction job |
| `GET` | `/v1/extract/jobs` | List extraction jobs |
### Batch
| Method | Endpoint | Description |
| -------- | -------------------- | -------------------- |
| `POST` | `/v1/batch` | Start a batch job |
| `GET` | `/v1/batch/{job_id}` | Retrieve a batch job |
| `DELETE` | `/v1/batch/{job_id}` | Cancel a batch job |
### Map
| Method | Endpoint | Description |
| ------ | --------- | ---------------------------- |
| `POST` | `/v1/map` | Discover URLs from a website |
### Crawl
| Method | Endpoint | Description |
| -------- | -------------------- | -------------------- |
| `POST` | `/v1/crawl` | Start a crawl job |
| `GET` | `/v1/crawl/{job_id}` | Retrieve a crawl job |
| `DELETE` | `/v1/crawl/{job_id}` | Cancel a crawl job |
### Statistics
| Method | Endpoint | Description |
| ------ | ---------------- | ------------------------- |
| `GET` | `/v1/statistics` | Retrieve usage statistics |
### Webhooks
| Method | Endpoint | Description |
| -------- | ----------------------------------------- | ----------------------- |
| `POST` | `/v1/webhooks` | Create a webhook |
| `GET` | `/v1/webhooks` | List webhooks |
| `GET` | `/v1/webhooks/{webhook_id}` | Retrieve a webhook |
| `PATCH` | `/v1/webhooks/{webhook_id}` | Update a webhook |
| `DELETE` | `/v1/webhooks/{webhook_id}` | Delete a webhook |
| `POST` | `/v1/webhooks/{webhook_id}/rotate-secret` | Rotate a webhook secret |
| `GET` | `/v1/webhooks/{webhook_id}/deliveries` | List webhook deliveries |
## Next Steps
Choose the endpoint category that matches your use case and open its guides or endpoint reference pages.
# Some Apps Are Not Using the Proxy (/docs/proxies/additional-resources/fixing-common-issues/apps-bypassing-proxy)
Some applications may bypass proxy configurations, causing inconsistent behavior or incorrect IP routing.\
This guide explains why that happens and how to fix it.
***
## Common Symptoms
| Issue | Description |
| --------------------------- | -------------------------------------------------------------------------------- |
| **App ignores proxy** | The app connects directly to the internet instead of routing through your proxy. |
| **Inconsistent IP results** | IP changes work in the browser but not in the app. |
| **Authentication errors** | The app doesn’t support proxy login credentials. |
***
## Steps to Fix
### 1. Use a Proxy-Compatible Browser
Some built-in browsers (like Edge WebView or internal app browsers) ignore system proxy settings.\
✅ Try using **Google Chrome**, **Mozilla Firefox**, or **Brave**, which fully support proxies.
***
### 2. Check App Proxy Settings
Certain apps have their **own proxy configuration** separate from the system.\
🛠 Open the app’s **Network** or **Connection** settings and enter your proxy details manually.
***
### 3. Confirm Proxy Authentication Type
If your proxy requires **username/password authentication**, make sure the app supports it.\
Some mobile or desktop apps only support IP whitelisting instead.
***
### 4. Use a VPN as an Alternative
If the app completely ignores proxy settings, use a **VPN** service instead.\
VPNs tunnel all traffic system-wide, ensuring every app is routed through the same network path.
***
## FAQs
Some apps do not support proxy connections or require manual configuration.
{" "}
Only by using a VPN or a system-level proxy manager that intercepts all
connections.
Usually no — browsers like Chrome and Firefox respect system or manual proxy settings.
***
## Summary
* Some apps ignore system proxy settings by design.
* Always check if the app has its own proxy configuration field.
* Use Chrome or Firefox for consistent proxy testing.
* If nothing works, switch to a VPN or a proxy manager for full traffic control.
# Frequent Authentication Pop-Ups (/docs/proxies/additional-resources/fixing-common-issues/authentication-popups)
If you keep getting authentication pop-ups while using a proxy — like this:
— it usually means there’s an issue with authentication or IP whitelisting.\
Follow these steps to fix it.
***
## Troubleshooting Steps
### 1. Whitelist Your IP Address
If your IP is not whitelisted, the proxy will repeatedly ask for credentials.\
👉 [How to Whitelist My IP](/docs/proxies/getting-started/setup_and_configuration/advance-configuration/whitelist-ip)
***
### 2. Check Your Login Credentials
Ensure that you’re using the correct **username** and **password** for your Geonode proxy.\
Incorrect or expired credentials will trigger the authentication window every time.
***
### 3. Use an Authentication-Supported App
Some browsers or apps do not handle proxy authentication correctly.\
✅ Try using **Google Chrome**, **Mozilla Firefox**, or another modern browser that supports proxy logins.
***
## FAQs
Your IP may not be whitelisted, or you might be using an incompatible browser.
No, authentication is required for secure proxy usage.\
To avoid pop-ups, whitelist your IP so the proxy doesn’t ask for credentials.
***
## Summary
* Repeated pop-ups usually mean your IP isn’t whitelisted.
* Double-check your proxy username and password.
* Use a browser that supports proxy authentication.
* Whitelisting your IP eliminates most repeated login prompts.
# Proxy Connection Keeps Dropping (/docs/proxies/additional-resources/fixing-common-issues/connection-dropping)
If your proxy connection keeps disconnecting or timing out, this guide will help you identify the cause and fix it quickly.
***
## Troubleshooting Steps
### 1. Ensure Your Internet Connection Is Stable
Unstable or weak internet can cause frequent proxy disconnections.\
✅ Try switching to a wired connection or move closer to your Wi-Fi router.
***
### 2. Reconfigure Proxy Settings
Incorrect or outdated proxy settings may cause the connection to drop.\
🛠 Go to your **Network or Wi-Fi Settings**, remove the existing proxy configuration, and re-enter your **Geonode proxy details**.
***
### 3. Restart Your Router
If you’re on a home or shared Wi-Fi network, restart your router.\
This helps clear temporary network conflicts that might interrupt proxy communication.
***
### 4. Test With Another App or Browser
Try connecting your proxy in a different browser or app.\
If the issue doesn’t repeat, it may be caused by the original app’s network configuration.
***
## FAQs
The most common causes are unstable internet connections or incorrect proxy configuration.
{" "}
Sometimes — switching to a different port can bypass temporary routing issues.
Not necessarily. Check if your connection is stable and test the same proxy on another device before assuming downtime.
***
## Summary
* Check and stabilize your internet connection.
* Re-enter your proxy credentials and configuration.
* Restart your router to clear temporary issues.
* Test on a different app or browser to isolate the problem.
# Proxy Not Changing My IP (/docs/proxies/additional-resources/fixing-common-issues/ip-not-changing)
If your IP address doesn’t change after setting up a proxy, it’s usually due to app configuration or proxy setup issues.\
Follow these steps to verify and fix the problem.
***
## Troubleshooting Steps
### 1. Check Your IP on a Verification Site
Visit [ShowMyIP](https://www.showmyip.com/) or [ip-api.com](https://ip-api.com/) to confirm whether your IP has actually changed.\
Sometimes, caching or DNS delays can make it look like your IP is the same.
***
### 2. Restart Your Browser or App
After updating proxy settings, restart your browser or app to apply the configuration.\
⚙️ Some apps only load proxy settings at startup — a restart ensures they take effect.
***
### 3. Ensure Your App Supports Proxies
Certain applications bypass system proxy settings entirely.\
✅ Try testing your proxy using **Google Chrome**, **Firefox**, or another proxy-compatible browser.
***
### 4. Confirm Manual Configuration
Double-check that your proxy details were entered correctly under **Manual Proxy Setup** (IP, Port, Username, and Password).\
A missing field or typo can prevent the proxy from activating.
***
### 5. Test Another Proxy Type
If your IP still doesn’t change, switch between **HTTP** and **SOCKS5** proxies to see if the app or service supports one better than the other.
***
## FAQs
Some apps bypass system proxy settings. Test using a proxy-compatible browser like Chrome or Firefox.
{" "}
Yes — some websites detect your previous session via cookies or cache, even
after the IP changes.
Yes. If a VPN is active, it overrides proxy routing. Disable VPNs before testing proxy connections.
***
## Summary
* Verify your IP with an external site like ShowMyIP or ip-api.com.
* Restart your app to apply new proxy settings.
* Ensure the proxy is configured correctly and supported by your application.
* Disable VPNs or conflicting network tools before testing.
# No Save Option in Proxy Settings (/docs/proxies/additional-resources/fixing-common-issues/no-save-option)
If your device doesn’t show a **“Save”** button when setting up a proxy, don’t worry — some systems handle proxy saving automatically.\
Follow the steps below to confirm your settings are applied correctly.
***
## Troubleshooting Steps
### 1. Exit Settings After Configuration
On many devices, proxy changes are **auto-saved** as soon as you leave the settings page.\
Try exiting the menu normally — the configuration is likely already active.
***
### 2. Restart Your Device
If changes don’t take effect immediately, restart your device.\
A quick reboot often forces the system to apply pending network configurations.
***
### 3. Try Another Network
Some Wi-Fi networks or administrators **block proxy changes** for security reasons.\
Connect to a different network and try setting up the proxy again.
***
## FAQs
Some devices automatically save proxy configurations when you exit the settings menu.
{" "}
Usually no — but restarting can help apply the settings if the proxy doesn’t
activate immediately.
Yes. Most operating systems save proxy settings silently once you leave the configuration screen.
***
## Summary
* Many systems auto-save proxy settings when you close the menu.
* Restarting your device applies the new configuration if it doesn’t activate right away.
* If settings still don’t apply, try switching to a different Wi-Fi network.
# Pages Not Loading After Proxy Setup (/docs/proxies/additional-resources/fixing-common-issues/pages-not-loading)
If web pages fail to load or load slowly after setting up a proxy, this guide will help you identify and fix the issue.
***
## Troubleshooting Steps
### 1. Check Proxy Server Status
Make sure your proxy is **active and running** on the [Geonode Dashboard](https://app.geonode.com/).\
If it’s inactive or expired, your connection requests won’t go through.
***
### 2. Disable and Re-enable the Proxy
Sometimes, reapplying the configuration helps.\
Go to your **Wi-Fi or Network Settings**, disable the proxy, save changes, and then enable it again.
***
### 3. Clear Browser Cache
Cached or outdated data can conflict with new proxy settings.\
🧹 Clear your browser cache and cookies, then restart the browser or app.
***
### 4. Try a Different Network
Some Wi-Fi networks — especially in offices or public places — block proxy connections via firewalls.\
✅ Test your proxy setup using a different Wi-Fi or mobile hotspot.
***
### 5. Test the Proxy
Visit [ShowMyIP](https://www.showmyip.com/) or [ip-api.com](https://ip-api.com/) to confirm whether your proxy is active and your IP address has changed.
***
## FAQs
Some websites block certain proxy servers or IP ranges. Try switching to another proxy location or type (HTTP/SOCKS5).
{" "}
Visit [ShowMyIP](https://www.showmyip.com/) or
[ip-api.com](https://ip-api.com/) to confirm your proxy IP and location.
Yes. Cached DNS and cookies can store old network data that conflicts with new proxy settings.
***
## Summary
* Verify your proxy is active and properly configured.
* Reapply proxy settings if pages aren’t loading.
* Clear your browser cache and restart the app.
* Test with another network or proxy location if the issue persists.
# Proxy Not Working (/docs/proxies/additional-resources/fixing-common-issues/proxy-not-working)
If your proxy isn’t working or failing to connect, follow these steps to diagnose and fix the issue.
***
## Troubleshooting Steps
### 1. Check Proxy Details
Make sure you’ve entered the **correct Proxy IP and Port** from your [Geonode Dashboard](https://app.geonode.com/).\
⚠️ A small typo in the IP or port number can prevent the connection entirely.
***
### 2. Restart Your Device
After configuring the proxy, **restart your device** to ensure the settings are properly applied.
***
### 3. Switch Networks
Try switching between **Wi-Fi and mobile data**.\
Some networks — especially corporate or public ones — block proxy traffic for security reasons.
***
### 4. Verify Proxy Compatibility
Not all applications support proxies equally.\
✅ Test your proxy with a browser like **Google Chrome** or **Firefox** to confirm it works as expected.\
If it does, the issue might be app-specific.
***
### 5. Check Proxy Authentication
If your proxy requires credentials, double-check your **username** and **password**.\
Incorrect authentication can prevent the proxy from connecting.
***
## FAQs
Ensure that you entered the correct Proxy IP, Port, Username, and Password. Also, confirm that the proxy is active in your Geonode Dashboard.
{" "}
Yes. In most cases, you’ll need to manually enter proxy details in your
device’s or browser’s network settings.
Some apps bypass system proxy settings. Try using a proxy-compatible browser or configure the proxy directly within the app.
***
## Summary
* Double-check Proxy IP, Port, and authentication credentials.
* Restart your device to apply settings.
* Try a different network or browser to test compatibility.
* If the proxy works elsewhere, the issue is likely app-specific.
# Troubleshooting Proxy Issues on Android (/docs/proxies/additional-resources/fixing-common-issues/troubleshooting-common-issue)
Setting up a proxy on Android can surface the same issues covered in other troubleshooting guides (no connection, pages not loading, repeated auth prompts, etc.). To avoid duplication, this page highlights only Android-specific checks and points you to the detailed articles for each issue.
***
## Quick Android Checks
* Ensure the proxy is set under **Wi‑Fi → Advanced → Proxy → Manual** for the network you’re using.
* Some Android apps bypass system proxies; test with **Chrome** or **Firefox** first.
* Toggle Wi‑Fi off/on or reboot the device after changing proxy settings.
* If using IP whitelist, confirm your current IP is added in the dashboard.
***
## Issue Index (Android)
* **Proxy not working** — follow the main guide and re-check Wi‑Fi manual proxy config on Android: [Proxy Not Working](/docs/proxies/additional-resources/fixing-common-issues/proxy-not-working)
* **Pages not loading** — see the primary steps for slow/no load; on Android also re-apply the proxy on the current Wi‑Fi: [Pages Not Loading](/docs/proxies/additional-resources/fixing-common-issues/pages-not-loading)
* **Frequent authentication pop-ups** — confirm IP whitelist or credentials, and use a browser that supports auth prompts: [Authentication Popups](/docs/proxies/additional-resources/fixing-common-issues/authentication-popups)
* **IP not changing** — verify the active network has the proxy set and test with a browser: [IP Not Changing](/docs/proxies/additional-resources/fixing-common-issues/ip-not-changing)
* **No Save option in proxy settings** — many Android builds auto-save when you back out of settings: [No Save Option](/docs/proxies/additional-resources/fixing-common-issues/no-save-option)
* **Connection keeps dropping** — apply the general stability steps and re-enter proxy details on your Wi‑Fi: [Connection Dropping](/docs/proxies/additional-resources/fixing-common-issues/connection-dropping)
* **Apps bypassing proxy** — some apps ignore system proxies; use proxy-aware browsers or app-level proxy fields: [Apps Bypassing Proxy](/docs/proxies/additional-resources/fixing-common-issues/apps-bypassing-proxy)
***
## Conclusion
Most Android proxy issues are resolved by confirming Wi‑Fi manual proxy settings, using a proxy-aware app, and applying the detailed steps in the linked articles above. If problems persist, visit the [Geonode Documentation](/) or contact **Geonode Support**.
# Overview (/docs/proxies/api-reference/geo-targeting/geo-targeting-options)
Geo-targeting is one of the most powerful features of the Geonode Proxy API. It allows you to route your proxy requests through specific geographic locations, giving you precise control over where your traffic appears to originate from. This is essential for location-specific testing, content access, market research, and compliance with regional requirements.
## What is Geo-Targeting?
Geo-targeting enables you to specify the geographic location of the IP address that will be used for your proxy requests. Instead of getting a random IP from anywhere in the world, you can target:
* **Countries**: Route traffic through specific countries (e.g., United States, United Kingdom, Germany)
* **States/Regions**: Narrow down to specific states or provinces within a country
* **Cities**: Target specific cities for even more precise location control
* **ISPs/ASNs**: Route through specific Internet Service Providers or Autonomous System Numbers
## Why Use Geo-Targeting?
Geo-targeting allows you to route your proxy requests through specific geographic locations, giving you control over where your traffic appears to originate from.
## Targeting Levels
Geonode supports multiple levels of geo-targeting, from broad to highly specific:
### Country-Level Targeting
The broadest level of targeting. Simply append `-country-` to your username to route traffic through a specific country. This is ideal when you need traffic from a particular country but don't need more specific location control.
**Example**: `username-country-US` routes traffic through the United States.
### State-Level Targeting
For countries with states or provinces, you can target specific regions. This is useful when you need traffic from a particular state but don't need city-level precision.
**Example**: `username-country-US-state-california` routes traffic through California.
You cannot target both state and city at the same time. Choose either
state-level or city-level targeting for your requests.
### City-Level Targeting
The most precise geographic targeting option. Target specific cities within a country for maximum location accuracy.
**Example**: `username-country-US-city-newyork` routes traffic through New York City.
### ISP/ASN-Level Targeting
For advanced use cases, you can target specific Internet Service Providers or Autonomous System Numbers. This is useful when you need traffic from a particular ISP or network infrastructure.
**Example**: `username-type-residential-country-US-asn-12345` routes traffic through a specific ASN in the United States.
## Available Endpoints
This section provides endpoints for different geo-targeting options:
* **[Perform Country Targeting](/docs/proxies/api-reference/geo-targeting/get-country)**: Route traffic through specific countries
* **[Perform State Targeting](/docs/proxies/api-reference/geo-targeting/get-state)**: Target specific states or regions
* **[Perform City Targeting](/docs/proxies/api-reference/geo-targeting/get-city)**: Target specific cities
* **[Perform ISP/ASN Targeting](/docs/proxies/api-reference/geo-targeting/get-isp)**: Route through specific ISPs or ASNs
## Finding Available Locations
Before you can target a location, you need to know what locations are available. Use the [Retrieve Available Geo-locations](/docs/proxies/api-reference/available-geo-locations) endpoint to get a comprehensive list of:
* Available countries and their codes
* Cities within each country
* States/regions within each country
* ISPs and ASNs available in each location
## Best Practices
Before targeting a location, verify that it's available using the available geo-locations endpoint. Choose the appropriate level of targeting for your needs—use country-level if you don't need more specific location control.
Use ISO 3166-1 alpha-2 country codes (e.g., `US`, `GB`, `DE`) for country
targeting. City and state names should match the exact format provided in the
available locations list.
# Perform City Targeting (/docs/proxies/api-reference/geo-targeting/get-city)
Route your proxy requests through a specific city by including both the country code and city name in your username string.
You cannot target both state and city at the same time. Choose either
state-level or city-level targeting.
Append `-city-` after `-country-` to target a specific city.
For example: `username-country-US-city-newyork` targets New York City in the United States.
## Request
```bash
curl --request GET \
-x "http://proxy.geonode.io:" \
--user "-country--city-:" \
--url "http://ip-api.com/json"
```
## Response
### 200 Success
Returns detailed geolocation information about the IP address assigned to your proxy connection.
#### Response Fields
| Field | Type | Description |
| --------------- | ------ | ------------------------------------------- |
| `status` | string | The status of the request (e.g., "success") |
| `continent` | string | The name of the continent |
| `continentCode` | string | The two-letter continent code |
| `country` | string | The full name of the country |
| `countryCode` | string | The two-letter country code |
| `region` | string | The region code |
| `regionName` | string | The full name of the region |
| `city` | string | The name of the city |
| `district` | string | The district name, if available |
| `zip` | string | The postal code associated with the IP |
| `query` | string | The queried IP address |
#### Example Response
```json
{
"status": "success",
"continent": "North America",
"continentCode": "NA",
"country": "United States",
"countryCode": "US",
"region": "AL",
"regionName": "Alabama",
"city": "Decatur",
"district": "",
"zip": 35601,
"query": "68.191.141.86"
}
```
# Perform Country Targeting (/docs/proxies/api-reference/geo-targeting/get-country)
Route your proxy requests through a specific country by including the country code in your username string.
Append `-country-` after your `` to target a specific country.
For example: `username-country-US` targets the United States.
## Request
```bash
curl --request GET \
-x "http://proxy.geonode.io:" \
--user "-country-:" \
--url "http://ip-api.com/json"
```
## Response
### 200 Success
Returns detailed geolocation information about the IP address assigned to your proxy connection.
#### Response Fields
| Field | Type | Description |
| --------------- | ------- | ------------------------------------------------ |
| `status` | string | The status of the request (e.g., "success") |
| `continent` | string | The name of the continent |
| `continentCode` | string | The two-letter continent code |
| `country` | string | The full name of the country |
| `countryCode` | string | The two-letter country code (ISO 3166-1 alpha-2) |
| `region` | string | The region code |
| `regionName` | string | The full name of the region |
| `city` | string | The name of the city |
| `district` | string | The district name, if available |
| `zip` | string | The postal code associated with the IP |
| `lat` | number | The latitude coordinate |
| `lon` | number | The longitude coordinate |
| `timezone` | string | The timezone of the IP location |
| `offset` | integer | The time offset in seconds from UTC |
| `currency` | string | The currency code of the country |
| `isp` | string | The name of the Internet Service Provider |
| `org` | string | The organization that owns the IP address |
| `as` | string | The Autonomous System (AS) number |
| `query` | string | The queried IP address |
#### Example Response
```json
{
"status": "success",
"continent": "North America",
"continentCode": "NA",
"country": "Canada",
"countryCode": "CA",
"region": "QC",
"regionName": "Quebec",
"city": "Montreal",
"district": "",
"zip": "H2Y",
"lat": 45.5088,
"lon": -73.5878,
"timezone": "America/Toronto",
"offset": -18000,
"currency": "CAD",
"isp": "Bell Canada",
"org": "Bell Canada",
"as": "AS577 Bell Canada",
"query": "207.134.47.124"
}
```
# Perform ISP/ASN Targeting (/docs/proxies/api-reference/geo-targeting/get-isp)
Route your proxy requests through a specific ISP or Autonomous System Number (ASN) by including the ASN number in your username string.
The username format supports multiple targeting options:
* **Country**: `-country-` - Specify the country
* **ASN**: `-asn-` - Filter by ASN number
* **IP Type**: `-type-` - Choose residential, datacenter, or mix
Refer to our [Geo-Locations List](https://docs.geonode.com/docs/proxies/api-reference/available-geo-locations) to find available ASN numbers and locations.
## Request
```bash
curl -x "http://proxy.geonode.io:" \
--user "-type--country--asn-:" \
--url "http://ip-api.com/json"
```
## Response
### 200 Success
Returns detailed geolocation information about the IP address assigned to your proxy connection.
#### Response Fields
| Field | Type | Description |
| --------------- | ------- | ------------------------------------------------------ |
| `status` | string | The status of the request (e.g., "success") |
| `continent` | string | The continent where the IP is located |
| `continentCode` | string | The continent code |
| `country` | string | The country where the IP is registered |
| `countryCode` | string | The country code in ISO 3166-1 alpha-2 format |
| `region` | string | The regional subdivision (state/province) |
| `regionName` | string | The full name of the region |
| `city` | string | The city associated with the IP address |
| `district` | string | The district or subdivision of the city |
| `zip` | string | The postal or ZIP code of the location |
| `lat` | number | Latitude coordinate of the location |
| `lon` | number | Longitude coordinate of the location |
| `timezone` | string | Time zone in which the IP is located |
| `offset` | integer | Time offset from UTC in seconds |
| `currency` | string | Local currency used in the country |
| `isp` | string | The name of the Internet Service Provider (ISP) |
| `org` | string | The name of the organization associated with the IP |
| `as` | string | The Autonomous System (AS) number and name |
| `asname` | string | The full Autonomous System (AS) name |
| `mobile` | boolean | Indicates whether the IP is from a mobile network |
| `proxy` | boolean | Indicates whether the IP is being used as a proxy |
| `hosting` | boolean | Indicates whether the IP belongs to a hosting provider |
| `query` | string | The IP address queried in the request |
#### Example Response
```json
{
"status": "success",
"continent": "Europe",
"continentCode": "EU",
"country": "Russia",
"countryCode": "RU",
"region": "VGG",
"regionName": "Volgograd Oblast",
"city": "Volgograd",
"district": "",
"zip": "",
"lat": 48.5044,
"lon": 44.5838,
"timezone": "Europe/Volgograd",
"offset": 14400,
"currency": "RUB",
"isp": "JSC ER-Telecom Holding Volgograd branch",
"org": "JSC Columbia-Telecom",
"as": "AS50543 JSC ER-Telecom Holding",
"asname": "SARATOV-AS",
"mobile": false,
"proxy": false,
"hosting": false,
"query": "83.167.79.185"
}
```
# Perform OS Targeting (/docs/proxies/api-reference/geo-targeting/get-os-targeting)
Route your requests through proxy IPs associated with devices running a specific operating system by adding the `-os-` parameter to your proxy username.
Append `-os-` to your proxy username to target devices running a specific operating system.
For example, `-os-ios-` routes traffic through iOS devices.
## Supported Operating Systems
The following operating systems are supported:
| Operating System | Username Value |
| ---------------- | -------------- |
| Windows | `windows` |
| Android | `android` |
| iOS | `ios` |
| macOS | `mac` |
## Request
```bash
curl --request GET \
-x "http://proxy.geonode.io:10009" \
--user "geonode_username-os-ios-lifetime-30-session-randomIOS:YOUR_PASSWORD" \
--url "http://ip-api.com/json"
```
### Example
```bash
curl -x proxy.geonode.io:10009 \
-U geonode_username-os-ios-lifetime-30-session-randomIOS:YOUR_PASSWORD \
http://ip-api.com/json
```
In this example, the operating system target is specified in the proxy username:
```text
geonode_username-os-ios-lifetime-30-session-randomIOS
```
The important part is:
```text
-os-ios-
```
You can replace `ios` with any of the supported operating systems:
```text
-os-windows-
-os-android-
-os-ios-
-os-mac-
```
## Response
### 200 Success
Returns information about the IP address assigned to your proxy connection.
#### Response Fields
| Field | Type | Description |
| ------------- | ------ | -------------------------------------------- |
| `status` | string | Status of the request. |
| `country` | string | Country of the proxy IP address. |
| `countryCode` | string | Two-letter country code. |
| `region` | string | Region or state code. |
| `regionName` | string | Full name of the region or state. |
| `city` | string | City associated with the proxy IP. |
| `zip` | string | ZIP or postal code. |
| `lat` | number | Latitude of the proxy IP. |
| `lon` | number | Longitude of the proxy IP. |
| `timezone` | string | Time zone of the proxy IP. |
| `isp` | string | Internet service provider. |
| `org` | string | Organization associated with the IP address. |
| `as` | string | Autonomous System (AS) information. |
| `query` | string | Public IP address returned by the proxy. |
#### Example Response
```json
{
"status": "success",
"country": "United States",
"countryCode": "US",
"region": "PA",
"regionName": "Pennsylvania",
"city": "Philadelphia",
"zip": "19133",
"lat": 39.9934,
"lon": -75.1425,
"timezone": "America/New_York",
"isp": "T-Mobile USA, Inc.",
"org": "T-Mobile USA, Inc.",
"as": "AS21928 T-Mobile USA, Inc.",
"query": "172.56.216.58"
}
```
# Perform State Targeting (/docs/proxies/api-reference/geo-targeting/get-state)
Route your proxy requests through a specific state or region by including both the country code and state name in your username string.
You cannot target both state and city at the same time. Choose either
state-level or city-level targeting.
Append `-state-` after `-country-` to target a specific state.
For example: `username-country-US-state-california` targets California in the United States.
## Request
```bash
curl --request GET \
-x "http://proxy.geonode.io:" \
--user "-country--state-:" \
--url "http://ip-api.com/json"
```
## Response
### 200 Success
Returns detailed geolocation information about the IP address assigned to your proxy connection.
#### Response Fields
| Field | Type | Description |
| --------------- | ------ | ------------------------------------------- |
| `status` | string | The status of the request (e.g., "success") |
| `continent` | string | The name of the continent |
| `continentCode` | string | The two-letter continent code |
| `country` | string | The full name of the country |
| `countryCode` | string | The two-letter country code |
| `region` | string | The region code |
| `regionName` | string | The full name of the region |
| `city` | string | The name of the city |
| `district` | string | The district name, if available |
| `zip` | string | The postal code associated with the IP |
| `query` | string | The queried IP address |
#### Example Response
```json
{
"status": "success",
"continent": "North America",
"continentCode": "NA",
"country": "Canada",
"countryCode": "CA",
"region": "QC",
"regionName": "Quebec",
"city": "Granby",
"district": "",
"zip": "J2H",
"query": "192.168.1.1"
}
```
# Perform strict matching (/docs/proxies/api-reference/geo-targeting/perform-strict-matching)
### Overview
Previously, when you requested an IP address from a specific geolocation and none were available, the system would silently fall back to a similar location in the same country. We are changing this behavior to make it explicit and give you more control:
1. **Strict Matching (Default)**
* By default, if your requested geolocation is unavailable, you will receive an error indicating that no IP address is available (`proxy-error`).
* This means the system will **not** automatically provide an alternative IP from another location.
2. **Fallback with Flag**
* If you **want** to allow fallback (i.e., you want an IP from another location in the same country when your exact requested location is unavailable), you must explicitly enable it by setting the flag `-strict-off`.
### URL Flag Usage
You can now include one of the following flags in your username string:
* `-strict-on`
* Forces strict matching. If there are no available IPs for the specified location, the system returns a `proxy-error`.
* This is the **default** if you do **not** specify any strict flag.
* `-strict-off`
* Allows fallback. If the exact location is unavailable, the system will return another IP from the same country (if available).
#### Syntax
```
geonode_-type--country--asn--strict-on:
```
or
```
geonode_-type--country--asn--strict-off:
```
> **Note:**
>
> * You can apply `-strict-on` or `-strict-off` to any combination of geo-targeting flags (e.g., `country`, `city`, `state`, `asn`), but we recommend always including a `-country-` in your query.
> * The `asn` flag must be used **with** the `country` flag; otherwise, you will receive an error.
### Example Usage
#### 1. Fallback Enabled (`-strict-off`)
If you want to **allow** fallback to a different location in the same country when your requested location is unavailable:
```bash
curl -x : \
-U "geonode_username-type-residential-country-fr-asn-3215-strict-off:your_password" \
--url "http://ip-api.com"
```
* Here, you requested a French (`country-fr`) IP that belongs to ASN `3215`.
* With `-strict-off`, if no IP is available exactly for ASN 3215 in France, the system will return a **different** IP from France.
#### 2. Strict Matching (`-strict-on`)
If you want **strict** matching for the exact country/ASN combination:
```bash
curl -x : \
-U "geonode_username-type-residential-country-fr-asn-3215-strict-on:your_password" \
--url "http://ip-api.com"
```
* If no IP is available that exactly matches France + ASN 3215, the system will return:
```json
{ "proxy-error": "country-fr-asn-3215 target was not found" }
```
### Error Cases
1. **No Exact Location Found (Strict Matching)**
* **Error**: `{"proxy-error":"country-fr-asn-3215 target was not found"}`
* Occurs when using `-strict-on` (or default strict matching) and the system cannot find an IP for that location.
2. **`asn` Without `country`**
* **Error**: `{"proxy-error":"ASN should be used with country."}`
* The `asn` flag must be paired with a specific country code.
3. **`-strict-on` or `-strict-off` Without Any Geo-Targeting Flag**
* **Error**: If you use `-strict-on` or `-strict-off` **without** specifying any location or ASN.
* You must include at least one geo-targeting parameter (e.g. `-country-xx`, `-asn-xxxxx`, etc.) for the request to make sense.
### Summary
* **Default**: Strict matching is enforced (the system does **not** fall back to another location).
* **Use `-strict-off`**: to allow fallback to a similar location within the same country.
* **Always Pair `asn` With `country`**: `-asn-50543` must be accompanied by `-country-xx`.
* **Error Handling**: You will receive JSON error responses if no IP is found under strict conditions or if flags are incorrectly combined.
By explicitly controlling strict matching via `-strict-on` or `-strict-off`, you can now decide whether to **always** require a specific location or to **allow** fallback within the same country in case your requested location is unavailable.
# Perform ASN Exclusion (/docs/proxies/api-reference/exclude-targeting/exclude-asn)
Route your proxy requests while excluding specific ASNs (Autonomous System Numbers) from the routing.
You cannot mix ASN exclusions with other exclusion types (country, city, or
state) in the same request. Use only one exclusion type per request.
Append `-not.asn-,` after your ``.
* **Single ASN**: `-not.asn-31898`
* **Multiple ASNs**: `-not.asn-31898,12345,67890`
* **With country targeting**: `-country-us-not.asn-31898`
## Request
```bash
curl --request GET \
-x "http://proxy.geonode.io:9000" \
--user "-country-us-not.asn-:" \
--url "http://ip-api.com/json"
```
## Response
### 200 Success
Returns detailed geolocation information about the IP address, excluding the specified ASNs.
#### Response Fields
| Field | Type | Description |
| ------------- | ------ | ------------------------------------------- |
| `status` | string | The status of the request (e.g., "success") |
| `country` | string | The full name of the country |
| `countryCode` | string | The two-letter country code |
| `region` | string | The region code |
| `regionName` | string | The full name of the region |
| `city` | string | The name of the city |
| `zip` | string | The postal code associated with the IP |
| `lat` | number | The latitude coordinate |
| `lon` | number | The longitude coordinate |
| `timezone` | string | The timezone of the IP location |
| `isp` | string | The name of the Internet Service Provider |
| `org` | string | The organization that owns the IP address |
| `as` | string | The Autonomous System (AS) number |
| `query` | string | The queried IP address |
#### Example Response
```json
{
"status": "success",
"country": "United States",
"countryCode": "US",
"region": "MA",
"regionName": "Massachusetts",
"city": "Springfield",
"zip": "01101",
"lat": 42.0986,
"lon": -72.5931,
"timezone": "America/New_York",
"isp": "RingSquared CC",
"org": "",
"as": "AS7849 RingSquared CC",
"query": "161.77.215.31"
}
```
# Perform City Exclusion (/docs/proxies/api-reference/exclude-targeting/exclude-city)
Route your proxy requests while excluding specific cities from the routing.
You can only exclude cities. You cannot combine city exclusions with country,
state, or ASN exclusions in the same request.
Append `-not.city-,` after your ``.
* **Single city**: `-not.city-charlotte`
* **Multiple cities**: `-not.city-charlotte,newyork,houston`
* **With country**: `-country-US-not.city-charlotte,newyork`
## Request
```bash
curl --request GET \
-x "http://proxy.geonode.io:" \
--user "-country--not.city-,:" \
--url "http://ip-api.com/json"
```
## Response
### 200 Success
Returns detailed geolocation information about the IP address, excluding the specified cities.
#### Response Fields
| Field | Type | Description |
| ------------- | ------ | ------------------------------------------- |
| `status` | string | The status of the request (e.g., "success") |
| `country` | string | The full name of the country |
| `countryCode` | string | The two-letter ISO country code |
| `region` | string | The region code |
| `regionName` | string | The full name of the region |
| `city` | string | The name of the city |
| `zip` | string | The postal code associated with the IP |
| `lat` | number | The latitude coordinate |
| `lon` | number | The longitude coordinate |
| `timezone` | string | The timezone of the IP location |
| `isp` | string | The name of the Internet Service Provider |
| `org` | string | The organization that owns the IP address |
| `as` | string | The Autonomous System (AS) number |
| `query` | string | The queried IP address |
#### Example Response
```json
{
"status": "success",
"country": "United States",
"countryCode": "US",
"region": "NC",
"regionName": "North Carolina",
"city": "Charlotte",
"zip": "28202",
"lat": 35.2327,
"lon": -80.8461,
"timezone": "America/New_York",
"isp": "FiberPower LLC",
"org": "FiberPower LLC",
"as": "AS214483 FiberPower LLC",
"query": "38.13.166.129"
}
```
# Perform Country Exclusion (/docs/proxies/api-reference/exclude-targeting/exclude-country)
Route your proxy requests while excluding specific countries from the routing.
You can only exclude countries. You cannot combine country exclusions with
city, state, or ASN exclusions in the same request.
Append `-not.country-,` after your ``.
* **Single country**: `-not.country-US`
* **Multiple countries**: `-not.country-US,CA,MX`
## Request
```bash
curl --request GET \
-x "http://proxy.geonode.io:" \
--user "-not.country-,:" \
--url "http://ip-api.com/json"
```
## Response
### 200 Success
Returns detailed geolocation information about the IP address, excluding the specified countries.
#### Response Fields
| Field | Type | Description |
| ------------- | ------ | ------------------------------------------- |
| `status` | string | The status of the request (e.g., "success") |
| `country` | string | The full name of the country |
| `countryCode` | string | The two-letter ISO country code |
| `region` | string | The region code |
| `regionName` | string | The full name of the region |
| `city` | string | The name of the city |
| `zip` | string | The postal code associated with the IP |
| `lat` | number | The latitude coordinate |
| `lon` | number | The longitude coordinate |
| `timezone` | string | The timezone of the IP location |
| `isp` | string | The name of the Internet Service Provider |
| `org` | string | The organization that owns the IP address |
| `as` | string | The Autonomous System (AS) number |
| `query` | string | The queried IP address |
#### Example Response
```json
{
"status": "success",
"country": "Mexico",
"countryCode": "MX",
"region": "AGU",
"regionName": "Aguascalientes",
"city": "Aguascalientes",
"zip": "20326",
"lat": 21.9419,
"lon": -102.2756,
"timezone": "America/Mexico_City",
"isp": "Uninet S.A. de C.V.",
"org": "UNINET",
"as": "AS8151 UNINET",
"query": "187.232.239.178"
}
```
# Perform State Exclusion (/docs/proxies/api-reference/exclude-targeting/exclude-state)
Route your proxy requests while excluding specific states or regions from the routing.
You can only exclude states. You cannot combine state exclusions with city,
country, or ASN exclusions in the same request.
Append `-not.state-,` after your ``.
* **Single state**: `-not.state-california`
* **Multiple states**: `-not.state-california,newyork,texas`
* **With country**: `-country-US-not.state-california,newyork`
## Request
```bash
curl --request GET \
-x "http://proxy.geonode.io:" \
--user "-country--not.state-,:" \
--url "http://ip-api.com/json"
```
## Response
### 200 Success
Returns detailed geolocation information about the IP address, excluding the specified states.
#### Response Fields
| Field | Type | Description |
| ------------- | ------ | ------------------------------------------- |
| `status` | string | The status of the request (e.g., "success") |
| `country` | string | The full name of the country |
| `countryCode` | string | The two-letter ISO country code |
| `region` | string | The region code |
| `regionName` | string | The full name of the region |
| `city` | string | The name of the city |
| `zip` | string | The postal code associated with the IP |
| `lat` | number | The latitude coordinate |
| `lon` | number | The longitude coordinate |
| `timezone` | string | The timezone of the IP location |
| `isp` | string | The name of the Internet Service Provider |
| `org` | string | The organization that owns the IP address |
| `as` | string | The Autonomous System (AS) number |
| `query` | string | The queried IP address |
#### Example Response
```json
{
"status": "success",
"country": "United States",
"countryCode": "US",
"region": "NY",
"regionName": "New York",
"city": "Queens",
"zip": "11436",
"lat": 40.6744,
"lon": -73.8016,
"timezone": "America/New_York",
"isp": "Charter Communications",
"org": "Spectrum",
"as": "AS12271 Charter Communications Inc",
"query": "72.227.174.55"
}
```
# Remove Whitelisted IPs (/docs/proxies/api-reference/whitelisting-ip/delete)
You can remove up to 10 IP addresses per request.
Remove IP addresses from your whitelist. Once removed, these IPs will no longer have automatic access to your proxy services.
## Request
```bash
curl -X DELETE "https://app-api.geonode.com/api/configuration/active/whitelisted-ips" \
-H "Authorization: Basic base64(username:password)" \
-H "Content-Type: application/json" \
-d '{"ips":[{"ip":"161.142.148.121"},{"ip":"161.142.148.147"}]}'
```
### Request Body Parameters
| Field | Type | Required | Description |
| ---------- | ------ | -------- | --------------------------------------- |
| `ips` | array | Yes | List of IP objects to remove |
| `ips[].ip` | string | Yes | The IP address to remove from whitelist |
### Example Request Body
```json
{
"ips": [
{
"ip": "161.142.148.121"
},
{
"ip": "161.142.148.147"
}
]
}
```
## Response
### 200 Success
The IP addresses have been successfully removed from your whitelist.
#### Response Fields
| Field | Type | Description |
| -------------------- | ------ | ------------------------------------------------------ |
| `data` | array | A list of whitelisted IPs that have been removed |
| `data[].ip` | string | The IP address that was removed from the whitelist |
| `data[].description` | string | The user-defined label for the removed IP |
| `data[]._id` | string | A unique identifier assigned to the removed IP |
| `message` | object | Provides status details about the removal operation |
| `message.title` | string | A short message indicating the status of the operation |
| `message.body` | string | A detailed message about the removal operation |
| `message.variant` | string | The status variant indicating success or failure |
#### Example Response
```json
{
"data": [
{
"ip": "161.142.148.141",
"description": "mac-m1",
"_id": "67a875f693afe58d40f2e93d"
}
],
"message": {
"title": "Updated",
"body": "Whitelisted IPs removed.",
"variant": "success"
}
}
```
### Error Responses
#### 400 Bad Request
Invalid input data provided.
| Field | Type | Description |
| ------- | ------ | ------------- |
| `error` | string | Error message |
```json
{
"error": "Invalid input data."
}
```
#### 401 Unauthorized
Invalid API credentials provided.
| Field | Type | Description |
| ------- | ------ | ------------- |
| `error` | string | Error message |
```json
{
"error": "Unauthorized - Invalid credentials."
}
```
# Retrieve Whitelisted IPs (/docs/proxies/api-reference/whitelisting-ip/get)
Retrieve all IP addresses currently on your whitelist.
## Request
```bash
curl -X GET "https://app-api.geonode.com/api/configuration/active/whitelisted-ips" -u "username:apiKey" \
-H "Authorization: Basic base64(username:password)"
```
## Response
### 200 Success
Returns a list of all active whitelisted IP addresses.
#### Response Fields
| Field | Type | Description |
| -------------------- | ------ | ---------------------------------------------- |
| `data` | array | List of whitelisted IP addresses |
| `data[].ip` | string | The IP address of the whitelisted entity |
| `data[].description` | string | A brief description associated with the IP |
| `data[]._id` | string | The unique identifier assigned to the IP entry |
#### Example Response
```json
{
"data": [
{
"ip": "161.142.148.150",
"description": "updated-description",
"_id": "67a0a5bbb412ddaf6f139b3b"
}
]
}
```
### Error Responses
#### 400 Bad Request
Invalid parameters provided.
| Field | Type | Description |
| ------- | ------ | ------------- |
| `error` | string | Error message |
```json
{
"error": "Invalid parameters provided."
}
```
#### 401 Unauthorized
Invalid API credentials provided.
| Field | Type | Description |
| ------- | ------ | ------------- |
| `error` | string | Error message |
```json
{
"error": "Unauthorized - Invalid API key or credentials."
}
```
#### 500 Internal Server Error
An internal server error occurred.
# Add Whitelisted IPs (/docs/proxies/api-reference/whitelisting-ip/post)
You can add up to 10 IP addresses per request.
Add IP addresses to your whitelist to allow them to access your Geonode proxy services without authentication.
## Request
```bash
curl -X POST "https://app-api.geonode.com/api/configuration/active/whitelisted-ips" \
-H "Authorization: Basic base64(username:password)" \
-H "Content-Type: application/json" \
-d '{"ips": [{"ip": "161.142.148.140", "description": "mac-m1"}]}'
```
### Request Body Parameters
| Field | Type | Required | Description |
| ------------------- | ------ | -------- | --------------------------------------------- |
| `ips` | array | Yes | List of IP objects to add to the whitelist |
| `ips[].ip` | string | Yes | The IP address being added to the whitelist |
| `ips[].description` | string | No | A user-defined label or identifier for the IP |
### Example Request Body
```json
{
"ips": [
{
"ip": "161.142.148.140",
"description": "mac-m1"
}
]
}
```
## Response
### 200 Success
The IP addresses have been successfully added to your whitelist.
#### Response Fields
| Field | Type | Description |
| -------------------- | ------ | -------------------------------------------------- |
| `data` | array | A list of successfully whitelisted IP addresses |
| `data[].ip` | string | The IP address that was whitelisted |
| `data[].description` | string | A user-defined label or identifier for the IP |
| `data[]._id` | string | A unique identifier assigned to the whitelisted IP |
| `message` | object | Details about the status of the whitelist update |
| `message.title` | string | A short title summarizing the update status |
| `message.body` | string | A message detailing the outcome of the operation |
| `message.variant` | string | The status variant indicating success or failure |
#### Example Response
```json
{
"data": [
{
"ip": "161.142.148.150",
"description": "updated-description",
"_id": "67a0a5bbb412ddaf6f139b3b"
}
],
"message": {
"title": "Updated",
"body": "Whitelisted IPs saved.",
"variant": "success"
}
}
```
### Error Responses
#### 400 Bad Request
Invalid input data provided.
| Field | Type | Description |
| ------- | ------ | ------------- |
| `error` | string | Error message |
```json
{
"error": "Invalid input data."
}
```
#### 401 Unauthorized
Invalid API credentials provided.
| Field | Type | Description |
| ------- | ------ | ------------- |
| `error` | string | Error message |
```json
{
"error": "Unauthorized - Invalid API key or credentials."
}
```
# Update Whitelisted IP Description (/docs/proxies/api-reference/whitelisting-ip/put)
Update the description or label associated with a whitelisted IP address.
## Request
```bash
curl -X PUT "https://app-api.geonode.com/api/configuration/active/whitelisted-ips" \
-H "Authorization: Basic base64(username:password)" \
-H "Content-Type: application/json" \
-d '{"ip":"161.142.148.150","description":"updated-description"}'
```
### Request Body Parameters
| Field | Type | Required | Description |
| ------------- | ------ | -------- | ------------------------------------------ |
| `ip` | string | Yes | The IP address to update |
| `description` | string | Yes | The new description for the whitelisted IP |
### Example Request Body
```json
{
"ip": "161.142.148.150",
"description": "updated-description"
}
```
## Response
### 200 Success
The IP description has been successfully updated.
#### Response Fields
| Field | Type | Description |
| -------------------- | ------ | --------------------------------------------------- |
| `data` | array | A list of updated whitelisted IPs |
| `data[].ip` | string | The IP address that was updated |
| `data[].description` | string | The updated description for the whitelisted IP |
| `data[]._id` | string | A unique identifier assigned to the whitelisted IP |
| `message` | object | Provides status details about the update operation |
| `message.title` | string | A short message indicating the status of the update |
| `message.body` | string | A detailed message about the update operation |
| `message.variant` | string | The status variant indicating success or failure |
#### Example Response
```json
{
"data": [
{
"ip": "161.142.148.150",
"description": "just update",
"_id": "67a0a5bbb412ddaf6f139b3b"
}
],
"message": {
"title": "Updated",
"body": "Whitelisted IPs updated.",
"variant": "success"
}
}
```
### Error Responses
#### 400 Bad Request
Invalid input data provided.
| Field | Type | Description |
| ------- | ------ | ------------- |
| `error` | string | Error message |
```json
{
"error": "Invalid input data."
}
```
#### 401 Unauthorized
Invalid API credentials provided.
| Field | Type | Description |
| ------- | ------ | ------------- |
| `error` | string | Error message |
```json
{
"error": "Unauthorized - Invalid API key or credentials."
}
```
# Overview (/docs/proxies/api-reference/whitelisting-ip/whitelisting-ips)
IP whitelisting is a security feature that allows you to control which IP addresses can access your Geonode proxy services without requiring authentication. This is particularly useful for securing your proxy infrastructure and ensuring that only authorized IPs can connect to your account.
## What is IP Whitelisting?
When you whitelist an IP address, you're essentially creating a trusted list of IPs that can bypass the standard authentication process. This is ideal for scenarios where:
* You have a fixed server or application that always connects from the same IP
* You want to enhance security by restricting access to specific IP addresses
* You need to simplify authentication for automated systems
* You want to prevent unauthorized access from unknown locations
## Key Features
You can manage up to 10 IP addresses per request, making it easy to bulk
update your whitelist.
* **Add IPs**: Add up to 10 IP addresses at once with optional descriptions for easy identification
* **List IPs**: Retrieve all currently whitelisted IP addresses for your account
* **Update Descriptions**: Modify the labels associated with whitelisted IPs for better organization
* **Remove IPs**: Remove IP addresses from your whitelist when they're no longer needed
## Use Cases
IP whitelisting is useful when you need to allow specific IP addresses to access your proxy services without authentication. Common scenarios include automated scripts, server applications, and controlled access environments.
## Getting Started
To start using IP whitelisting, you'll need to:
1. **Add IPs to your whitelist** using the [Add Whitelisted IPs](/docs/proxies/api-reference/whitelisting-ip/post) endpoint
2. **View your current whitelist** with the [Retrieve Whitelisted IPs](/docs/proxies/api-reference/whitelisting-ip/get) endpoint
3. **Update IP descriptions** as needed using the [Update Whitelisted IP Description](/docs/proxies/api-reference/whitelisting-ip/put) endpoint
4. **Remove IPs** when they're no longer needed via the [Remove Whitelisted IPs](/docs/proxies/api-reference/whitelisting-ip/delete) endpoint
Once an IP is whitelisted, it can access your proxy services without
authentication. Make sure to only whitelist trusted IP addresses and regularly
review your whitelist to remove any IPs that are no longer needed.
# ASN/ISP Targeting (/docs/proxies/getting-started/knowledge-base/asn-isp-targeting)
This guide explains what ASN/ISP targeting is and how to use it in Geonode for precise proxy control and optimized network performance.
## What Is ASN/ISP Targeting?
**ASN/ISP targeting** allows you to filter proxy connections based on an Internet Service Provider (ISP) or **Autonomous System Number (ASN)**.\
This helps you choose proxies from specific ISPs to improve reliability, compliance, and performance for targeted use cases.
## How ASN/ISP Targeting Works
* Every ISP is assigned a unique ASN (Autonomous System Number).
* Geonode enables users to select proxies by **country** and **ASN**.
* This is especially useful for:
* Accessing geo-restricted content
* Market research
* Increasing connection consistency and anonymity
### Example: Targeting a Specific ASN
To target a specific ISP, include both the country code and the ASN in your request:
```text
country-us-asn-7018
```
# Proxy Endpoint Formats (/docs/proxies/getting-started/knowledge-base/endpoint-formats)
This guide explains what proxy endpoints are, the types available in Geonode, and how to choose the most suitable format for your specific use case.
## What Is a Proxy Endpoint?
A **proxy endpoint** is the address used to connect to a proxy server.\
It defines how your requests are routed through Geonode’s network, ensuring secure, reliable, and efficient data transmission.
## DNS vs. IP-Based Endpoints
Geonode allows users to connect either via a **DNS-based hostname** or a **direct IP address**.
### 1. DNS-Based Endpoints (Recommended)
* Use a domain name such as `proxy.geonode.io`.
* Easier to maintain — DNS automatically updates IPs when servers change.
* Reduces connection issues caused by IP rotation.
### 2. IP-Based Endpoints
* Use a direct IP address, e.g. `123.45.67.89:9000`.
* Skips DNS resolution, which can be slightly faster in some setups.
* Ideal for apps or devices that **don’t support DNS hostnames**.
## Available Endpoint Formats
Geonode supports multiple endpoint formats to ensure compatibility with different systems and authentication methods.
### 1. `hostname:port`
* Example: `proxy.geonode.io:9000`
* Best for: Simple connections without authentication.
### 2. `hostname:port:username:password`
* Example: `proxy.geonode.io:9000:user:pass`
* Includes authentication credentials for secure access.
* Best for: APIs and applications requiring basic authentication.
### 3. `hostname:port@username:password`
* Example: `proxy.geonode.io:9000@user:pass`
* Alternative authentication format.
* Best for: Legacy or custom proxy clients.
### 4. `username:password@hostname:port`
* Example: `user:pass@proxy.geonode.io:9000`
* Credentials placed before the host for compatibility.
* Best for: Apps that authenticate before establishing a connection.
### 5. `http://username:password@server:port`
* Example: `http://user:pass@proxy.geonode.io:9000`
* Explicit HTTP/HTTPS proxy format.
* Best for: Secure browsing and authenticated API requests.
## Choosing the Right Proxy Endpoint Format
| Format | Best For |
| -------------------------------------- | --------------------------------------------- |
| `hostname:port` | Standard connections without authentication |
| `hostname:port:username:password` | Secure authentication for APIs & applications |
| `hostname:port@username:password` | Custom or legacy proxy configurations |
| `username:password@hostname:port` | Systems needing pre-authentication |
| `http://username:password@server:port` | Secure browsing, authenticated API access |
## Tips and Best Practices
* **Start simple:** Use `hostname:port` if no authentication is required.
* **Prefer DNS:** It’s more reliable and updates automatically.
* **Use IP endpoints** only for tools that can’t resolve hostnames.
* **Always use HTTPS** or encrypted channels when handling sensitive credentials.
* Test different formats — some apps accept only specific syntaxes.
## Summary
Proxy endpoints define how your connection routes through Geonode’s network.\
Choosing the correct format ensures:
* Smooth compatibility with your app or client,
* Secure authentication where required,
* Stable performance with minimal downtime.
# Gateways (/docs/proxies/getting-started/knowledge-base/gateway)
This guide explains what proxy gateways are and how Geonode uses them to optimize proxy routing, improve performance, and maintain anonymity.
## What Is a Proxy Gateway?
A **proxy gateway** is a server that routes your internet traffic through a specific geographic location.\
By selecting a gateway, you define the **entry point** for your proxy requests — controlling how and where your traffic appears to originate.
Geonode provides multiple gateway locations to:
* Access region-restricted content
* Improve browsing speed and stability
* Maintain privacy and network anonymity
## Geonode’s Available Gateways
Geonode currently offers the following gateway locations:
* **France**
* **United States**
* **Singapore**
Each gateway routes your traffic through a regional hub, providing faster speeds and greater accessibility for local or restricted online services.
## Why Use a Gateway?
### 1. Access Geo-Restricted Content
* Browse the internet as if you’re in another country
* Useful for streaming, e-commerce, and international research
### 2. Improve Connection Speed
* Routes traffic through optimized, low-latency locations
* Reduces delay and improves access to nearby content
### 3. Enhance Privacy and Security
* Masks your real IP and location
* Adds a layer of protection when using public or sensitive networks
## Choosing the Right Gateway
| Gateway | Best For |
| ----------------- | -------------------------------------------------------------- |
| **France** | EU-based content, GDPR-compliant data collection |
| **United States** | Streaming, U.S. market research, and e-commerce |
| **Singapore** | Low-latency connections in Asia, accessing region-locked sites |
## Tips and Best Practices
* **Use the closest gateway** to minimize latency and maximize performance.
* **Select gateways strategically** based on your target region or content source.
* **Switch gateways** if your connection feels slow — load balancing can vary by region.
## Summary
Proxy gateways determine where your connection enters Geonode’s network.\
By selecting the right gateway, you can:
* Improve connection speed,
* Access region-specific content,
* And enhance your overall privacy and anonymity online.
# Geo-Targeting (/docs/proxies/getting-started/knowledge-base/geo-targeting)
This guide explains what Geo-Targeting is and how Geonode allows you to choose proxies based on country, state, and city for precise connection control and regional flexibility.
## What Is Geo-Targeting?
**Geo-Targeting** lets you filter and select proxy locations by **country**, **state**, and **city**.\
It enables businesses, developers, and marketers to access content and services as if they were browsing directly from a specific geographic area.
## How Geo-Targeting Works in Geonode
Geonode supports **three levels of geo-targeting**, shown in the example below:
1. **Country Targeting** – Select a country for your proxy connection.
2. **State Targeting** – Choose a particular state or region within that country.
3. **City Targeting** – Narrow it down further by selecting a specific city, or set it to **Any** for broader access.
## Why Use Geo-Targeting?
### 1. Access Geo-Restricted Content
* Bypass region-based restrictions on websites, apps, and streaming services.
* View local search engine results, prices, and ads as a user from that area.
### 2. Improve Localized Testing and Marketing
* Run **ad verification** and A/B tests by region.
* Test **localized websites, apps, and payment systems** from multiple markets.
### 3. Ensure Compliance and Security
* Simulate user behavior from different regions for compliance testing.
* Avoid detection by using realistic, location-accurate IPs.
## Choosing the Right Geo-Targeting Option
| Geo-Targeting Level | Best For |
| --------------------- | --------------------------------------------------------------------- |
| **Country Targeting** | General browsing, international research, region-based content access |
| **State Targeting** | Regional services, ad verification, localized e-commerce |
| **City Targeting** | Precision testing, local SEO, hyper-targeted ad campaigns |
## Notes and Recommendations
* The availability of **state** and **city** targeting depends on the selected country.
* Some cities may have fewer proxy IPs — use **“Any”** for better coverage.
* For faster performance, select a location closer to your target audience or testing region.
## Best Practices
* Use **Country Targeting** for broad, region-specific access.
* Choose **State** or **City Targeting** for more precise, location-based control.
* Start with **Country Targeting** and refine your selection as needed for campaigns or tests.
## Summary
Geo-Targeting gives you granular control over your proxy location — from country down to city level.\
By choosing the right targeting depth, you can:
* Access localized content,
* Run accurate regional tests, and
* Optimize performance for specific markets.
# IP Address (/docs/proxies/getting-started/knowledge-base/ip-address)
This guide explains what an IP address is, how it functions, and why it’s important when using proxies.
## What Is an IP Address?
An **IP address (Internet Protocol Address)** is a unique identifier assigned to every device connected to the internet.\
Think of it as a mailing address for your device — without it, data wouldn’t know where to go.
When you visit a website, your IP address tells the site where to send the information you requested.\
It can also reveal your **approximate location**, **internet provider**, and **network type**.
## How IP Addresses Work
* Every device gets an IP address from its **Internet Service Provider (ISP)**.
* When you make a request (like opening a website), your IP acts as a return address.
* The server sends data back to that address — completing the connection.
## Types of IP Addresses
There are several types of IPs depending on their use, assignment method, and format.
### 1. Public vs. Private IP Addresses
| Type | Description |
| ---------------------- | --------------------------------------------------------------------------- |
| **Public IP Address** | Assigned by your ISP and used to communicate over the internet. |
| **Private IP Address** | Used within local networks (e.g., home Wi-Fi) to identify internal devices. |
### 2. Static vs. Dynamic IP Addresses
| Type | Description |
| ---------------------- | -------------------------------------------------------------------- |
| **Static IP Address** | Fixed and unchanging. Common for servers, businesses, and VPNs. |
| **Dynamic IP Address** | Changes periodically. Used by most home connections for flexibility. |
### 3. IPv4 vs. IPv6
| Type | Description |
| -------- | -------------------------------------------------------------------------------- |
| **IPv4** | Classic format with four number sets, e.g. `192.168.1.1`. Still the most common. |
| **IPv6** | Newer format supporting many more devices, e.g. `2001:db8::ff00:42:8329`. |
## Why IP Addresses Matter in Proxies
When using proxies, your **real IP** is replaced with another one — masking your location and identity.\
This provides privacy, allows region-based access, and reduces the risk of detection or blocking.
### 1. Residential vs. Datacenter IPs
| Type | Description |
| ------------------- | ------------------------------------------------------------------------------- |
| **Residential IPs** | Provided by ISPs to real devices. Highly trusted and less likely to be flagged. |
| **Datacenter IPs** | Generated by data centers or hosting services. Faster, but easier to detect. |
### 2. Rotating vs. Sticky IPs
| Type | Description |
| ---------------- | ------------------------------------------------------------------------------------ |
| **Rotating IPs** | Change with every request or at timed intervals — ideal for scraping and automation. |
| **Sticky IPs** | Remain constant for a session — best for logins, account management, or testing. |
## How to Check Your IP Address
You can easily find your public IP:
* Search **“What is my IP”** on Google.
* Visit [whatismyip.com](https://www.whatismyip.com).
* Check network details in your router or device settings.
## Tips and Best Practices
* Use proxies to protect your identity and location online.
* **Residential IPs** → more privacy and authenticity.
* **Datacenter IPs** → more speed and cost efficiency.
* **IPv6** is expanding, but **IPv4** remains dominant for most users.
* Choose a reliable provider like **Geonode** to balance **security, anonymity, and performance**.
## Summary
An IP address is your device’s online identifier — essential for routing internet traffic.\
Using a proxy lets you control or hide that identity, giving you:
* More **privacy**,
* Better **regional access**, and
* Stronger **protection** online.
# IP Types (/docs/proxies/getting-started/knowledge-base/ip-type)
This guide explains what IP types are and how Geonode provides them for different proxy use cases.
## What Are IP Types?
**IP types** define how an IP address is sourced and used within proxy networks.\
Geonode offers three main categories:
* **Residential IPs** – Real-user IPs assigned by Internet Service Providers (ISPs).
* **Datacenter IPs** – IPs generated by third-party data centers.
* **Mixed IPs** – A combination of Residential and Datacenter IPs for flexibility.
Each serves a unique purpose — from bypassing geo-restrictions to powering high-speed automation.
## 1. Residential IPs
### What Are Residential IPs?
Residential IPs come from **real user connections** provided by ISPs.\
Because they look like genuine home users, websites see them as legitimate traffic.
### Best For
* Web scraping with minimal detection risk
* Accessing geo-restricted content
* Managing e-commerce or social media accounts
* Secure streaming and gaming
### Pros
✔️ High trust level — less likely to be blocked\
✔️ Works on strict, detection-sensitive sites
### Cons
❌ Slower than datacenter IPs\
❌ More expensive due to limited availability
## 2. Datacenter IPs
### What Are Datacenter IPs?
Datacenter IPs are **server-based IPs** generated by hosting providers, not ISPs.\
They offer higher speed and lower cost but are easier to detect.
### Best For
* Bulk web scraping and data collection
* SEO monitoring and automation
* High-performance or bot-driven tasks
### Pros
✔️ Fast and reliable\
✔️ Cost-effective for large-scale operations
### Cons
❌ Easier to detect and block on strict websites\
❌ Limited success with region-locked or sensitive content
## 3. Mixed IPs
### What Are Mixed IPs?
Mixed IPs combine both **Residential** and **Datacenter** sources — balancing realism and performance.
### Best For
* Scalable web scraping and automation
* Testing and development environments
* Multi-purpose e-commerce or social media tasks
### Pros
✔️ Balanced speed, security, and price\
✔️ Greater flexibility across use cases
### Cons
❌ Not as anonymous as pure Residential IPs\
❌ Some tasks may require dedicated IP types
## Comparison: Residential vs Datacenter vs Mixed
| Feature | Residential IPs | Datacenter IPs | Mixed IPs |
| ------------- | ----------------------------------- | ------------------------------ | ------------------------------------------------ |
| **Speed** | Moderate | Fast | Balanced |
| **Anonymity** | High (real users) | Low (easily detected) | Moderate |
| **Cost** | Expensive | Affordable | Mid-range |
| **Best For** | Geo access, social media, streaming | Bulk scraping, SEO, automation | Versatile tasks needing both reliability & speed |
## How to Choose the Right IP Type
* **Residential IPs** → for high anonymity and geo-specific access.
* **Datacenter IPs** → for fast, large-scale automation.
* **Mixed IPs** → for a balanced approach between performance and stealth.
## Tips and Best Practices
* Start with **Mixed IPs** to test both speed and reliability.
* Use **Residential IPs** for sensitive or geo-restricted sites.
* Use **Datacenter IPs** for speed-critical automation or scraping.
* Always verify which type performs best for your target websites.
## Summary
Each IP type serves a distinct purpose:
* **Residential** for trust and authenticity,
* **Datacenter** for speed and scale,
* **Mixed** for flexibility.
Choosing the right type ensures stable, efficient, and secure proxy performance for your specific needs.
# IP Whitelisting (/docs/proxies/getting-started/knowledge-base/ip-whitelist)
**IP whitelisting** is a security feature that allows only approved IP addresses to connect to your network or service.\
It blocks unauthorized access and ensures safer connections.
When using proxies, **whitelisted IPs** let you connect without a password — making access both **simple** and **secure**.
## Benefits of IP Whitelisting
* **Better Security** – Only trusted IPs can connect.
* **More Control** – You decide who can access your system.
* **Easy Access** – No need for manual logins or credentials.
## Best Practices
To maximize security and efficiency, follow these recommendations:
* **Add Only Trusted IPs** – Whitelist secure, verified networks.
* **Use Clear Labels** – Name IPs (e.g., “Home,” “Office”) for easy identification.
* **Review Regularly** – Remove unused or outdated IPs.
* **Combine with Other Tools** – Use firewalls and authentication for layered protection.
## Troubleshooting Common Issues
Ensure your proxy IP is whitelisted — connections may fail if it isn’t.
Add your new IP address from your current location to regain access.
* Verify that you entered the correct **public IP address**.
* Check if your IP has changed and update it in the whitelist.
## Summary
**IP whitelisting** strengthens your proxy setup by restricting access to trusted sources only.\
It simplifies authentication and prevents unauthorized use — ideal for both personal and enterprise environments.
For a detailed setup guide, see:\
➡️ [How to Whitelist Your IP Address](/docs/proxies/getting-started/setup_and_configuration/advance-configuration/whitelist-ip)
# Overview (/docs/proxies/getting-started/knowledge-base/overview)
Welcome to the **Proxy Service Guide** — your complete reference for understanding and configuring Geonode proxies.
This guide covers all the essential concepts for effective proxy usage, including:
* **IP Addresses** – Learn how IPs work and why they matter in proxy networks.
* **Proxy Types** – Understand the differences between residential, datacenter, and mixed IPs.
* **Authentication** – Explore IP whitelisting, username/password access, and security best practices.
* **Geo-Targeting** – Configure country, state, and city-level targeting.
* **Ports & Sessions** – Manage connections, rotation, and session persistence.
By mastering these topics, you’ll be able to:
* Optimize connection performance,
* Maintain high security and privacy,
* Avoid detection across different platforms and use cases.
Explore each section to gain a clear, practical understanding of how Geonode proxies work — and how to use them efficiently for your specific needs.
# Ports (/docs/proxies/getting-started/knowledge-base/port-type)
This guide explains what ports are, how they work, and why they play a key role in proxy connections.
## What Is a Port?
Imagine your device as an apartment building (your **IP address**) with many rooms inside — these rooms are **ports**.\
Each room serves a different purpose, helping data reach the right application.
* Every device connected to the internet has a unique IP address.
* **Ports** act as doorways that direct data traffic to the correct app or service.
* Each service (browsing, messaging, streaming, etc.) uses its own port number to avoid confusion.
Without ports, your device wouldn’t know which application incoming data belongs to — everything would arrive at the same “door.”
## How Ports Work in Practice
Let’s say you open several browser tabs at once:
* One tab has **WhatsApp Web**,
* Another shows **Facebook**,
* And the third streams **YouTube**.
When a WhatsApp message arrives, your computer needs to know which app it’s for.\
Ports make this possible:
* WhatsApp Web might use **port 5222**,
* Facebook might use **port 443**,
* YouTube might use another.
Your system reads the port number, matches it with the right app, and delivers the data correctly.\
When that session closes, the port becomes available again — these are called **ephemeral ports** (temporary ports for short-lived connections).
## Ports in Proxy Configuration
In proxy setups, ports function like **dedicated lanes** for your traffic.\
Different ports connect to different types of proxy behavior — such as rotating or sticky sessions.
### Common Proxy Ports in Geonode
| Proxy Type | Port Range |
| ------------------- | ----------- |
| **SOCKS5 Rotating** | 11000–11010 |
| **SOCKS5 Sticky** | 12000–12010 |
| **HTTP Rotating** | 9000–9010 |
| **HTTP Sticky** | 10000–10900 |
Each range corresponds to a specific connection type, giving you control over how often your proxy IP changes.
## How to Configure Proxy Ports
When setting up a proxy, choose a port according to your goal:
* **Rotating proxies** → Use ports from the *rotating* range to get a new IP for each request.
* **Sticky proxies** → Use ports from the *sticky* range to keep the same IP for a set duration.
➡️ For detailed setup steps, see:\
[Proxy Port Configuration](/docs/proxies/getting-started/setup_and_configuration/advance-configuration/port-configuration)
## Best Practices
* Treat each **port** as a separate communication channel.
* Choose **rotating** or **sticky** ports depending on your use case.
* Once a port is assigned to a specific country, **remove the assignment** before reusing it.
* Understanding ports helps ensure a stable and efficient proxy connection.
## Summary
**Ports** are the internal “routes” that direct internet traffic where it needs to go.\
In proxy configurations, they define how your connection behaves — rotating for dynamic IPs, sticky for consistent sessions.\
Selecting the right port ensures smooth, secure, and optimized proxy performance.
# Protocol Type (/docs/proxies/getting-started/knowledge-base/protocol-type)
Choosing the right **protocol type** is essential for optimizing proxy performance, ensuring security, and maintaining reliable connections.
This guide explains what a protocol is in the context of proxies and helps you decide between **HTTP/HTTPS** and **SOCKS5** for your specific use case.
## What Is a Protocol in Proxy Usage?
A **protocol** is a set of rules that defines how data is transmitted between devices across a network.\
When you use a proxy, the protocol determines **how requests are sent, processed, and returned**.
## Types of Proxy Protocols
Geonode supports two main protocol types:
* **HTTP / HTTPS**
* **SOCKS5**
Let’s explore how they differ — and when to use each.
***
## 1. HTTP / HTTPS Proxies
**HTTP/HTTPS proxies** handle standard web traffic.\
They forward web requests between your browser (or app) and the destination website.\
The **HTTPS** variant encrypts the connection, providing extra privacy and protection.
### Ideal For
* General web browsing
* APIs and web services
* Web scraping and automation
* Managing multiple online accounts
### Key Features
* Easy to configure — compatible with most browsers and apps.
* HTTPS ensures encrypted data transmission.
* Works seamlessly with tools like **Postman**, **cURL**, and browser extensions.
### Port Range
* **Rotating Proxy:** `9000–9010`
* **Sticky Proxy:** `10000–10900`
### Common Use Cases
* **Web Scraping:** Automate data collection while reducing detection risk.
* **Account Management:** Maintain consistent login sessions.
* **API Requests:** Manage large volumes of secure, authenticated requests.
***
## 2. SOCKS5 Proxies
**SOCKS5** is a more flexible and powerful proxy protocol.\
Unlike HTTP, it works at a lower network level and doesn’t alter the transmitted data — making it ideal for **non-web traffic** and privacy-focused applications.
### Ideal For
* Privacy and anonymity
* P2P (peer-to-peer) connections
* Bypassing firewalls and geo-restrictions
* Streaming, gaming, and VoIP
### Key Features
* Supports **both TCP and UDP**, enabling real-time communication.
* Offers **higher anonymity** — does not inject identifying headers.
* Handles all types of traffic: HTTP, FTP, VoIP, gaming, etc.
### Port Range
* **Rotating Proxy:** `11000–11010`
* **Sticky Proxy:** `12000–12010`
### Common Use Cases
* **Bypassing Restrictions:** Access blocked content securely.
* **Torrenting:** Stable and fast peer-to-peer transfers.
* **VoIP & Streaming:** Lower latency and improved connection quality.
***
## How to Choose the Right Protocol
| Criteria | **HTTP / HTTPS** | **SOCKS5** |
| --------------- | ------------------------------ | ----------------------------------- |
| **Security** | HTTPS encryption protects data | High anonymity, supports encryption |
| **Speed** | Fast for web-based traffic | Faster for non-HTTP traffic |
| **Flexibility** | Limited to web requests | Supports all internet protocols |
| **Best For** | Websites, APIs, automation | P2P, VoIP, bypassing firewalls |
***
## Summary
* Use **HTTP/HTTPS** for simplicity, compatibility, and secure web requests.
* Choose **SOCKS5** for advanced use cases, full traffic support, and stronger anonymity.
* If you’re just getting started, start with HTTP/HTTPS — you can always switch to SOCKS5 later as your needs evolve.
# Protocols (/docs/proxies/getting-started/knowledge-base/protocols)
Geonode’s proxy network supports several protocols, each designed for different use cases and levels of security.\
Choosing the right protocol ensures compatibility, performance, and privacy when connecting through Geonode.
## Supported Protocols
Geonode currently supports the following proxy protocols:
1. **HTTP** — Hypertext Transfer Protocol
2. **HTTPS** — Hypertext Transfer Protocol Secure
3. **SOCKS5** — Socket Secure version 5
Each protocol serves a specific purpose — from standard web browsing to high-security, high-speed connections.
***
## Protocol Comparison
| Protocol | Security Level | Best For | Availability |
| ---------- | ------------------------- | ----------------------------------------- | ------------ |
| **HTTP** | ❌ No encryption | General web scraping, browsing, APIs | ✅ Supported |
| **HTTPS** | ✅ SSL/TLS encrypted | Secure websites, automation, transactions | ✅ Supported |
| **SOCKS5** | ✅ High anonymity & secure | Streaming, gaming, bypassing restrictions | ✅ Supported |
***
## Unsupported Protocols
The following protocols are **not supported** by Geonode proxies:
* **UDP** (User Datagram Protocol)
* **Email protocols** — such as SMTP, IMAP, and POP3
These protocols operate differently from HTTP/HTTPS and SOCKS5 and are not part of Geonode’s proxy infrastructure.
***
## Best Practices
* Use **HTTP** for general data scraping and standard web requests.
* Choose **HTTPS** for secure, encrypted communication and automated workflows.
* Select **SOCKS5** for streaming, gaming, or tasks requiring maximum privacy and flexibility.
Geonode focuses on stable, encrypted proxy connections — therefore, **UDP and email-related traffic are excluded** for security and performance reasons.
***
## Summary
Geonode supports the three most widely used proxy protocols — **HTTP**, **HTTPS**, and **SOCKS5** — giving users the flexibility to choose between speed, compatibility, and security.\
Pick the protocol that fits your use case and connection needs to ensure optimal performance and reliability.
# Proxy Authentication Methods (/docs/proxies/getting-started/knowledge-base/proxy-authentication)
***
This guide will help you understand what the Proxy Authentication method is and the different ways geonode provides to authenticate it.
## **What is Proxy Authentication?**
Proxy authentication is a way to verify your identity before accessing a proxy server. It helps keep your connection secure and ensures only authorized users can use the proxy.
Geonode provides two ways to authenticate:
1. Using a username and password **(Basic Authentication)**
2. Whitelisting your IP address **(IP Authentication)**
Each method has its advantages, depending on your use case.
***
## **1. Authenticating with Username and Password (Basic Authentication)**
This method requires you to send your Geonode proxy username and password with every request. Many applications, APIs, and automation tools support this authentication method.
### **How Basic Authentication Works**
1. Format your credentials as:
```
username:password
```
2. Convert this string to Base64 format (for security).
3. Add the encoded credentials to your request header like this:
```
Authorization: Basic BASE64_ENCODED_STRING
```
➡️ For a full setup guide, check: [Set Up Proxy Authentication in API](/docs/proxies/getting-started/setup_and_configuration/advance-configuration/proxy-authentication-in-api)
***
## **2. Authenticating with IP Whitelisting**
IP whitelisting allows you to skip entering your username and password by authorizing a specific IP address to access the proxy.
➡️[What is IP Whitelisting?](/docs/proxies/getting-started/setup_and_configuration/advance-configuration/whitelist-ip)
### **When to use IP Whitelisting**
* You have a static IP and want seamless authentication.
* You are automating tasks and don't want to include credentials in every request.
* You want better security by restricting access to trusted IPs.
➡️ [Learn How to Whitelist Your IP Address](/docs/proxies/getting-started/setup_and_configuration/advance-configuration/whitelist-ip)
***
## **Choosing the Right Authentication Method**
| Authentication Method | Best For | Requires Credentials? |
| ----------------------- | -------------------------------------------------------- | --------------------- |
| **Username & Password** | Most API calls, scripts, and browser extensions | ✅ Yes |
| **IP Whitelisting** | Trusted networks, automation, and security-focused users | ❌ No |
***
## **Final Tips**
* Use Username & Password for flexible authentication across different devices.
* Use IP Whitelisting if you have a static IP and want to avoid entering credentials.
* Always ensure your authentication details are kept secure and not exposed in scripts or public repositories.
# Port Usage (/docs/proxies/getting-started/knowledge-base/proxy-usage)
This guide explains how ports function when configuring **SOCKS5** and **HTTP** proxies in Geonode.\
Geonode provides **unlimited port ranges**, meaning you can generate as many proxies as needed — without performance limits.
## How Port Ranges Work
Geonode assigns specific port ranges to organize different proxy types and session behaviors:
| Proxy Type | Session Type | Port Range |
| ---------- | ------------ | ------------- |
| **HTTP** | Rotating | `9000–9010` |
| **HTTP** | Sticky | `10000–10900` |
| **SOCKS5** | Rotating | `11000–11010` |
| **SOCKS5** | Sticky | `12000–12010` |
These ranges act as **identifiers**, not limitations.\
You can reuse the same port as often as you like — each connection request automatically generates a **unique proxy IP**.
***
## Key Points to Remember
### Same Performance Across All Ports
* Whether you use port `9000` or `9010`, performance and stability remain identical.
* Reusing a port does **not** affect connection speed or session quality.
### Efficient Port Reuse
* You don’t need to change ports for every connection.
* Thousands of requests can run on the **same port** without issues or conflicts.
* Port selection mainly helps organize your setup (for example, by type or location).
***
## Common Use Cases
### Using Multiple Proxies
If you need multiple proxies for automation or scraping:
* You can assign **all requests to the same port** (e.g., `9000`).
* Each request will still produce a **unique IP address**, ensuring anonymity and diversity.
* There’s no impact on speed or reliability.
### Rotating vs. Sticky Sessions
* **Rotating Proxies** — Ports `9000–9010` (HTTP) and `11000–11010` (SOCKS5):\
Each request gets a **new IP address** automatically.
* **Sticky Proxies** — Ports `10000–10900` (HTTP) and `12000–12010` (SOCKS5):\
The same IP remains active for a set duration (useful for account logins, testing, etc.).
### Country-Specific Port Assignments
* When you assign a port to a specific **country**, it cannot be reused for another country until the previous assignment is **deleted**.
* This ensures accurate geo-routing and prevents proxy conflicts.
***
## Common Misconceptions
> “Fewer ports mean fewer proxies.”
❌ False.\
The **number of ports** has no effect on how many proxies you can use.\
Ports are simply **access points** — Geonode dynamically assigns IPs behind them.
***
## Summary
* You can reuse the same port indefinitely — performance stays the same.
* Rotating ports change IPs automatically; sticky ports keep the same IP for a set session.
* Assign ports carefully if you’re working with geo-targeted proxies.
* Geonode’s architecture allows **unlimited proxy generation** across all ports.
By understanding how port usage works, you’ll be able to manage proxy sessions efficiently — without worrying about limits or performance degradation.
# Proxy (/docs/proxies/getting-started/knowledge-base/proxy)
This guide explains what a proxy is, how it works, and why it’s essential for online security, privacy, and automation.
## What Is a Proxy?
A **proxy server** acts as a bridge between your device and the internet.\
Instead of connecting directly to a website, your request first passes through the proxy, which forwards it on your behalf.
### Analogy: A Proxy as a Messenger
Imagine you’re ordering food but don’t want the restaurant to know your home address.\
You ask a **friend (proxy)** to pick it up and deliver it. The restaurant only sees your friend’s address — not yours.
Likewise, when using a proxy:
* Your real IP is hidden.
* The proxy server communicates with websites for you.
* Websites see the proxy’s IP instead of yours.
***
## How a Proxy Works
When you connect through a proxy, your data follows four simple steps:
1. **Request Sent** — You request access to a website.\
The request goes to the proxy server first.
2. **IP Replaced** — The proxy swaps your IP with its own and sends the request onward.
3. **Response Received** — The target website sends data back to the proxy.
4. **Response Delivered** — The proxy forwards that data back to you, keeping your IP hidden.
***
## Types of Proxies
### 1. Forward vs. Reverse Proxies
| Type | Description |
| ----------------- | -------------------------------------------------------------------------------------------------------- |
| **Forward Proxy** | Protects the user by hiding their IP when browsing or scraping the internet. |
| **Reverse Proxy** | Protects servers by managing inbound traffic, improving security and load balancing for hosted websites. |
***
### 2. Residential vs. Datacenter Proxies
| Type | Description |
| --------------------- | ----------------------------------------------------------------- |
| **Residential Proxy** | Uses real IPs from ISPs — highly trusted and difficult to detect. |
| **Datacenter Proxy** | Uses IPs from data centers — faster but easier to identify. |
***
### 3. Rotating vs. Sticky Proxies
| Type | Description |
| ------------------ | ------------------------------------------------------------------------------------------------- |
| **Rotating Proxy** | Changes IP for every request or after a set time — great for scraping and large-scale automation. |
| **Sticky Proxy** | Keeps the same IP for a session — best for logins, account management, and session-based tasks. |
***
## Why Use a Proxy?
### 1. Privacy & Anonymity
* Hides your real IP and online identity.
* Prevents tracking and profiling by websites.
### 2. Geo-Unblocking
* Access region-restricted content (streaming, marketplaces, apps).
* Simulate browsing from specific countries.
### 3. Security & Protection
* Avoid IP bans and throttling.
* Reduce exposure to threats by masking your origin.
### 4. Automation & Data Collection
* Collect public data at scale without detection.
* Power SEO, price-tracking, and analytics tools.
***
## How to Choose the Right Proxy
| Use Case | Recommended Proxy Type |
| ------------------------------- | ------------------------- |
| **Browsing & Privacy** | Residential Proxy |
| **Web Scraping & Automation** | Rotating Datacenter Proxy |
| **Geo-Blocked Content** | Residential Proxy |
| **Multiple Account Management** | Sticky Residential Proxy |
***
## Summary
* Proxies act as intermediaries that protect your identity online.
* Choose **rotating** proxies for scraping, **sticky** ones for stable sessions.
* Residential proxies offer trust and geo-access, while datacenter proxies focus on speed.
* Understanding proxy types helps you stay secure, anonymous, and efficient across all use cases.
# Rotating Proxies (/docs/proxies/getting-started/knowledge-base/rotating-proxies)
This guide explains what rotating proxies are, how they work, and how they help maintain anonymity and stability during high-volume requests.
## What Are Rotating Proxies?
A **rotating proxy** automatically assigns a new IP address after each request — or after a defined time interval.\
This rotation helps prevent detection and blocking by websites that monitor for repeated activity from a single IP.
***
## How Rotating Proxies Work
* Each outgoing request is assigned a **unique IP address** from a large proxy pool.
* The IP rotates **automatically** after every request or based on a timer (e.g., every few minutes).
* This makes requests appear as if they are coming from different users around the world.
Rotating proxies are ideal for maintaining anonymity and avoiding IP bans in automation and scraping tasks.
***
## Key Benefits
| Benefit | Description |
| ------------------------ | ------------------------------------------------------------- |
| **High Anonymity** | Constantly changing IPs prevent tracking and detection. |
| **Bypass Rate Limits** | Allows multiple requests without triggering anti-bot systems. |
| **Improved Scalability** | Enables large-scale web scraping and data collection. |
| **Global Coverage** | Access data from various geographic locations automatically. |
***
## Common Use Cases
| Use Case | Why It’s Useful |
| --------------------------- | ---------------------------------------------------------- |
| **Web Scraping** | Collects large datasets without detection or bans. |
| **Ad Verification** | Checks ad placements from multiple IPs and regions. |
| **Market Research** | Gathers pricing and trend data from competitor sites. |
| **SEO Monitoring** | Tracks rankings without triggering search engine security. |
| **Social Media Automation** | Manages multiple accounts without getting flagged. |
***
## Limitations
While rotating proxies offer flexibility and anonymity, they’re not ideal for all scenarios:
| Limitation | Impact |
| ---------------------------------------- | --------------------------------------------------------- |
| **Unstable Sessions** | Frequent IP changes break login-based activities. |
| **Detection by Strict Websites** | Some systems still recognize automated behavior. |
| **Potential for Inconsistent Responses** | Different IPs may yield varying localized or cached data. |
***
## Summary
* **Rotating proxies** provide high anonymity and scalability for automation and data collection.
* Use them for **web scraping**, **ad verification**, and **SEO monitoring**.
* For login-based or session-sensitive tasks, switch to **sticky proxies** to maintain consistency.
* Adjust rotation frequency to balance anonymity with connection stability.
# Session Type (/docs/proxies/getting-started/knowledge-base/session-type)
Choosing the right **session type** is key to optimizing proxy performance, maintaining stable connections, and avoiding detection.\
This guide explains what sessions are, how they work, and when to use **rotating** or **sticky** sessions.
## What Is a Session?
In proxy configuration, a **session** refers to the period during which you use the same IP address.\
The session type determines whether your IP remains constant (sticky) or changes frequently (rotating).
Geonode supports two main session types:
* **Rotating Sessions**
* **Sticky Sessions**
***
## Rotating Sessions
A **rotating session** assigns a new IP address for every request or after a defined time interval.\
This setup is designed for maximum anonymity and large-scale operations where detection risk is high.
### Key Features
| Feature | Description |
| ------------------------- | ----------------------------------------------------------- |
| **Automatic IP Rotation** | Changes IP after each request or at a set interval. |
| **High Anonymity** | Prevents blocks by cycling through multiple IPs. |
| **Scalable** | Ideal for bulk data collection and high-frequency requests. |
### Best For
* Web scraping and data aggregation
* Market or SEO research
* Ad verification and automation tasks
***
## Sticky Sessions
A **sticky session** maintains the same IP for a longer duration, creating a stable and continuous connection.\
This type is useful when session consistency or login persistence is required.
### Key Features
| Feature | Description |
| ---------------------------- | ------------------------------------------- |
| **Consistent IP Address** | Keeps the same IP throughout the session. |
| **Custom Duration** | Control how long the IP remains active. |
| **Reduced Reauthentication** | Avoids frequent logouts and session resets. |
### Best For
* Account management and login sessions
* E-commerce or transactional activities
* Application and website testing
***
## Comparison: Rotating vs. Sticky Sessions
| Criteria | Rotating Sessions | Sticky Sessions |
| --------------- | ----------------------------------------- | --------------------------------------- |
| **IP Behavior** | Changes with each request or set interval | Remains the same for a defined duration |
| **Anonymity** | High – ideal for stealth and scaling | Moderate – consistent IP for stability |
| **Performance** | Suited for high-volume tasks | Suited for persistent sessions |
| **Use Case** | Scraping, automation, data collection | Logins, transactions, testing |
***
## Summary
* **Rotating sessions** provide higher anonymity and flexibility for automation, scraping, and large-scale operations.
* **Sticky sessions** ensure reliability for logins, testing, and tasks that depend on stable IPs.
* Combine both session types when needed — for instance, rotating proxies for data gathering and sticky proxies for account-based operations.\
Understanding how sessions work helps you balance **anonymity, performance, and connection stability** for your specific goals.
# Sticky Session (/docs/proxies/getting-started/knowledge-base/sticky-session)
A **sticky session** maintains the same IP address for a fixed period instead of changing it with every request.\
This allows for stable, consistent sessions — ideal for activities like account logins, e-commerce transactions, and automation tasks that rely on persistent identity.
***
## How Sticky Sessions Work
* When you start a connection, a unique IP address is assigned to your session.
* The IP remains active for a defined duration (e.g., 10–30 minutes, or longer).
* Once the session expires, a new IP is automatically assigned on the next connection.
* You can manually release a session early if you need to refresh your IP before expiration.
Sticky sessions offer balance — they maintain stability without locking you to a single IP indefinitely.
***
## Key Benefits
| Benefit | Description |
| ----------------------------- | ---------------------------------------------------------------------- |
| **Stable Identity** | Keeps the same IP during the session, ideal for login-based workflows. |
| **Fewer CAPTCHA Challenges** | Reduces interruptions caused by frequent IP changes. |
| **Persistent Access** | Prevents session resets on e-commerce and social media platforms. |
| **Better Automation Control** | Ensures bots and tools operate smoothly across multi-step actions. |
***
## Common Use Cases
| Use Case | Why It’s Useful |
| ------------------------------ | ----------------------------------------------------- |
| **Account Management** | Maintains session stability and prevents logouts. |
| **E-Commerce & Checkout Bots** | Avoids disruptions during multi-step transactions. |
| **Web Scraping (Login Sites)** | Enables consistent access for authenticated scraping. |
| **Streaming Services** | Prevents interruptions or reauthentication requests. |
| **SEO Monitoring** | Keeps IP identity consistent for ongoing checks. |
***
## Limitations
| Limitation | Impact |
| ----------------------------- | ---------------------------------------------------------------------------- |
| **Overuse of One IP** | Extended use can make IPs easier to flag or block. |
| **Limited Parallel Requests** | Using the same IP across many tasks can reduce efficiency. |
| **Session Expiration** | Once the duration ends, the IP changes automatically, interrupting sessions. |
***
## Summary
* **Sticky sessions** provide reliable, consistent connections — perfect for logins, form submissions, and stable workflows.
* If an IP becomes blocked, you can **release the session** and instantly obtain a new one.
* For **high anonymity and frequent IP changes**, use **rotating sessions** instead.
* Choosing between sticky and rotating proxies depends on whether your task needs **stability or stealth**.
# Threads in Proxy Usage (/docs/proxies/getting-started/knowledge-base/thread)
Threads determine how many tasks can run at the same time when using proxies.\
Understanding how they work helps optimize performance, avoid bans, and make the most out of your proxy setup.
***
## What Are Threads?
A **thread** is a unit of execution within a program.\
In proxy usage, threads allow you to run multiple actions—such as web requests or scrapers—**simultaneously**, improving efficiency and speed.
### Analogy: Threads Are Like Checkout Counters
Imagine a supermarket with one cashier: everyone waits in line.\
Open 10 counters, and customers check out faster.
* **Single-threaded** → one cashier, one queue
* **Multi-threaded** → multiple cashiers, faster service
That’s exactly how threads improve proxy-based automation and scraping.
***
## How Threads Work in Proxies
Threads are used to send many requests at once, without waiting for each one to finish before starting the next.
### Without Threads (Single-Threaded)
* Only one request runs at a time.
* The next request waits for the previous one to finish.
* Very slow when dealing with large datasets.
### With Threads (Multi-Threaded)
* Many requests run at the same time.
* No waiting between requests.
* Data collection and automation happen much faster.
***
## Why Threads Matter in Proxy Usage
| Benefit | Explanation |
| ------------------------------- | ----------------------------------------------------------------- |
| **Faster Data Collection** | Executes multiple requests in parallel. |
| **Efficient Proxy Utilization** | Distributes requests across multiple proxies, reducing detection. |
| **Large-Scale Capability** | Ideal for scraping thousands of pages or bulk API requests. |
| **Lower Ban Risk** | Threads allow balanced load across IPs to minimize blocking. |
***
## Choosing the Right Number of Threads
The ideal number of threads depends on your **proxy type**, **hardware resources**, and **task intensity**.
| Scenario | Recommended Threads |
| ------------------------------------- | -------------------------------- |
| Small-scale scraping (few pages) | 5–10 threads |
| Medium-scale scraping (moderate data) | 20–50 threads |
| Large-scale scraping (massive data) | 100+ threads |
| Using residential proxies | Fewer threads (avoid bans) |
| Using datacenter proxies | More threads (faster processing) |
⚠️ **Too many threads** → may cause IP bans or overload your proxies.\
🐢 **Too few threads** → slows down operations.\
🎯 **Balance is key** — test and adjust based on performance.
***
## Threads vs. Concurrent Connections
Threads and concurrent connections are related but not identical.
| Feature | Threads | Concurrent Connections |
| -------------- | --------------------------------------- | ------------------------------------------------- |
| **Definition** | Execution units within a program | Multiple network connections happening at once |
| **Purpose** | Controls how many tasks run in parallel | Defines how many requests are sent simultaneously |
| **Example** | Running multiple scrapers in parallel | Opening 50 browser tabs at once |
***
## Summary
* Threads enable multitasking and speed up proxy workflows.
* Start with fewer threads and scale gradually.
* Match thread count to your proxy pool — more proxies can safely support more threads.
* Monitor CPU, RAM, and network limits to prevent system overload.
Understanding threading is crucial for stable, efficient, and scalable proxy performance.
# Access Credentials (/docs/proxies/getting-started/prerequisites/access-credentials)
***
To use Geonode's proxy services or API, you need authentication credentials. These credentials include a **username** and **password**, which allow you to securely connect to the service.
This guide will help you:\
✔️ Access your credentials from the Geonode dashboard.\
✔️ Secure them properly to prevent unauthorized access.\
✔️ Use them for authentication in API requests.
### Access Your Dashboard
* **Login to Geonode Dashboard:**
Visit [Geonode Dashboard](https://app.geonode.com/) and log in using your credentials.
* **Navigate to Proxy Section:**
Go to the **Proxies** section in the dashboard.
* **Get Your User Credentials:**
* **Username:** Copy your unique API username.
* **Password:** Copy your API password.
* **Secure Your Credentials:**
Keep your API credentials safe and never share them publicly
Once you've obtained your API credentials, you're ready to make your first API request or configure your proxy.
# Proxy Server Information (/docs/proxies/getting-started/prerequisites/proxy-server-information)
***
To connect to Geonode's proxy network, you need the **Proxy IP** and **Port**.
These details allow you to configure your applications, scripts, or browser to route traffic through Geonode's servers.
## Steps: Get Proxy Credentials from Geonode
1. Login to the Geonode Dashboard
Visit [Geonode Dashboard](https://app.geonode.com/) and sign in.
2. Go to the Proxy Section\
Navigate to the Proxies section to access your assigned proxy details.
3. Select the proxy format as: `hostname:port:username:password`
4. Copy the proxy details from the Geonode dashboard.
* Proxy IP: This is the server address you will use.
* Port: The port number required to establish a connection.
OR
Copy Your Proxy dns and Port
5. **Save These Details Securely**\
Do not share your proxy details publicly to avoid unauthorized usage.
## Next Steps
Once you have your proxy IP and port, you can:\
✔️ Use them in API requests.\
✔️ Configure them in your browser, terminal, or scripts.
# Verify Proxy Connection (/docs/proxies/getting-started/setup_and_configuration/verify-proxy-connection)
import ProxyOs from "../../../../snippets/proxy-os.mdx";
import BrowsersFaqs from "../../../../snippets/browsers-faqs.mdx";
import SupportParagraph from "../../../../snippets/support-paragraph.mdx";
You can verify your proxy connection in two ways:
* Using an online verification tool
* Using the command line (cURL)
***
## Method 1: Using an Online Tool
### Step 1: Check Your Current IP Address
Before enabling your proxy, check your current IP address.
1. Visit [IP API](https://ip-api.com/).
2. Note your IP address for reference.
### Step 2: Connect to Your Proxy
Set up and configure your proxy based on your device or browser.\
Follow the setup guides for your platform below:
### Step 3: Verify Your New IP Address
Once connected to the proxy:
1. Go back to [IP API](https://ip-api.com/).
2. Verify if the displayed IP address matches your proxy location.
3. If the IP has changed, your proxy connection is active.
***
## Method 2: Using the Command Line (cURL)
You can also verify the proxy connection through the command line using `curl`.
```bash
curl -x proxy.geonode.io:9000 http://ip-api.com
```
# Overview (/docs/proxies/guides/datacenter-proxies/overview)
Rotating Datacenter Proxies give you access to high-speed datacenter IPs for browsing, automation, and data collection. They are billed pay-as-you-go by traffic (per GB), and you configure targeting, protocol, and session type from the Geonode dashboard.
The dashboard experience is similar to [Residential Proxies](/docs/proxies/guides/residential-proxies/overview). The main difference is that traffic uses **datacenter** IPs instead of residential IPs.
## How Rotating Datacenter Proxies Work
Rotating Datacenter Proxies use a traffic-based plan model:
* Usage is billed per GB.
* Your remaining bandwidth is shown in GB in the dashboard.
* Traffic is consumed as you send requests through the proxy.
* Pricing starts from the rate shown on your Rotating Datacenter Proxies plan.
* You can configure endpoints, targeting, and session type before connecting.
* You can top up bandwidth manually or enable auto top-up from the dashboard.
See [Pricing](/docs/proxies/guides/datacenter-proxies/pricing) for Pay As You Go and Subscription plans.
## Access Your Credentials
1. Log in to the [Geonode Dashboard](https://app.geonode.com/).
2. Go to **Proxies → Rotating Datacenter Proxies**.
3. Open **Proxy configuration**.
4. Find your proxy credentials.
Your credentials include:
* **Username**
* **Password**
* **Host**
* **Port**
## Configure Your Proxy
You can configure Rotating Datacenter Proxies from the Geonode dashboard under **Proxies → Rotating Datacenter Proxies → Proxy configuration**.
From Proxy Configuration, you can set:
* Gateway
* Country, state, and city targeting
* ASN/ISP targeting
* OS targeting
* Protocol
* Session type
* Endpoint count and format
* API credentials for username and password authentication
The dashboard also includes **Port configuration**, **Active sticky sessions**, **Statistics**, **Block list**, and **Reseller**.
See [Block List](/docs/proxies/guides/residential-proxies/block-list) to block outlier domains, IPs, or wildcards so the proxy never fetches them.
### Configuration Options
| Option | Description |
| --------------------- | -------------------------------------------------------- |
| **Gateway** | Select the gateway region for your proxy traffic |
| **Country targeting** | **Any** by default, or a specific country |
| **State targeting** | **Any** by default, or a specific state or region |
| **City targeting** | **Any** by default, or a specific city |
| **ASN/ISP targeting** | **Any** by default, or a specific ASN/ISP when available |
| **OS targeting** | Filter by operating system when available |
| **Protocol** | HTTP/HTTPS or SOCKS5 |
| **Session type** | Rotating or Sticky |
| **Endpoints** | Number of endpoints and output format |
### Bandwidth
The top of the Rotating Datacenter Proxies page shows your available bandwidth in GB.
From there you can:
* Check remaining bandwidth
* Use **Manual top-up** to add traffic
* Enable **auto top-up**
* Open **See pricing** to review plan rates
For full Pay As You Go and Subscription pricing, see [Pricing](/docs/proxies/guides/datacenter-proxies/pricing).
### Endpoint Format
Use this format for tools that accept a single proxy string:
```text
hostname:port:username:password
```
Example:
```text
proxy.geonode.io:9000:USERNAME:PASSWORD
```
Replace `USERNAME` and `PASSWORD` with your dashboard credentials.
### Port Ranges
For rotating sessions, Geonode uses these common port ranges:
| Protocol | Port range |
| -------------- | ------------- |
| **HTTP/HTTPS** | `9000–9010` |
| **SOCKS5** | `11000–11010` |
The dashboard shows the active port range for your selected protocol under the generated endpoint list.
### Basic Connection
#### HTTP
```bash
curl -x http://USERNAME:PASSWORD@proxy.geonode.io:9000 https://ipinfo.io
```
#### SOCKS5
```bash
curl --proxy socks5://USERNAME:PASSWORD@proxy.geonode.io:11000 https://ipinfo.io
```
## Targeting
Rotating Datacenter Proxies support the same style of geo and network targeting as Residential Proxies:
* **Country targeting** — route traffic through a specific country
* **State targeting** — narrow traffic to a state or region
* **City targeting** — target a specific city
* **ASN/ISP targeting** — route through a specific ISP or ASN when available
* **OS targeting** — filter by operating system when available
For detailed targeting workflows, see:
* [Target Specific Location](/docs/proxies/guides/residential-proxies/target-specific-location)
* [Exclude Specific Location](/docs/proxies/guides/residential-proxies/exclude-specific-location)
* [Switch Between Different Proxy Locations](/docs/proxies/guides/residential-proxies/switch-between-different-proxy-locations)
## Sessions
Rotating Datacenter Proxies support both rotating and sticky sessions:
* **Rotating** — a new IP is used across requests, typically over rotating ports such as HTTP `9000–9010`
* **Sticky** — keep the same IP for a session when you need continuity
For session workflows, see:
* [Rotating Proxies](/docs/proxies/guides/residential-proxies/rotating-proxies)
* [New Sticky Session](/docs/proxies/guides/residential-proxies/new-sticky-session)
* [How to Release Sticky Sessions](/docs/proxies/guides/residential-proxies/release-a-sticky-session)
* [How to Set Proxy Session Lifetime](/docs/proxies/guides/residential-proxies/identifying-session-lifetime)
## Endpoints and Credentials
From **Proxy configuration**, you can:
1. Copy your API username and password.
2. Choose an endpoint format such as `hostname:port:username:password`.
3. Generate one or more proxy endpoints.
4. Use the generated endpoints in your tools or scripts.
For code samples generated from your configuration, see [API Code Generator](/docs/proxies/guides/residential-proxies/api-code-generator).
## Monitoring
Use the dashboard **Statistics** view and usage guides to monitor traffic and performance.
See [Usage Stats & Analytics](/docs/proxies/guides/residential-proxies/usage-stats-analytics) for monitoring proxy usage.
See [Latency Testing](/docs/proxies/guides/residential-proxies/success-latency) when you need to test response performance.
## Getting Started
To start using Rotating Datacenter Proxies:
1. Open the Geonode dashboard.
2. Go to **Rotating Datacenter Proxies**.
3. Open **Proxy configuration**.
4. Set your targeting, protocol, and session type.
5. Copy your credentials and generated endpoints.
If you are new to Geonode proxies, start with the [Quick Start Guide](/docs/proxies/getting-started/quick-start).
# Pricing (/docs/proxies/guides/datacenter-proxies/pricing)
Rotating Datacenter Proxies are billed by traffic (per GB). In the dashboard, you can choose between **Pay As You Go** and **Subscription**.
Subscription plans save about **10%** compared with Pay As You Go rates for the same bandwidth tier.
## Billing Options
| Option | Description |
| ----------------- | ------------------------------------------------- |
| **Pay As You Go** | Buy a fixed amount of bandwidth when you need it |
| **Subscription** | Monthly bandwidth plans with a lower price per GB |
Across both options:
* Bandwidth is measured in GB.
* Unused bandwidth rolls over until cancellation.
* Crypto payment is available where shown in the dashboard.
* Enterprise pricing is available through sales.
## Pay As You Go
Pay As You Go lets you purchase bandwidth without a monthly subscription commitment.
| Plan | Price per GB | Total price | Bandwidth |
| -------------- | -----------: | ----------: | --------: |
| **10 GB Plan** | $0.53 | $5.30 | 10 GB |
| **25 GB Plan** | $0.49 | $12.25 | 25 GB |
| **50 GB Plan** | $0.47 | $23.50 | 50 GB |
### Flexible Plan
The **Flexible Plan** lets you choose a custom amount.
| Detail | Value |
| ------------------- | ----------------------------- |
| **Price per GB** | From $0.53 to $0.16 |
| **Minimum order** | 10 GB |
| **Monthly minimum** | None |
| **Bandwidth** | Rolls over until cancellation |
The more bandwidth you buy, the lower the price per GB.
## Subscription
Subscription plans bill monthly and include a set amount of bandwidth each month. Unused bandwidth rolls over until cancellation.
| Plan | Price per GB | Monthly price | Bandwidth included monthly |
| --------------- | -----------: | ------------: | -------------------------: |
| **10 GB Plan** | $0.477 | $4.77 | 10 GB |
| **25 GB Plan** | $0.441 | $11.03 | 25 GB |
| **50 GB Plan** | $0.423 | $21.15 | 50 GB |
| **100 GB Plan** | $0.396 | $39.60 | 100 GB |
| **250 GB Plan** | $0.360 | $90.00 | 250 GB |
| **500 GB Plan** | $0.333 | $166.50 | 500 GB |
| **1 TB Plan** | $0.306 | $306.00 | 1,000 GB |
| **2 TB Plan** | $0.288 | $576.00 | 2,000 GB |
| **3 TB Plan** | $0.270 | $810.00 | 3,000 GB |
| **5 TB Plan** | $0.252 | $1,260.00 | 5,000 GB |
| **7.5 TB Plan** | $0.234 | $1,755.00 | 7,500 GB |
| **10 TB Plan** | $0.225 | $2,250.00 | 10,000 GB |
| **15 TB Plan** | $0.216 | $3,240.00 | 15,000 GB |
| **20 TB Plan** | $0.198 | $3,960.00 | 20,000 GB |
| **25 TB Plan** | $0.198 | $4,950.00 | 25,000 GB |
| **30 TB Plan** | $0.189 | $5,670.00 | 30,000 GB |
| **40 TB Plan** | $0.180 | $7,200.00 | 40,000 GB |
| **50 TB Plan** | $0.171 | $8,550.00 | 50,000 GB |
| **60 TB Plan** | $0.162 | $9,720.00 | 60,000 GB |
| **75 TB Plan** | $0.153 | $11,475.00 | 75,000 GB |
| **100 TB Plan** | $0.144 | $14,400.00 | 100,000 GB |
Higher-volume plans have a lower price per GB.
## Choosing a Plan
| If you need | Consider |
| ----------------------------- | ------------------------- |
| Occasional or irregular usage | Pay As You Go |
| A custom bandwidth amount | Flexible Plan |
| Predictable monthly usage | Subscription |
| Lower price per GB at scale | Larger Subscription tiers |
## How Bandwidth Works
* Your remaining bandwidth is shown in GB on the Rotating Datacenter Proxies dashboard.
* Traffic is consumed as you send requests through the proxy.
* You can top up bandwidth manually or enable auto top-up.
* Unused bandwidth rolls over until the plan is cancelled.
## Need Custom Terms?
If you need enterprise-grade solutions or a unique offer, [contact sales](https://geonode.com/contact).
# Overview (/docs/proxies/guides/isp-proxies/isp-proxies)
ISP Proxies provide proxy IPs associated with Internet Service Providers (ISPs). The Geonode dashboard lets you manage your assigned ISP proxy IPs, view their network details, organize them with tags and notes, and monitor usage.
If you are new to proxies, see the [Proxy Service Guide](/docs/proxies/getting-started/knowledge-base/overview) to learn the basic proxy concepts before continuing.
## ISP Proxies Dashboard
Open **ISP Proxies** from the **Proxies** section of the Geonode dashboard.
The dashboard provides two main areas:
* **Proxy List** — View and manage your assigned ISP proxy IPs.
* **Statistics** — Monitor proxy usage and performance.
The top section also shows information about your current ISP Proxies subscription and the number of assigned IPs.
## Subscription Overview
The subscription section shows what your current ISP Proxies package includes.
Depending on your account, you can see:
* Total number of assigned IPs
* Number of **Shared** proxies
* Number of **Dedicated** proxies
* Current subscription status
* Subscription end or renewal information
* Options to view plans, buy more IPs, or manage the subscription
The number of IPs in the **Proxy List** matches your package. For example, if your plan includes **4 IPs** with **2 Shared** and **2 Dedicated** proxies, the dashboard shows **4 IPs** at the top and lists those four proxies below.
Your package can include Shared IPs only, Dedicated IPs only, or a mix of both. The Shared and Dedicated counts in the summary explain why each IP appears in your list.
Based on your subscription, IPs may be removed when the plan ends on a specific date. If your subscription allows it, you can re-activate the plan to restore access instead of losing the IPs permanently.
For plan types and pricing, see [ISP Proxy Pricing](/docs/proxies/guides/isp-proxies/pricing).
## Proxy List
The **Proxy List** contains the ISP proxy IPs assigned to your account based on your current package.
The proxy list provides information about each IP, including:
| Field | Description |
| ---------------- | ------------------------------------------------------- |
| **IP** | IP address of the proxy. |
| **Port** | Port used to connect to the proxy. |
| **Username** | Username associated with the proxy. |
| **Password** | Password associated with the proxy. |
| **Country** | Country associated with the proxy IP. |
| **City** | City associated with the proxy IP. |
| **State** | State or region associated with the proxy IP. |
| **ISP** | Internet Service Provider associated with the proxy IP. |
| **Last checked** | Most recent time the proxy information was checked. |
| **Created at** | Date when the proxy was created. |
| **Type** | Whether the proxy is Shared or Dedicated. |
| **Status** | Current status of the proxy. |
| **Tags** | Tags assigned to the proxy. |
## Search Your Proxies
Use the **Search** field above the proxy list to find a specific proxy.
Search can help you locate an IP in a larger proxy list without manually checking each row.
## View IP Details
Select a proxy from the list to open its **IP details**.
The IP details view shows the full information for the selected proxy, including:
* IP address
* Port
* Country
* State
* City
* ISP
* Status (for example, Active)
* **Type** — Shared or Dedicated
* Tags
* Notes
Use **Type** to confirm whether the selected IP is **Shared** or **Dedicated**. This matches the Shared and Dedicated counts shown in your subscription summary.
For example, in a package with 2 Shared and 2 Dedicated proxies, opening any IP in the list shows whether that specific IP is Shared or Dedicated.
You can also add notes to the proxy.
### Add Notes
Use the notes field to store information about the selected IP.
For example, you can use notes to keep track of how you use a particular proxy within your own workflow.
After adding or updating a note, select **Save Changes**.
### Delete or Re-activate an IP
The IP details window provides a **Delete IP** option when you want to remove a selected proxy from your account manually.
Separately, your package controls how long IPs stay available:
* When a subscription ends on a specific date, IPs from that plan may be removed automatically.
* If re-activation is available for your plan, you can re-activate the subscription to restore access to your ISP proxies.
Deleting an IP removes it from your proxy list. Make sure you no longer need the proxy before deleting it. Package end dates can also remove IPs unless you renew or re-activate the subscription.
## Organize Proxies with Tags
You can assign tags to your ISP proxy IPs from the proxy list.
Select **add tag** for a proxy to open the Tags dialog.
You can:
* Choose a tag color.
* Enter a tag name.
* Save the tag.
Tags can help you organize proxies according to your own workflow.
For example, you could create tags based on an internal project, application, or other grouping that is useful to you.
## Export Proxy List
The proxy list includes an **Export CSV** option.
Use **Export CSV** to export the proxy list as a CSV file for use outside the Geonode dashboard.
This can be useful when you need to work with your proxy information in another application or maintain your own records.
The proxy list contains sensitive connection information, including usernames and passwords. Store exported files securely and do not share them publicly.
## Statistics
Open the **Statistics** tab to view your ISP Proxy usage.
The Statistics section allows you to select a time period and review your proxy activity.
Available time ranges include:
* **Last 24 hours**
* **7 days**
* **30 days**
* **90 days**
You can also select a custom date range.
## Usage Metrics
The statistics dashboard provides an overview of your proxy activity for the selected period.
The dashboard displays metrics such as:
* **Total** — Total number of requests.
* **Success Rate** — Percentage of successful requests.
* **Average Duration** — Average request duration.
* **Requests Used** — Number of requests used during the selected period.
The dashboard also provides charts for request activity and request usage.
## Using Your ISP Proxies
Once you have your ISP proxy IP, port, username, and password, you can configure the proxy in your application or tool.
For information about proxy protocols and connection methods, see:
* [Protocols](/docs/proxies/getting-started/knowledge-base/protocols)
* [Proxy Endpoint Formats](/docs/proxies/getting-started/knowledge-base/endpoint-formats)
* [Port Usage](/docs/proxies/getting-started/knowledge-base/proxy-usage)
If you need to target a specific ISP or ASN, see [ASN/ISP Targeting](/docs/proxies/getting-started/knowledge-base/asn-isp-targeting).
## What's Next?
You now know how to manage ISP Proxies from the Geonode dashboard.
You can:
* Understand how many Shared and Dedicated IPs your package includes.
* View your assigned IPs in the proxy list.
* Open IP details to confirm whether an IP is Shared or Dedicated.
* Add tags and notes.
* Delete IPs or re-activate a subscription when needed.
* Export your proxy list.
* Monitor proxy usage and performance.
For plan options, see [ISP Proxy Pricing](/docs/proxies/guides/isp-proxies/pricing).
For more information about using ISP and ASN targeting, see [ASN/ISP Targeting](/docs/proxies/getting-started/knowledge-base/asn-isp-targeting).
# ISP Proxy Pricing (/docs/proxies/guides/isp-proxies/pricing)
ISP Proxies are available in **Shared** and **Dedicated** options. The price per IP depends on the selected proxy type and the number of IPs in your plan.
## Choose Your Proxy Type
When creating or changing an ISP Proxy subscription, you can choose between two proxy types.
| Proxy Type | Description |
| ------------------------- | ----------------------------------- |
| **Shared ISP Proxies** | Shared by up to 3 users. |
| **Dedicated ISP Proxies** | Exclusive access, used only by you. |
The selected proxy type determines which pricing tiers apply to your subscription.
## Shared ISP Proxies
Shared ISP Proxies are shared by up to 3 users.
The price per IP decreases as you increase the number of IPs in your plan.
### Shared ISP Pricing Tiers
| Number of IPs | Price per IP |
| ------------- | -----------: |
| 3–4 | $1.75 |
| 5–19 | $1.70 |
| 20–29 | $1.65 |
| 30–39 | $1.60 |
| 40–49 | $1.55 |
| 50–74 | $1.50 |
| 75–99 | $1.45 |
| 100–149 | $1.40 |
| 150–199 | $1.35 |
| 200–999 | $1.30 |
| 1000–10000 | $1.25 |
### Example
If you select **15 Shared ISP IPs**, the dashboard shows the **5–19** pricing tier.
The price is:
```text
15 × $1.70 = $25.50/month
```
The dashboard may also show the price for your current subscription alongside the price of the new subscription when you change the number of IPs.
## Dedicated ISP Proxies
Dedicated ISP Proxies provide exclusive access to you.
Dedicated ISP Proxies use a separate pricing structure from Shared ISP Proxies.
### Dedicated ISP Pricing Tiers
| Number of IPs | Price per IP |
| ------------- | -----------: |
| 3–4 | $3.50 |
| 5–29 | $3.00 |
| 30–74 | $2.85 |
| 75–299 | $2.75 |
| 300–999 | $2.50 |
| 1000–10000 | $2.25 |
### Example
If you select **15 Dedicated ISP IPs**, the dashboard shows the **5–29** pricing tier.
The price is:
```text
15 × $3.00 = $45.00/month
```
## How Pricing Is Calculated
The monthly price is based on:
1. The selected proxy type.
2. The number of IPs.
3. The price per IP for the applicable pricing tier.
For example, selecting 15 IPs results in:
| Proxy Type | IPs | Price per IP | Monthly Price |
| ---------- | --: | -----------: | ------------: |
| Shared | 15 | $1.70 | $25.50 |
| Dedicated | 15 | $3.00 | $45.00 |
## Configure Your Plan
When configuring an ISP Proxy subscription:
1. Select **Shared ISP Proxies** or **Dedicated ISP Proxies**.
2. Set the **Number of IPs** using the number field or the `+` and `−` controls.
3. Review the price per IP shown by the dashboard.
4. Review the monthly subscription amount.
5. Select **Proceed to checkout** to continue.
The dashboard shows the current subscription and the new subscription separately when you are changing an existing plan.
## Increasing Your IP Count
Increasing the number of IPs can move your subscription into a different pricing tier.
For example, Shared ISP pricing changes as follows:
```text
3–4 IPs → $1.75/IP
5–19 IPs → $1.70/IP
20–29 IPs → $1.65/IP
30–39 IPs → $1.60/IP
...
1000–10000 IPs → $1.25/IP
```
The same tier-based pricing applies to Dedicated ISP Proxies using their dedicated pricing table.
## Pricing Tiers
The pricing tables show the price **per IP** for each IP range.
This means adding more IPs can move the subscription into a lower per-IP pricing tier.
For example:
```text
Shared ISP
15 IPs
↓
5–19 tier
↓
$1.70 per IP
↓
$25.50/month
```
And:
```text
Dedicated ISP
15 IPs
↓
5–29 tier
↓
$3.00 per IP
↓
$45.00/month
```
## Current Subscription vs New Subscription
When modifying an existing subscription, the configuration screen can show both:
* **Current subscription**
* **New subscription**
This allows you to compare the existing IP allocation and price with the configuration you are about to purchase.
For example, the dashboard can show an existing subscription with 2 IPs and a new configuration with 17 IPs after adding 15 more IPs.
Review the **New subscription** section before continuing to checkout.
## Choosing Between Shared and Dedicated
Choose **Shared ISP Proxies** when you need ISP proxy IPs that can be shared by up to 3 users.
Choose **Dedicated ISP Proxies** when you need exclusive access to the assigned proxy IPs.
The two options have different prices, so select the proxy type based on your access requirements.
## What's Next?
After selecting your ISP Proxy type and configuring the number of IPs, continue to checkout to create or update your subscription.
Once your proxies are available, you can manage them from the [ISP Proxies Dashboard](/docs/proxies/guides/isp-proxies/isp-proxies).
# Overview (/docs/proxies/guides/unlimited-residential-proxies/00_unlimited_residential_proxies)
Unlimited Residential Proxies are speed-based plans designed for high-usage customers. They are billed as a monthly subscription, not per GB, and provide unlimited traffic within the limits of your selected plan.
## How Unlimited Residential Proxies Work
Unlimited Residential Proxies use a speed-based plan model:
* Plans are based on a speed limit such as 200 Mbps, 400 Mbps, or 1000 Mbps.
* Traffic is unlimited.
* Pricing is charged monthly.
* Usage is not billed per GB.
* Performance depends on your setup and the speed cap of your plan.
Unlimited Residential currently supports the **US gateway only**.
There are no thread or concurrency limits enforced by the plan. Performance depends on your setup and the speed limit of your selected plan.
## Configure Your Proxy
You can configure your Unlimited Residential Proxy from the Geonode dashboard.
Available configuration options include:
* Gateway
* Country targeting
* OS targeting
* Protocol
* Session type and session time
* Endpoints
See [Configure Proxy Settings](/docs/proxies/guides/unlimited-residential-proxies/03_configure_proxy_settings) for setup and connection details.
## Plans
Unlimited Residential plans are available at different speed levels. Each plan provides unlimited traffic with a different monthly price and speed limit.
See [Pricing](/docs/proxies/guides/unlimited-residential-proxies/01_unlimited_residential_proxy_pricing) for the available plans and pricing.
## Monitoring
To monitor requests, success rate, and data usage, see [Usage Stats & Analytics](/docs/proxies/guides/residential-proxies/usage-stats-analytics).
To stop the proxy from fetching outlier domains you define, see [Block List](/docs/proxies/guides/residential-proxies/block-list).
## Getting Started
To start using Unlimited Residential Proxies:
1. Open the Geonode dashboard.
2. Go to **Unlimited Residential Proxies**.
3. Open **Proxy Configuration**.
4. Configure your proxy settings and access your credentials.
See [Configure Proxy Settings](/docs/proxies/guides/unlimited-residential-proxies/03_configure_proxy_settings) for the setup steps.
# Pricing (/docs/proxies/guides/unlimited-residential-proxies/01_unlimited_residential_proxy_pricing)
## Available Plans
Unlimited Residential Proxies use a speed-based monthly pricing model. Traffic is unlimited, and the monthly price depends on the selected speed.
| Speed | Monthly Price | Traffic |
| --------: | ------------: | --------- |
| 200 Mbps | $1,800 | Unlimited |
| 400 Mbps | $2,400 | Unlimited |
| 600 Mbps | $3,000 | Unlimited |
| 800 Mbps | $3,400 | Unlimited |
| 1000 Mbps | $3,800 | Unlimited |
Higher-speed plans provide a higher speed limit, while all listed plans provide unlimited traffic.
Unlimited Residential plans are not billed per GB. Your monthly price is based on the selected speed tier.
## Choosing a Plan
Choose a plan based on the throughput your workload requires.
| If you need | Consider |
| --------------- | -------------- |
| Up to 200 Mbps | 200 Mbps plan |
| Up to 400 Mbps | 400 Mbps plan |
| Up to 600 Mbps | 600 Mbps plan |
| Up to 800 Mbps | 800 Mbps plan |
| Up to 1000 Mbps | 1000 Mbps plan |
Unlimited refers to traffic volume. Your selected plan still has a defined speed limit.
## Monthly Subscription
Unlimited Residential plans are billed as a monthly subscription.
The plans are not charged based on the amount of data used. Instead, the monthly price is determined by the selected speed tier.
## Need Custom Terms?
If you need a different duration or custom terms, [contact sales](https://geonode.com/contact) for an enterprise-grade solution.
# Configure Proxy Settings (/docs/proxies/guides/unlimited-residential-proxies/03_configure_proxy_settings)
To get started with Unlimited Residential Proxies, access your proxy credentials from the Geonode dashboard.
## Access Your Credentials
1. Log in to the [Geonode Dashboard](https://app.geonode.com/).
2. Go to **Residential Proxies**.
3. Open **Proxy Configuration**.
4. Find your proxy credentials.
Your credentials include:
* **Username**
* **Password**
* **Host**
* **Port**
## Proxy Configuration
You can configure Unlimited Residential Proxies from the Geonode dashboard under **Proxies → Unlimited Residential Proxies → Proxy configuration**.
From Proxy Configuration, you can set:
* Gateway
* Country targeting
* OS targeting
* Protocol
* Session type and session time
* Endpoint count and format
### Configuration Options
| Option | Description |
| --------------------- | ---------------------------------------------------------- |
| **Gateway** | Currently **United States** only for Unlimited Residential |
| **Country targeting** | **Any** by default, or a specific country |
| **OS targeting** | Filter by operating system when available |
| **Protocol** | HTTP/HTTPS or SOCKS5 |
| **Session type** | Rotating or Sticky |
| **Session time** | How long a sticky session keeps the same IP |
| **Endpoints** | Number of endpoints and output format |
Unlimited Residential currently supports the **US gateway only**.
### Endpoint Format
Use this format for tools that accept a single proxy string:
```text
hostname:port:username:password
```
Example:
```text
residential-unlimited-us-01-proxy.geonode.io:13003:USERNAME:PASSWORD
```
Replace `USERNAME` and `PASSWORD` with your dashboard credentials.
### Basic Connection
#### HTTP
```bash
curl -x http://USERNAME:PASSWORD@residential-unlimited-us-01-proxy.geonode.io:13003 https://ipinfo.io
```
#### SOCKS5
```bash
curl --proxy socks5://USERNAME:PASSWORD@residential-unlimited-us-01-proxy.geonode.io:13003 https://ipinfo.io
```
## Related Guides
For detailed workflows, use the Residential Proxies guides:
* [Target Specific Location](/docs/proxies/guides/residential-proxies/target-specific-location) — country and geo targeting
* [Rotating Proxies](/docs/proxies/guides/residential-proxies/rotating-proxies) — rotating session setup
* [New Sticky Session](/docs/proxies/guides/residential-proxies/new-sticky-session) — create sticky sessions
* [How to Set Proxy Session Lifetime](/docs/proxies/guides/residential-proxies/identifying-session-lifetime) — session time
* [How to Release Sticky Sessions](/docs/proxies/guides/residential-proxies/release-a-sticky-session) — release sticky sessions
* [Usage Stats & Analytics](/docs/proxies/guides/residential-proxies/usage-stats-analytics) — monitor usage and statistics
* [Block List](/docs/proxies/guides/residential-proxies/block-list) — block outlier domains, IPs, or wildcards
# Troubleshooting Unlimited Residential Proxies (/docs/proxies/guides/unlimited-residential-proxies/06_troubleshooting)
In this guide, you’ll learn how to troubleshoot common issues with Unlimited Residential Proxies and what to try when your results are slow, blocked, or failing.
## Slow Results
If your results are slow, try:
* Reduce local concurrency.
* Release sessions to refresh IPs.
* Check your local ISP or server limits.
## Target Blocks You
If the target website blocks you, try:
* Release or rotate the session.
* Adjust country targeting.
* Use sticky sessions for login or checkout flows.
## High Failure Rate
If you experience a high failure rate, try:
* Release sessions.
* Try different country targeting.
* Use fewer concurrent connections.
Review your proxy configuration and statistics to identify whether the issue is related to sessions, country targeting, concurrency, or your local network setup.
# API Code Generator (/docs/proxies/guides/residential-proxies/api-code-generator)
The API Code Generator allows you to instantly create code snippets for different programming languages based on your configured proxy parameters.\
This helps you test, integrate, and automate API calls quickly and efficiently.
***
## Step 1: Configure Your Proxy
Before generating code, set up your proxy endpoint.
Follow the guide:\
➡️ [How to Use the Endpoint Generator to Configure a Proxy](/docs/proxies/getting-started/setup_and_configuration/advance-configuration/proxy-configuration)
This ensures you have the correct proxy details ready for code generation.
***
## Step 2: Access the API Code Generator
Once your endpoint is created:
1. Scroll down to the **API Code Generator** section in your dashboard.
2. You will see code automatically generated based on your proxy configuration.
3. The code is available in multiple programming languages, including:
* Python
* Node.js
* Go
* And others
***
## Step 3: Copy and Paste the Code
1. Choose your preferred language.
2. Copy the generated code snippet.
3. Paste it into your development environment or editor (e.g., VS Code, PyCharm, GoLand).
***
## Step 4: Run the Code
Example: Running the Python code in **VS Code**.
If your snippet uses Python’s `requests` package, install it first:
```bash
pip install requests
```
Then run your file:
`python app.py`
***
## Step 5: Verify the Output
After running the script, check the console or terminal.
The output should display the expected data from your API call, confirming that your proxy and code are properly configured.
You’re now ready to generate, customize, and run API code seamlessly with Geonode.
# Block List (/docs/proxies/guides/residential-proxies/block-list)
The **Block list** lets you prevent your Geonode proxies from accessing specific domains, IPs, or wildcards.
## When to use the Block list
Use the Block list when you want to prevent your proxy from accessing specific destinations, such as unwanted or background requests.
For example, if your browser sends requests to `mozilla.dev` that you do not need, you can add it to the Block list.
## How it works
Once a target is added to the Block list, requests to that target return **468 Target Forbidden**. Other destinations continue to work normally.
## Block list vs. Location Exclude
These features control different things:
| Feature | Controls | Example |
| -------------------- | --------------------------------------- | ------------------- |
| **Block list** | Which destinations the proxy can access | Block `mozilla.dev` |
| **Location exclude** | Which proxy locations can be used | Exclude Germany |
Use the **Block list** to control **where your proxy can connect** and [Location Exclude](/docs/proxies/guides/residential-proxies/exclude-specific-location) to control **which proxy locations you receive**.
Image paths below are placeholders. Drop the dashboard captures into
`/images/functionalities/how-to/block-list/` using the filenames in each step.
***
## Step 1: Open the Block list
1. Log in to the [Geonode Dashboard](https://app.geonode.com/).
2. Go to **Proxies → Residential Proxies**.
3. Open the **Block list** tab.
Direct link: [app.geonode.com/proxies?tab=block-list](https://app.geonode.com/proxies?tab=block-list)
The same tab is also available from **Rotating Datacenter Proxies** (and other products that use this proxy dashboard).
If the list is empty, you will see **No data yet**.
***
## Step 2: What you can add
The input accepts:
| Type | Example | Use when |
| -------- | --------------- | ------------------------------- |
| Domain | `mozilla.dev` | Block that host |
| IP | `192.168.0.1` | Block a specific address |
| Wildcard | `*.domain2.com` | Block a host and its subdomains |
The placeholder in the dashboard shows the same formats: `domain1.com`, `192.168.0.1`, `*.domain2.com`.
Do **not** add the site you actually need (for example your scrape target or `ip-api.com` if you use it to test). Only add outlier hosts you want the proxy to drop.
***
## Step 3: Add a single entry
1. Type a domain, IP, or wildcard in the field (for example `mozilla.dev`).
2. Click **Add entries**.
3. Confirm the row appears in the table.
Wait a few seconds, then send a request (see [Verify the block](#verify-the-block)).
{/*  */}
***
## Step 4: Bulk add
Use **Bulk add** when you have many hosts at once (trackers, helper domains, a list from a HAR file).
1. Click **Bulk add**.
2. Paste the domains, IPs, or wildcards.
3. Confirm. The new rows appear in the table.
You can also type several values in the main field (comma-separated, matching the placeholder examples) and click **Add entries**.
***
## Step 5: Search, delete, and export CSV
### Search
Use **Search entries** to filter a long list by domain, IP, or wildcard.
### Delete
* Select one or more checkboxes, then click **Delete selected**.
* Or use the trash icon on a single row.
After you delete an entry, wait a few seconds. That host is allowed again.
### Export CSV
Click **Export CSV** to download the current list. Use it as a backup, to review entries offline, or to share the list with your team.
***
## Verify the block
After `mozilla.dev` is on the list, send two requests through the same proxy: one allowed host, one blocked host.
Replace `USERNAME` and `PASSWORD` with the values from **Proxy configuration**. For rotating residential HTTP, use port `9000` and a username that includes `-type-residential`.
### cURL
```bash
curl -x "http://USERNAME:PASSWORD@proxy.geonode.io:9000" -I "http://ip-api.com/json"
curl -x "http://USERNAME:PASSWORD@proxy.geonode.io:9000" -I "http://mozilla.dev"
```
### Python
```python
import requests
username = "geonode_demouser-type-residential"
password = "demopass"
proxy = {
"http": f"http://{username}:{password}@proxy.geonode.io:9000",
"https": f"http://{username}:{password}@proxy.geonode.io:9000",
}
allowed = requests.get("http://ip-api.com/json", proxies=proxy, timeout=30)
print(allowed.status_code, allowed.text[:200])
blocked = requests.get("http://mozilla.dev", proxies=proxy, timeout=30)
print(blocked.status_code, blocked.reason)
```
### What you should see
| Target | Result |
| ------------------------ | ------------------------------------------ |
| `http://ip-api.com/json` | **200** — still allowed |
| `http://example.com` | **200** — still allowed |
| `http://mozilla.dev` | **468 Target Forbidden** |
| `https://mozilla.dev` | Tunnel fails with **468 Target Forbidden** |
The proxy refuses the blocked host immediately. It does not fetch the page.
Platform security policies can return **464**. A host on *your* Block list
returns **468 Target Forbidden**. If `mozilla.dev` still loads, wait 10–30
seconds and retry, or confirm the row is still in the table.
Delete `mozilla.dev` from the list, wait, and run the same request again. It should succeed like any other allowed host.
***
## Tips
* Block only outliers, not the destination you are trying to reach.
* Browser sessions benefit the most: one page load can trigger many extra hosts.
* After add or delete, wait a moment before testing.
* Export CSV before a large cleanup so you can restore the list.
* Failed Block list hits can show up as failed requests in [Usage Stats & Analytics](/docs/proxies/guides/residential-proxies/usage-stats-analytics).
For other proxy error codes, see [Error Handling](/docs/proxies/api-reference/error-handling).
# Exclude Specific Location (/docs/proxies/guides/residential-proxies/exclude-specific-location)
import ProxyInfoFromGeonode from "../../../../snippets/get-proxy-info-from-geonode.mdx";
import RegionSpecificFAQs from "../../../../snippets/region-specific-faqs.mdx";
import SupportParagraph from "../../../../snippets/support-paragraph.mdx";
***
This guide will help you understand how to exclude proxies from specific cities, countries, states, or ISPs using the Geonode API.
***
## Prerequisites
Before you begin, make sure you:
* Have active Geonode proxy credentials.
* Understand how to make API calls using tools like cURL or Python.
***
## What is Location Exclusion?
Geonode allows you to exclude certain locations while routing traffic through proxies. This is helpful when you want to avoid specific regions due to content restrictions, compliance, or testing needs.
➡️ Want to route traffic through a location instead? [See Targeting Locations](/docs/proxies/guides/residential-proxies/target-specific-location)
***
## What Can You Exclude?
You can exclude proxies based on:
* Country
* State
* City
* ASN (Autonomous System Number)
You can't combine different exclusion types in a single request. For example,
you can't exclude cities and ASNs together.
***
## Format for Exclusion
To exclude a location, modify your proxy username like this:
```bash
curl --request GET \
-x "http://proxy.geonode.io:" \
--user "-not.country-:" \
--url "http://ip-api.com/json"
```
Supported location types:
* `not.country`
* `not.city`
* `not.state`
* `not.asn`
You can pass:
* A single exclusion: `-not.city-tokyo`
* Multiple exclusions: `-not.city-tokyo,kyoto,osaka`
***
## Get Proxy with Exclusions from Dashboard
You can easily get the correct country codes, city names, state names, and ASN numbers directly from the Geonode Dashboard.
Just go to the right-hand filters for Country, City, State, or ASN targeting. Once selected, they will appear in the proxy string for reference.
***
## Calling API with Location Exclusions
### 1. Exclude by Country
Use `-not.country-xx` or multiple like `-not.country-xx,yy,zz`.
```bash
curl --request GET \
-x "http://proxy.geonode.io:" \
--user "-not.country-,:" \
--url "http://ip-api.com/json"
```
📄 [See full API doc for country exclusion](/docs/proxies/api-reference/exclude-targeting/exclude-country)
***
### 2. Exclude by State
Use `-not.state-`.
```bash
curl --request GET \
-x "http://proxy.geonode.io:" \
--user "-country--not.state-:" \
--url "http://ip-api.com/json"
```
📄 [See full API doc for state exclusion](/docs/proxies/api-reference/exclude-targeting/exclude-state)
***
### 3. Exclude by City
Use `-not.city-`.
```bash
curl --request GET \
-x "http://proxy.geonode.io:" \
--user "-country--not.city-:" \
--url "http://ip-api.com/json"
```
📄 [See full API doc for city exclusion](/docs/proxies/api-reference/exclude-targeting/exclude-city)
***
### 4. Exclude by ASN (ISP)
Use `-not.asn-`. You can also pass multiple values like `-not.asn-31898,12271`.
```bash
curl --request GET \
-x "http://proxy.geonode.io:" \
--user "-country--not.asn-:" \
--url "http://ip-api.com/json"
```
📄 [See full API doc for ASN exclusion](/docs/proxies/api-reference/exclude-targeting/exclude-asn)
***
## Example Response
Here's an example API response when city or ASN is excluded:
```json
{
"status": "success",
"country": "United States",
"countryCode": "US",
"region": "NC",
"regionName": "North Carolina",
"city": "Charlotte",
"zip": "28202",
"lat": 35.2327,
"lon": -80.8461,
"timezone": "America/New_York",
"isp": "FiberPower LLC",
"org": "FiberPower LLC",
"as": "AS214483 FiberPower LLC",
"query": "38.13.166.129"
}
```
This shows the API successfully excluded the targeted city, and routed through an allowed location instead.
***
## Troubleshooting Tips
* Make sure exclusions use correct names (e.g., "newyork" not "New York").
* City/state names should not have spaces.
* Double-check that you are not mixing exclusion types in one request.
* If you get `407` errors, check username/password.
***
# How to Set Proxy Session Lifetime (/docs/proxies/guides/residential-proxies/identifying-session-lifetime)
This guide explains how to configure session lifetime in Geonode to control how long each proxy session stays active.
***
## What Is Session Lifetime
Session lifetime defines how long a proxy connection remains active before resetting.\
Setting the right duration helps to:
* Maintain session consistency for account management.
* Optimize proxy usage for scraping, automation, and security.
* Prevent detection by avoiding overly frequent IP changes.
***
## Ways to Configure Session Lifetime
You can set session lifetime in two ways:
1. Through the Geonode Dashboard (covered in this guide).
2. Via an API request → [Set Session Lifetime via API](/docs/proxies/api-reference/sticky-session/get-sticky-session/get-create).
***
## Step 1: Select a Service or Product
Make sure you have an active **Residential Proxies** plan in your Geonode account.\
From the sidebar, open **Proxy Configuration** to start setting up your proxies.
***
## Step 2: Configure Your Proxy
1. In **Proxy Configuration**, select the **Sticky** proxy type.
2. Choose a proxy protocol: **HTTP/HTTPS** or **SOCKS5**.
* Changing session lifetime automatically updates the generated proxy list.
* The new duration will apply to all selected sessions.
***
## Step 3: Set the Session Lifetime
By default, session lifetime is **10 minutes**, but you can adjust it as needed.
* **Minimum:** 3 minutes
* **Maximum:** 24 hours (1440 minutes)
You can enter the value in either minutes or hours, depending on your preference.
You can change this setting directly in the Dashboard or through the API.
***
* Use shorter sessions for frequent IP changes (for example, scraping). -
Choose longer sessions for stability in login or account-based workflows. -
Adjust settings as needed to balance performance and anonymity.
# New Sticky Session (/docs/proxies/guides/residential-proxies/new-sticky-session)
import VerifyProxyConnectionComponent from "../../../../snippets/verify-proxy-connection-component.mdx";
import SupportParagraph from "../../../../snippets/support-paragraph.mdx";
import MonitorProxyUsage from "../../../../snippets/monitor-proxy-usage-component.mdx";
This guide explains how to create a new sticky session with Geonode.
***
## What Are Sticky Proxies
Sticky proxies allow you to keep the same IP address for a set duration, making them ideal for activities that require stable connections — such as managing social media accounts, scraping sites with login sessions, or performing long-running data collection.
➡️ See the detailed explanation here: [What Are Sticky Session Proxies](/docs/proxies/guides/residential-proxies/new-sticky-session)
***
## Geonode's Sticky Proxy Ports
Geonode assigns specific port ranges for sticky sessions to maintain consistent IPs over time:
* **HTTP:** 10000–10900
* **SOCKS5:** 12000–12010
***
## How to Create a Sticky Session via the Geonode API
Before creating a session, configure your proxy settings in the Geonode dashboard.
### Step 1: Configure Your Proxy Settings
1. Log in to the **Geonode Dashboard**.
2. Open the **Proxies** section from the sidebar.
3. Select **Sticky** as your session type.
4. Set the session duration:
* Minimum: 3 minutes
* Maximum: 24 hours (1440 minutes)
5. Leave other fields as default unless you have specific preferences.
Example configuration:
* IP Type: *Residential*
* Gateway: *France*
* Port Range: *HTTP 10000–10900*
* Session Type: *Sticky Session*
Once configured, copy your proxy endpoint, which looks like this:
```
92.204.164.15:10000:geonode_demouser-type-residential-lifetime-3-RAOnwR
```
Where:
* IP → `92.204.164.15`
* Port → `10000`
* Username → `geonode_demouser-type-residential-lifetime-3-RAOnwR`
* Password → `demopass`
You can use this configuration for API access or in third-party tools such as
Chrome extensions.
***
### Step 2: Create a Sticky Session via API
To start a sticky session through the API, make a request to the `:` endpoint using your configured username and password.
Example using **cURL**:
```bash
curl -x 92.204.164.15:10000 \
--user "geonode_demouser-type-residential-lifetime-3-RAOnwR:demopass" \
--url "http://ip-api.com/json" \
--header "Accept: application/json"
```
Follow this Sticky session API Guide to learn more about the API [Create a New Sticky Session](/docs/proxies/api-reference/sticky-session/get-sticky-session/get-create)
***
***
***
***
## FAQs
{" "}
{" "}
Sticky sessions maintain the same IP for a defined period, while rotating proxies
assign a new IP with every request.{" "}
{" "}
The session duration depends on your configuration — from 3 minutes to 24 hours.
You set this when creating the session.{" "}
{" "}
No, session parameters like IP and type cannot be changed once created. You can,
however, release the current session and create a new one with updated settings.{" "}
{" "}
# Overview (/docs/proxies/guides/residential-proxies/overview)
Residential Proxies give you access to real residential IPs for browsing, automation, and data collection. They are billed pay-as-you-go by traffic (per GB), and you configure targeting, protocol, and session type from the Geonode dashboard.
## How Residential Proxies Work
Residential Proxies use a traffic-based plan model:
* Usage is billed per GB.
* Your remaining balance is shown in GB in the dashboard.
* Traffic is consumed as you send requests through the proxy.
* Pricing starts from the rate shown on your Residential Proxies plan.
* You can configure endpoints, targeting, and session type before connecting.
## Configure Your Proxy
You can configure your Residential Proxy from the Geonode dashboard under **Proxies → Residential Proxies → Proxy configuration**.
Available configuration options include:
* IP type set to Residential
* Gateway selection
* Country, state, and city targeting
* ASN/ISP targeting
* OS targeting
* Protocol selection such as HTTP/HTTPS
* Session type such as Rotating or Sticky
* Endpoint count and format generation
* API credentials for username and password authentication
The dashboard also includes **Port configuration**, **Active sticky sessions**, **Statistics**, **Block list**, and **Reseller**.
See [Block List](/docs/proxies/guides/residential-proxies/block-list) to block outlier domains, IPs, or wildcards so the proxy never fetches them.
## Targeting
Residential Proxies support granular geo and network targeting:
* **Country targeting** — route traffic through a specific country
* **State targeting** — narrow traffic to a state or region
* **City targeting** — target a specific city
* **ASN/ISP targeting** — route through a specific ISP or ASN
* **OS targeting** — filter residential devices by operating system
See [Target Specific Location](/docs/proxies/guides/residential-proxies/target-specific-location) for geo-targeting workflows.
See [Exclude Specific Location](/docs/proxies/guides/residential-proxies/exclude-specific-location) when you need to avoid certain locations.
## Sessions
Residential Proxies support both rotating and sticky sessions:
* **Rotating** — a new IP is used across requests, typically over rotating ports such as HTTP `9000–9010`
* **Sticky** — keep the same IP for a session when you need continuity
See [Rotating Proxies](/docs/proxies/guides/residential-proxies/rotating-proxies) to use rotating endpoints.
See [New Sticky Session](/docs/proxies/guides/residential-proxies/new-sticky-session) and [How to Release Sticky Sessions](/docs/proxies/guides/residential-proxies/release-a-sticky-session) for sticky session workflows.
## Endpoints and Credentials
From **Proxy configuration**, you can:
1. Copy your API username and password.
2. Choose an endpoint format such as `hostname:port:username:password`.
3. Generate one or more proxy endpoints.
4. Use the generated endpoints in your tools or scripts.
For code samples generated from your configuration, see [API Code Generator](/docs/proxies/guides/residential-proxies/api-code-generator).
## Monitoring
Use the dashboard **Statistics** view and usage guides to monitor traffic and performance.
See [Usage Stats & Analytics](/docs/proxies/guides/residential-proxies/usage-stats-analytics) for monitoring residential proxy usage.
See [Latency Testing](/docs/proxies/guides/residential-proxies/success-latency) when you need to test response performance.
## Getting Started
To start using Residential Proxies:
1. Open the Geonode dashboard.
2. Go to **Residential Proxies**.
3. Open **Proxy configuration**.
4. Set your targeting, protocol, and session type.
5. Copy your credentials and generated endpoints.
If you are new to Geonode proxies, start with the [Quick Start Guide](/docs/proxies/getting-started/quick-start).
# How to Release Sticky Sessions (/docs/proxies/guides/residential-proxies/release-a-sticky-session)
This guide explains how to release sticky sessions in Geonode to refresh your proxy connection, switch to a new IP, or free up resources.
***
## Why Release a Sticky Session
Releasing a sticky session ends your current connection and starts a new one.\
This helps when you need to:
* Switch to a different proxy IP.
* Free up resources such as ports or threads.
* Improve security by refreshing your active session.
***
## How to Release Selected Sticky Sessions
Follow these steps to release specific sticky sessions from your Geonode Dashboard.
### Step 1: Select the Sessions to Release
1. Go to the **Active Sticky Sessions** section in your Geonode Dashboard.
2. Select the sessions you want to release — you can select multiple at once.
***
### Step 2: Release Selected Sessions
1. After selecting sessions, click **Release Selected**.
2. Confirm by clicking **Release Sessions**.
3. A success notification will appear once the sessions are released.
***
## How to Release All Sticky Sessions
To disconnect all active sessions at once:
1. Click **Release All Sessions** in the Dashboard.
2. Confirm by selecting **Release All**.
3. All active connections will be immediately terminated.
***
* If a sticky session shows slow performance or fails to connect, release it
to get a new IP. - Use **Release Selected** for specific sessions and
**Release All** for bulk actions. - Refresh sticky sessions periodically to
maintain security and performance.
# Rotating Proxies (/docs/proxies/guides/residential-proxies/rotating-proxies)
import VerifyProxyConnectionComponent from "../../../../snippets/verify-proxy-connection-component.mdx";
import SupportParagraph from "../../../../snippets/support-paragraph.mdx";
import ExtensionFAQs from "../../../../snippets/extensions-faqs.mdx";
This guide walks you through how to use Rotating Proxies in Geonode.
***
## What Are Rotating Proxies
Rotating proxies automatically change your IP address with every request.\
This feature helps maintain anonymity, prevent detection, and ensure smoother automation.
➡️ For a detailed explanation, see [What Are Rotating Proxies](/docs/proxies/guides/residential-proxies/rotating-proxies)
***
## Geonode's Rotating Proxies Range
Geonode provides rotating proxies via HTTP and SOCKS5 protocols, with the following ports:
* **HTTP:** 9000–9010
* **SOCKS5:** 11000–11010
Each request you send will use a new IP address from a different location, helping to avoid IP-based blocking.
***
## How to Use Rotating Proxies with the Geonode API
### Step 1: Configure Your Proxy Settings
In your Geonode dashboard:
1. Open **Proxy Configuration**.
2. Select:
* IP Type: *Residential*
* Gateway: *France* (or your preferred location)
* Protocol: *HTTP/HTTPS* or *SOCKS5*
* Session Type: *Rotating*
Example configuration:
***
### Step 2: Generate Proxy Endpoints
Once configured, generate your proxy endpoints.\
They will typically look like this:
```
92.204.164.15:9000:geonode_demouser-type-residential:demopass
```
Each endpoint includes your IP, port, and authentication credentials (username and password).
***
### Step 3: Make API Calls Using Rotating Proxies
Here’s a Python example using the `requests` library:
```python
import requests
username = "geonode_demouser-type-residential"
password = "demopass"
GEONODE_DNS = "92.204.164.15:9000"
url = "http://ip-api.com"
proxy = {
"http": f"http://{username}:{password}@{GEONODE_DNS}"
}
response = requests.get(url, proxies=proxy)
print("Response:\n", response.text)
```
Each request will use a different IP address from Geonode’s proxy pool.
***
***
***
## FAQs
{" "}
{" "}
Yes. You can add multiple proxies and switch between them by selecting the desired
one and clicking **Connect**.{" "}
{" "}
Rotating proxies: - Maintain anonymity by changing IPs per request. - Prevent
blocks tied to a single IP. - Work perfectly for scraping, SEO, and geo-restricted
access.{" "}
{" "}
* **HTTP:** Best for browsing, scraping, and API requests (HTTP/HTTPS traffic).
* **SOCKS5:** More flexible, supporting FTP, VoIP, and P2P connections.{" "}
{" "}
Yes, you can configure rotating proxies to target countries, cities, or ISPs
to get region-specific IPs.{" "}
{" "}
If one proxy IP is blocked, Geonode automatically rotates to a new IP from the
pool on the next request.{" "}
{" "}
Absolutely. They’re ideal for scraping since each request uses a new IP, reducing
the risk of detection.{" "}
{" "}
You can track requests, performance, and errors directly from your Geonode dashboard.{" "}
{" "}
Usually, yes. Rotating proxies provide better anonymity and resistance to bans,
while static proxies are easier to track and block.{" "}
{" "}
# Latency Testing (/docs/proxies/guides/residential-proxies/success-latency)
This guide explains how to test the latency and success rate of your proxies using Python.\
You’ll learn how to measure proxy performance, analyze results, and visualize latency data.
***
## Overview
The provided script uses Python’s `requests` library to send concurrent requests and measure performance metrics like:
* **Latency (response time)**
* **Success rate (status code analysis)**
* **Error types (timeouts, connection issues)**
It supports both **SOCKS5** and **HTTPS** proxies and uses [ip-api.com](http://ip-api.com) as the default target URL.\
You can also test alternative endpoints like **Cloudflare trace** for comparison.
***
## What You’ll Learn
* How to run latency tests on multiple proxies.
* How to configure proxy protocols (SOCKS5 or HTTP).
* How to analyze average latency, median response time, and success rates.
* How to interpret graphical results to detect issues.
***
## Default Settings
| Parameter | Default Value | Description |
| -------------- | ------------------------------- | ------------------------------------------ |
| Protocol | HTTPS | Set `use_socks5 = True` for SOCKS5 proxies |
| Target | [ip-api.com](http://ip-api.com) | Default endpoint for IP tests |
| Concurrency | 5–10 threads | Recommended thread range |
| Testing Volume | 2000+ requests | Recommended for stable averages |
***
## Variable Settings
The script allows you to adjust several key variables:
* **Target Country** – optional (leave blank to use random pool)
* **Protocol** – choose between HTTP/HTTPS or SOCKS5
* **Session Type** – rotating (default)
* **Target URL** – endpoint to test (default: ip-api.com)
* **Number of Requests** – define test size
* **Threads** – set concurrent workers (recommended: 5–10)
Results may vary depending on which gateway you use. Geonode currently
provides three gateway locations.
***
## Port Ranges
| Session Type | Protocol | Port Range | Description |
| ------------ | ---------- | ----------- | ---------------------------------- |
| Rotating | HTTP/HTTPS | 9000–9010 | Changes IP for every request |
| Rotating | SOCKS5 | 11000–11010 | Changes IP for every request |
| Sticky | HTTP/HTTPS | 10000–10900 | Keeps same IP for session duration |
| Sticky | SOCKS5 | 12000–12010 | Keeps same IP for session duration |
***
## Steps to Test Proxy Latency
### Step 1: Install Required Libraries
Install dependencies before running the script:
```bash
pip install requests matplotlib numpy
```
### Libraries used
* `requests` — send HTTP requests
* `matplotlib` — visualize latency results
* `numpy` — calculate average and median latency
* `collections` — count occurrences of errors
***
## Step 2: Configure Proxy Settings
You can test **SOCKS5** or **HTTP(S)** proxies by adjusting the `use_socks5` flag.
### SOCKS5 Proxy Example
```python
use_socks5 = True
proxies = {
'http': 'socks5://username:password@proxy.geonode.io:11009',
'https': 'socks5://username:password@proxy.geonode.io:11009'
}
```
### HTTP Proxy Example
```python
use_socks5 = False
proxies = {
'http': 'http://username:password@proxy.geonode.io:9008',
'https': 'http://username:password@proxy.geonode.io:9008'
}
```
Replace `username` and `password` with your Geonode credentials
## Step 3: Configure Test Parameters
Set your test parameters in the script:
```python
url = 'http://ip-api.com' # Default test target
num_requests = 5000 # Number of requests
num_workers = 5 # Concurrent threads
```
For accuracy:
* Use **2000+ requests**
* Run with **5–10 threads**
***
## Step 4: Send Concurrent Requests
The script uses `ThreadPoolExecutor` to send multiple requests simultaneously:
```python
from concurrent.futures import ThreadPoolExecutor
import requests, time
def fetch_url(i):
try:
start_time = time.time()
response = requests.get(url, proxies=proxies, timeout=60)
latency = time.time() - start_time
return response.status_code, latency
except requests.exceptions.Timeout:
return 'Timeout', time.time() - start_time
except requests.exceptions.RequestException:
return 'Error', time.time() - start_time
```
Each request logs:
* **Status code**
* **Latency (seconds)**
* **Timeouts or errors**
***
## Step 5: Analyze and Visualize Results
After completing all requests, the script plots a latency graph.
| Color | Meaning |
| ----------- | -------------------------------- |
| **Blue** | Successful requests (status 200) |
| **Red** | Non-200 responses (404, 500) |
| **Green** | Timeouts |
| **Magenta** | Other errors |
***
## Step 6: View Test Statistics
Once the test completes, the script outputs key metrics:
```
Average Latency: 1.20 seconds
Median Latency: 0.89 seconds
Standard Deviation of Latency: 1.30 seconds
Status Code Percentages:
200: 99.54%
Error: 0.12%
401: 0.04%
402: 0.06%
500: 0.20%
502: 0.04%
Total Requests: 5000
Successful (Status 200): 4977
Timeouts: 0
Other Errors: 6
Error Messages:
Error: 6
```
These values help assess **reliability**, **consistency**, and **stability** of your proxy connections.
***
## Step 7: Interpret Results
Use the output to evaluate proxy quality:
| Metric | Meaning |
| ---------------------- | --------------------------------------- |
| **Success Rate** | Percentage of status 200 responses |
| **Latency** | Average response time per request |
| **Error Distribution** | Frequency of timeouts or other failures |
Rotating ports provide more accurate, diversified benchmarks than testing a single static proxy.
***
## Source Code
You can find the full source code for this script on GitHub:
[Geonode Proxy Testing Toolkit](https://github.com/geonodecom/proxy-testing-toolkit/tree/performance-testing/success-latency)
```
```
# Switch Between Different Proxy Locations (/docs/proxies/guides/residential-proxies/switch-between-different-proxy-locations)
import BrowsersFaqs from "../../../../snippets/browsers-faqs.mdx";
import VerifyProxyConnectionComponent from "../../../../snippets/verify-proxy-connection-component.mdx";
import SupportParagraph from "../../../../snippets/support-paragraph.mdx";
This guide explains how to switch between different proxy locations — for example, using proxies from various countries to access region-specific websites or services.
***
## Configure Proxy Based on Location
To switch proxies by country, state, or city, follow the detailed setup instructions here:
➡️ [How to Target Specific Location with Proxies](/docs/proxies/guides/residential-proxies/target-specific-location)
That guide explains how to choose and configure a proxy for a specific region, helping you set the right routing for your use case.
***
## Geonode Proxy Manager (Chrome Extension)
If you use Google Chrome, Geonode provides a dedicated browser extension to manage and switch proxies quickly.
➡️ [How to Use the Geonode Chrome Extension for Proxy Management](/docs/proxies/getting-started/setup_and_configuration/extensions/geonode-proxy-manager)
The extension lets you:
* Switch between multiple proxy locations easily.
* Save frequently used proxies.
* Test and verify active connections in one click.
***
## FoxyProxy (Alternative Browser Extension)
If you prefer a third-party tool, **FoxyProxy** is a popular extension that allows switching proxies by location or rule-based logic.
➡️ [How to Set Up Proxy in FoxyProxy](/docs/proxies/getting-started/setup_and_configuration/extensions/foxyproxy)
FoxyProxy supports:
* Multiple proxy profiles
* Location-based routing
* One-click switching between proxies
You can use any other proxy management extension as long as it supports location-based switching.
***
## Rotating Proxy
If you don’t need a specific location and prefer automatic IP rotation, use Geonode’s **Rotating Proxy Service**.\
It automatically assigns a new IP address from a different location for each request — ideal for tasks like scraping or multi-region testing.
➡️ [How to Set Up Rotating Proxies](/docs/proxies/guides/residential-proxies/rotating-proxies)
***
***
***
# Target Specific Location (/docs/proxies/guides/residential-proxies/target-specific-location)
import ProxyInfoFromGeonode from "../../../../snippets/get-proxy-info-from-geonode.mdx";
import RegionSpecificFAQs from "../../../../snippets/region-specific-faqs.mdx";
import SupportParagraph from "../../../../snippets/support-paragraph.mdx";
***
This guide will help you understand how to geo-target proxies based on locations like cities, countries, states, and ISPs using the Geonode API..
***
## Prerequisites
Before you begin, ensure that you:
* Have valid Geonode API credentials.
* Have basic knowledge of making API requests (e.g., using cURL or Python).
***
## What is Geo-Targeting?
Geonode's **geo-targeting** capabilities allow you to route requests through proxies located in specific countries, cities, states, or regions. This enables precise location-based proxy management for your needs.
➡️ **Learn more about Geo-Targeting**: [What is Geo-Targeting](/docs/proxies/getting-started/knowledge-base/geo-targeting)
***
## Ways to Target Specific Locations in Geonode
Geonode allows you to target proxies based on:
* **Country**
* **City**
* **State**
* **ISP/ASN**
You can't target both **state** and **city** at the same time.
***
## API Endpoint Structure for Geo-Targeting
Geonode's geo-targeting endpoints are designed to accept location parameters such as city, state, country, or ISP/ASN. Simply append the specific parameter (e.g., `-country-`) after the Geonode username.
```bash
curl --request GET \
-x "http://proxy.geonode.io:" \
--user "-country-:" \
--url "http://ip-api.com/json"
```
Replace the placeholder values with your Geonode username, password, and desired location.
***
## Generating User Credentials Based on Location
Since the location needs to be included in the username, Geonode makes it easy to generate a location-based proxy username via the Geonode Dashboard:
* Go to the Geonode Dashboard.
* Scroll down to Proxy Configuration.
* On the right, you'll find various location targeting options.
* Once you configure the location, you'll see the updated username, which you can use in API calls for geo-targeted requests.
***
## Calling Geo-Targetting API
To target a specific city, state, or country using the Geonode API, use the following endpoint formats:
### Geo-Targeting by Country
To target a proxy by **country**, append `-country-` to the username.
**Example**
```bash
curl --request GET \
-x "http://proxy.geonode.io:" \
--user "-country-:" \
--url "http://ip-api.com/json"
```
Check out the detailed API docs here ---> [Perform Country Targeting](/docs/proxies/api-reference/geo-targeting/get-country)
***
### Geo-Targeting by State
To target a proxy by **state**, append `-state-` to the username.
**Example**
```bash
curl --request GET \
-x "http://proxy.geonode.io:" \
--user "-country--state-:" \
--url "http://ip-api.com/json"
```
Check out the detailed API docs here ---> [Perform State Targeting](/docs/proxies/api-reference/geo-targeting/get-state)
***
### Geo-Targeting by City
To target a proxy by **city**, append `-city-` to the username.
**Example**
```bash
curl --request GET \
-x "http://proxy.geonode.io:" \
--user "-country--city-:" \
--url "http://ip-api.com/json"
```
Check out the detailed API docs here ---> [Perform City Targeting](/docs/proxies/api-reference/geo-targeting/get-city)
***
### Geo-Targeting by ISP/ASN
Geonode also allows you to target proxies based on the **ASN (Autonomous System Number)** of an ISP. To do this, append `-asn-` after specifying the country.
**Example**
```bash
curl --request GET \
-x "http://proxy.geonode.io:" \
--user "-country--asn-:" \
--url "http://ip-api.com/json"
```
Check out the detailed API docs here ---> [Perform ISP/ASN Targeting](/docs/proxies/api-reference/geo-targeting/get-isp)
***
## API Response
The API will respond with detailed information about the targeted location, including:
* **Region**, **City**, **ISP**, **Latitude**, **Longitude**, and more.
**Example Response:**
```json
{
"status": "success",
"city": "New York",
"region": "New York",
"country": "United States",
"latitude": "40.7128",
"longitude": "-74.0060",
"isp": "ISP Example",
"timezone": "America/New_York",
"postal_code": "10001"
}
```
This shows that the request was routed through the proxy located in **New York**, with corresponding geographical details.
***
## Troubleshooting
If you encounter issues when targeting a specific location:
* Double-check that the city, state, or country name is spelled correctly.
* Ensure that the proxy is available and accessible in the targeted location.
* Verify the authentication credentials if you receive errors related to access or authorization.
***
# UDP over SOCKS5 for Residential Proxies (/docs/proxies/guides/residential-proxies/udp_over_socks5)
UDP support is available for **Geonode Residential Proxies** when using the **SOCKS5** protocol.
This guide explains how to enable UDP support, configure your proxy, understand how UDP works over SOCKS5, and verify your setup using a simple DNS test.
## Availability
UDP support is available for **Geonode Residential Proxies** when using the SOCKS5 protocol.
To enable UDP support, the proxy username must include the following flag:
```text
-requireUdp-true
```
Use your normal proxy password. Do not change the password format.
## Configure Your Proxy
Use the following proxy settings:
```text
Protocol: SOCKS5
Host: proxy.geonode.io
Port: 12000
Username: -requireUdp-true
Password:
```
Example:
```text
socks5://-requireUdp-true:@proxy.geonode.io:12000
```
## How UDP Works
Geonode supports **UDP over SOCKS5** using the standard SOCKS5 `UDP ASSOCIATE` command.
This is not a direct raw UDP connection between the client and the destination server. Instead, the client first establishes a standard SOCKS5 TCP connection with the proxy. After authentication, the client requests the proxy to create a UDP relay.
The communication flow is:
```text
1. Client connects to the Geonode SOCKS5 proxy over TCP.
2. Client authenticates using the proxy username and password.
3. Client sends a SOCKS5 UDP ASSOCIATE request.
4. The proxy returns a UDP relay IP address and port.
5. Client sends UDP packets to the relay.
6. The proxy forwards the UDP packets to the destination server.
7. UDP responses return through the relay back to the client.
```
> The TCP control connection must remain open while UDP traffic is active. If the TCP connection closes, the UDP association may also terminate.
## Test UDP with DNS
DNS is one of the simplest ways to verify UDP support because standard DNS queries commonly use UDP port `53`.
You can use any SOCKS5 client, library, or tool that supports the SOCKS5 `UDP ASSOCIATE` command.
> Some SOCKS5 clients only support TCP proxying. Those tools can verify SOCKS5 authentication but cannot verify UDP support.
For example, send a DNS query through the SOCKS5 UDP relay:
```text
DNS Server: 1.1.1.1
Port: 53
Domain: example.com
```
If a DNS response is returned, UDP over SOCKS5 is working correctly.
## DNS Test Workflow
A typical DNS test follows these steps:
```text
1. Open a SOCKS5 TCP connection to proxy.geonode.io:12000.
2. Authenticate using the proxy username and password.
3. Send a SOCKS5 UDP ASSOCIATE request.
4. Receive a UDP relay IP address and port.
5. Send a UDP DNS query through the relay to 1.1.1.1:53.
6. Confirm that a DNS response is received.
```
## Expected Output
A successful test should produce results similar to the following:
```text
SOCKS5 TCP connection: successful
SOCKS5 authentication: successful
UDP ASSOCIATE: accepted
UDP relay returned: :
UDP DNS query sent to: 1.1.1.1:53
UDP response: received
DNS result: answer received for example.com
Result: UDP over SOCKS5 is working.
```
## Understanding the Results
```text
SOCKS5 authentication: successful
```
The proxy username and password are valid.
```text
UDP ASSOCIATE: accepted
```
The proxy accepted the SOCKS5 `UDP ASSOCIATE` request and created a UDP relay.
```text
UDP response: received
```
The UDP packet was successfully forwarded through the proxy relay, reached the DNS server, and the response was returned to the client.
```text
DNS result: answer received
```
The DNS server successfully resolved the requested domain.
## Troubleshooting
If authentication fails:
```text
Authentication failed
SOCKS5 username/password auth failed
status=1
```
Verify the username, password, service access, and account status. Also confirm that you are using valid **Residential Proxy** credentials.
If the UDP ASSOCIATE request fails:
```text
UDP ASSOCIATE failed
Command not supported
Connection not allowed by ruleset
SOCKS5 reply code 7
SOCKS5 reply code 2
```
Confirm that you are using:
* Residential Proxies
* SOCKS5
* Port `12000`
* The `-requireUdp-true` username flag
If the UDP request times out:
```text
timed out
UDP response timeout
No UDP response received
```
The SOCKS5 connection may have succeeded, and the proxy may have returned a UDP relay, but the UDP packet or response did not complete.
For Geonode Residential Proxies, confirm that your username includes:
```text
-requireUdp-true
```
If the flag is missing, authentication may still succeed, but the UDP relay test can time out because the session was not routed through the UDP-enabled path.
Timeouts may also occur because of firewall rules, local network restrictions, destination-side filtering, or an incorrect proxy configuration.
## Summary
To use UDP with Geonode Residential Proxies, use the **SOCKS5** protocol on port **12000** and include the `-requireUdp-true` flag in your proxy username.
Geonode implements UDP support through the SOCKS5 `UDP ASSOCIATE` command, allowing the proxy to establish a UDP relay after successful authentication.
A DNS query sent through the SOCKS5 proxy is the simplest way to verify that UDP has been configured correctly.
# Usage Stats & Analytics (/docs/proxies/guides/residential-proxies/usage-stats-analytics)
This guide explains how to track, analyze, and optimize your proxy usage using the Geonode dashboard.\
Monitoring usage statistics helps you manage bandwidth efficiently, improve performance, and troubleshoot issues early.
***
## Step 1: Access the Geonode Dashboard
1. Log in to your **Geonode account**.
2. Open the **Dashboard** to see an overview of your key information:
* Available bandwidth
* Billing cycle and renewal dates
***
## Step 2: View Usage Statistics
In the **Statistics** section, you can explore detailed proxy performance metrics.
### Request Metrics
A visual overview of your request activity, including:
* **Successful Requests** – Proxies that connected successfully.
* **Failed Requests** – Requests that encountered errors or blocks.
* **Total Requests Sent** – The complete count of requests made through your proxies.
***
### Data Usage Monitoring
Keep track of your bandwidth and usage trends:
* **Daily Bandwidth Consumption** – See how much data is used each day.
* **Hourly Usage Trends** – Identify peak usage hours to optimize resources and balance workloads.
This information helps you spot inefficiencies and adjust configurations for better proxy performance.
***
## Final Tips
* Check usage regularly to ensure efficient proxy operation.
* Review failed requests to identify network or configuration issues.
* Use hourly and daily analytics to optimize your bandwidth strategy.
By leveraging Geonode’s analytics tools, you can maintain stable connections, reduce errors, and optimize your proxy performance.
# How to Use the Secure EU Gateway (France – Whitelisted IP Only) (/docs/proxies/guides/residential-proxies/use-the-secure-eu-gateway)
To ensure maximum uptime and performance, Geonode provides a secure EU gateway available only to paying users.\
This gateway offers enhanced reliability and protection through IP whitelisting.
***
## Gateway Details
* **IP Address:** 92.204.164.13
* **Hostname:** prod-proxy.geonode.io
* **Location:** France
* **Name:** France (Whitelisted IP Only)
* **Access Type:** IP-whitelisted access for paid users only
***
## How to Use
1. Log in to your **Geonode Dashboard**.
2. Go to the **IP Whitelist** section.
3. Add your current IP address.
4. In your proxy configuration, select the gateway location: *France (Whitelisted IP Only)*.
5. Save your configuration and connect through the new gateway.
***
Access will be denied for any requests coming from non-whitelisted IP
addresses.
This gateway is recommended for users who need higher reliability and tighter access control.
# Crawl Overview (/docs/scraper-api/dashboard-guides/crawl/00_crawl_overview)
The Crawl dashboard provides a workspace for crawling a website, reviewing recent crawl jobs, and monitoring crawl usage.
## Account Overview
At the top of the Crawl dashboard, you can view your current request usage and plan.
The account overview shows:
* **Available Requests** — Number of requests currently available, or Unlimited depending on your plan.
* **Threads in use** — Current thread usage for your plan.
* **Used last 24 hours** — Number of requests used during the last 24 hours.
* **Plan** — Your current subscription plan and upgrade options.
You can also use **See plans** to review available pricing.
## Crawl Workspace
The main Crawl workspace is where you enter the website URL you want to crawl.
Enter the URL in the field and use:
* **Settings** to configure the crawl.
* **API & Integrations** to access API-related options.
* **Start Crawling** to submit the crawl.
For information about configuring and starting a crawl from the dashboard, see [Crawl a Site](/docs/scraper-api/dashboard-guides/crawl/01_crawl_a_site).
## Recent Crawls
The **Recent Crawls** section displays your previous crawl jobs.
Each crawl displays:
* **Status** — Current status of the crawl job.
* **URL** — Starting URL that was crawled.
* **Pages** — Number of pages crawled, such as `5 / 5`.
* **Created** — Date and time when the crawl was created.
* **Execution Time** — Time taken to complete the crawl.
* **Output** — Output format, such as Markdown, with a download action.
You can also:
* Search previous jobs by URL.
* Filter jobs by status.
* Change the number of rows displayed per page.
* Move between result pages.
* Download crawl output from completed jobs.
## Crawl Status
Crawl jobs can appear with different statuses depending on their current state.
Use the **Status** filter to narrow the list of recent crawls by status.
## Statistics
The Crawl dashboard also includes a **Statistics** tab for viewing crawl usage.
Open **Statistics** to view your crawl activity and usage information.
For more information about managing crawl jobs and monitoring usage, see [Manage Crawl Jobs](/docs/scraper-api/dashboard-guides/crawl/02_manage_crawl_jobs).
## What's Next?
You now know the main sections of the Crawl dashboard.
Continue to [Crawl a Site](/docs/scraper-api/dashboard-guides/crawl/01_crawl_a_site) to learn how to configure a crawl and start crawling from the dashboard.
# Crawl a Site (/docs/scraper-api/dashboard-guides/crawl/01_crawl_a_site)
Use the Crawl dashboard to enter a starting URL, configure crawl options, and start crawling a website.
## Enter the Website URL
Open **Crawl a site** from the Web Data section of the dashboard.
Enter the starting URL in the crawl field. This is the seed URL where the crawl begins.
## Configure Your Crawl
Open **Settings** to configure crawl options before starting.
Available settings include:
| Setting | Description |
| ----------------------- | ---------------------------------------------------------------------------- |
| **Page limit** | Maximum number of pages to crawl. Maximum value is `10000`. |
| **Depth** | How many link levels to follow from the starting URL. Maximum value is `10`. |
| **Stay on same domain** | Keep the crawl limited to the same domain as the starting URL. |
| **Include subdomains** | Allow crawling pages on subdomains of the starting domain. |
| **Output format** | Format used for crawl output, such as Markdown. |
| **JS Rendering** | Enable JavaScript rendering when the target pages require it. |
| **Proxy Type** | Proxy network used for the crawl, such as Residential. |
| **Proxy Country** | Optional country targeting when the target has geo-restrictions. |
For detailed information about crawl parameters and how they affect your request, see [Configuring Crawl Requests](/docs/scraper-api/guides/crawl/02_configuring-crawl-requests).
## Start Crawling
After entering the URL and configuring the available settings, select **Start Crawling**.
The crawl is submitted as a job and appears in the **Recent Crawls** section.
## Review the Crawl Job
Once the crawl is running or completed, find it in **Recent Crawls**.
Each completed crawl shows:
* Status
* Starting URL
* Pages crawled
* Created time
* Execution time
* Output format
Use the download action in the **Output** column to download the crawl results.
## What's Next?
You now know how to configure a crawl, start crawling, and download results from the dashboard.
Continue to [Manage Crawl Jobs](/docs/scraper-api/dashboard-guides/crawl/02_manage_crawl_jobs) to learn how to search, filter, and manage your previous crawls.
# Manage Crawl Jobs (/docs/scraper-api/dashboard-guides/crawl/02_manage_crawl_jobs)
The **Crawl dashboard** lets you manage previous crawls and monitor your crawl activity from one place.
## Recent Crawls
Open the **Recent Crawls** tab to view your previous crawl jobs.
The Recent Crawls section provides:
* **Search by URL** — Find a previous crawl by its starting URL.
* **Status** — Filter crawls by their current status.
* **URL** — View the starting URL used for each crawl.
* **Pages** — See how many pages were crawled.
* **Created** — See when the crawl was created.
* **Execution Time** — See how long the crawl took to complete.
* **Output** — View the output format and download the results.
You can also use the pagination controls to move between pages of crawl jobs and change the number of rows displayed per page.
## Filter Crawl Jobs
Use **Search by URL** to find a specific crawl.
For example, you can enter:
```text
docs.geonode.com
```
to find crawls for that domain.
You can also use the **Status** filter to narrow the list of crawls by their current status.
## Download Crawl Output
Each completed crawl job includes a download action in the **Output** column.
Select the download icon to export the crawl results for that job.
## Statistics
The **Statistics** tab provides an overview of your Crawl API usage.
Open **Statistics** to review crawl activity for a selected time period.
You can typically:
* Select a custom date range.
* Choose a predefined time range, such as the last 24 hours, 7 days, 30 days, or 90 days.
* View the total number of crawls.
* Monitor your success rate.
* Review the average crawl duration.
* See the number of requests used during the selected period.
## What's Next?
You now know how to review previous crawls, filter crawl jobs, download output, and monitor Crawl API usage.
For the API-level details of crawl jobs, see [Managing Crawl Jobs](/docs/scraper-api/guides/crawl/04_managing-crawl-jobs).
# Scraper Overview (/docs/scraper-api/dashboard-guides/extraction/00_extractor_overview)
The Scraper dashboard lets you extract content from web pages without writing code. It is organized into three main sections that help you monitor your account, create extraction requests, and review previous jobs.
This guide introduces each section of the dashboard. The next guides walk you through creating extraction requests and managing your jobs.
## 1. Account Overview
The top section provides information about your account and subscription.
Here you can:
* View the number of available requests.
* Check your free request allowance.
* See when your request quota renews.
* Review the number of requests used during the last 24 hours.
* Upgrade your plan if you need additional requests or higher usage limits.
This section gives you a quick overview of your current usage before submitting new requests.
***
## 2. Extraction Workspace
The center section is where you create new extraction requests.
From this workspace, you can:
* Enter the URL you want to extract.
* Switch between single and multiple URL extraction.
* Open the **Settings** panel to configure your request.
* Access **API & Integrations**.
* Start a new extraction.
This guide only introduces the workspace. The extraction workflows are covered in the following guides:
* Extract Content from a Single URL
* Extract Content from Multiple URLs
For a detailed explanation of each extraction option, refer to the Extraction Guides.
***
## 3. Recent Jobs
The bottom section displays your previous extraction requests.
From here, you can:
* View recent extraction jobs.
* Search previous jobs by URL.
* Filter jobs by status and output format.
* Review execution times.
* Open completed extraction results.
Detailed job management is covered in the **Manage Extraction Jobs** guide.
***
## Next Steps
Now that you're familiar with the Scraper dashboard, continue with **Extract Content from a Single URL** to create your first extraction request.
# Extract Content from a Single URL (/docs/scraper-api/dashboard-guides/extraction/01_extract_single_url)
import ScraperAPIRequestParameters from "../../snippets/requests-parameters.mdx";
The Scraper dashboard lets you extract structured content from a single web page without writing code. Enter the URL, configure the extraction settings if needed, and start the request. Once the extraction is complete, you can preview or download the results directly from the dashboard.
## Step 1. Enter the URL
Enter the URL you want to extract into the URL field.
After entering a valid URL, you can either start the extraction immediately or configure additional settings before submitting the request.
## Step 2. Configure the extraction (Optional)
Click **Settings** to configure the extraction before starting the request.
The Settings panel allows you to configure:
If required, select a specific proxy country or leave it set to **Any** to automatically choose the best available location.
The **Advanced settings** section provides additional request options for more advanced extraction scenarios.
## Step 3. Start the extraction
After reviewing your configuration, click **Start Extraction**.
The dashboard submits the request and begins processing the page.
## Step 4. View the results
When the extraction completes, select the job from **Recent Jobs** to open the result.
The result page includes three output tabs:
* **Preview** – Displays the extracted content in an interactive viewer.
* **Markdown** – Displays the extracted Markdown output and allows you to copy or download it.
* **HTML** – Displays the extracted HTML output and allows you to copy or download it.
Switch between the available output formats using the tabs at the top of the result page.
## Step 5. Download the output
You can download or copy the extracted content directly from the result page.
### Preview
The **Preview** tab lets you inspect the extracted content before downloading it.
### Markdown
The **Markdown** tab allows you to:
* Copy the Markdown output.
* Download the Markdown file.
### HTML
The **HTML** tab allows you to:
* Copy the HTML output.
* Download the HTML file.
## Next Step
Need to extract multiple pages in a single request?
Continue with **Extract Content from Multiple URLs**.
# Extract Content from Multiple URLs (/docs/scraper-api/dashboard-guides/extraction/02_extract_multiple_urls)
import ScraperAPIRequestParameters from "../../snippets/requests-parameters.mdx";
The Scraper dashboard lets you extract content from multiple web pages in a single batch request. You can add URLs manually or import them from a CSV file, configure the extraction settings, and download the results once processing is complete.
## Step 1. Open Batch Extraction
From the Scraper dashboard, click **Set Multiple URLs**.
The **Batch Extraction** dialog opens, allowing you to submit multiple URLs in a single request.
***
## Step 2. Add URLs
You can add URLs in two ways.
### Add URLs manually
Enter a URL into the input field and click **Add in table**. Repeat this step until all required URLs have been added.
### Import a CSV file
If you already have a list of URLs, click **Import CSV** to upload them in bulk.
You can also download the sample CSV template by selecting **Download CSV template**.
The **Batch size** section displays the maximum number of URLs that can be included in a single batch based on your current plan.
***
## Step 3. Configure the extraction (Optional)
Select **Settings** to configure the extraction request before starting the batch.
Available options include:
***
## Step 4. Start the batch extraction
After adding your URLs and reviewing the configuration, click **Start Batch Extraction**.
The dashboard creates a batch job and begins processing each URL.
***
## Step 5. Monitor the batch job
Return to the **Recent Jobs** section and switch to the **Batches** tab.
While the batch is running, you can monitor its progress.
The dashboard displays:
* Current job status
* Total URLs in the batch
* Completed URLs
* Failed URLs
* Creation time
* Selected output format
The completed and failed counters update as each URL finishes processing.
***
## Step 6. Download the results
When the batch status changes to **Completed**, click the **Download** icon.
The dashboard downloads the extraction results for every successfully processed URL.
Depending on the selected output format, each URL is saved as an individual file.
***
## Next Step
Learn how to search, monitor, and manage your extraction history in the **Manage Extraction Jobs** guide.
# Manage Extraction Jobs (/docs/scraper-api/dashboard-guides/extraction/03_manage_extraction_jobs)
The **Recent Jobs** section lets you monitor all extraction requests submitted through the Scraper dashboard. You can switch between single URL and batch jobs, search previous requests, filter jobs by status, download completed results, and review extraction statistics.
## Switch between job types
Open the **Recent Jobs** section and choose the type of jobs you want to view.
* **Single URL** displays individual extraction requests.
* **Batches** displays batch extraction requests.
## Search and filter jobs
For **Single URL** jobs, you can quickly find previous requests using the available filters.
The dashboard allows you to:
* Search by URL.
* Filter jobs by status.
* Filter jobs by output format.
## Filter by status
Select the **Status** filter to display only jobs with a specific status.
Available filters include:
* Queued
* Processing
* Completed
* Failed
* Cancelled
## Monitor batch jobs
When viewing **Batch** jobs, the dashboard displays additional information about each extraction request.
For every batch, you can see:
* Current status
* Number of URLs in the batch
* Completed URLs
* Failed URLs
* Creation date
* Execution time
* Output format
While a batch is running, the **Completed** and **Failed** counters update as each URL finishes processing.
Once processing is complete, select the **Download** icon to download the extraction results.
## View extraction statistics
Open the **Statistics** tab to review your extraction activity.
The Statistics page includes:
* Date range selection
* Total extractions
* Success rate
* Average extraction duration
* Requests used
* Request activity over time
* Request usage charts
Use these statistics to monitor extraction performance and understand how your requests are being used over the selected period.
# Map Overview (/docs/scraper-api/dashboard-guides/map/00_map_overview)
import Link from "next/link";
The Map dashboard helps you discover the URLs available on a website before extracting content. It provides a simple interface for creating mapping requests, reviewing previous jobs, and monitoring your account usage.
Before creating your first mapping request, we recommend reviewing the following guides to understand how the Map API works and the best practices for efficient website discovery.
Understanding Map
Map Workflows
Map Best Practices
These guides explain how the Map API discovers URLs, how mapping requests are processed, and the recommended workflows for achieving the best results.
## 1. Account Overview
The top section provides information about your account and subscription.
Here you can:
* View your available requests.
* Check your free request allowance.
* See when your requests renew.
* Review the number of requests used during the last 24 hours.
* Upgrade your plan if you need additional requests.
This section helps you monitor your available usage before submitting new mapping requests.
## 2. Mapping Workspace
The center section is where you create new mapping requests.
From this workspace, you can:
* Enter the website URL you want to map.
* Open the **Settings** panel to configure the request.
* Access **API & Integrations**.
* Start a new mapping request.
This guide introduces the dashboard only. The complete mapping workflow is covered in the **Create Your First Map** guide.
## 3. Recent Maps
The bottom section displays your previous mapping requests.
From here, you can:
* View recent mapping jobs.
* Search previous requests by URL.
* Filter jobs by status.
* Review execution times.
* View the number of discovered links.
* Open completed mapping results.
Managing previous mapping jobs is covered in the **Manage Map Jobs** guide.
## Next Step
Continue with **Create Your First Map** to submit your first mapping request using the dashboard.
# Discover URLs (/docs/scraper-api/dashboard-guides/map/01_map_discover_urls)
import Link from "next/link";
The Map dashboard helps you discover URLs from a website before extracting content. Enter the website URL, optionally configure the mapping settings, and start the mapping request. Once the mapping is complete, you can review, copy, or download the discovered URLs.
## Step 1. Enter the website URL
Enter the website you want to map into the URL field.
Optionally, you can configure the mapping request before starting it.
Available options include:
* Filter URLs
* Include subdomains
* Ignore query parameters
For a detailed explanation of these options and the recommended mapping workflow, refer to the following guides:
Understanding Map
Map Workflows
Map Best Practices
After reviewing the settings, click **Start Mapping**.
## Step 2. View the mapping job
Once the mapping request has been submitted, it appears in the **Recent Maps** section.
Each completed mapping request displays:
* Status
* Website URL
* Creation date
* Execution time
* Number of discovered links
Click the **View** icon to open the mapping results.
## Step 3. Review the discovered URLs
The Map Preview displays all discovered URLs for the completed mapping request.
From the preview, you can:
* Review all discovered URLs.
* Copy the complete list to your clipboard.
* Download the discovered URLs as a file.
## Next Step
Continue with **Manage Map Jobs** to learn how to search, filter, and review previous mapping requests.
# Manage Map Jobs (/docs/scraper-api/dashboard-guides/map/02_map_manage_jobs)
The **Recent Maps** section lets you review all mapping requests submitted through the Map dashboard. You can search previous requests, filter them by status, open completed results, and monitor your mapping activity using the built-in statistics.
## Search and filter map jobs
Open the **Recent Maps** tab to view your mapping history.
The dashboard allows you to:
* Search previous mapping requests by URL.
* Filter jobs by status.
* Open completed mapping results.
The available status filters include:
* Queued
* Processing
* Completed
* Failed
* Cancelled
Click the **View** icon to open the discovered URLs for a completed mapping request.
## View mapping statistics
Select the **Statistics** tab to review your mapping activity.
The Statistics page provides an overview of your mapping requests for the selected time period.
You can:
* Select a custom date range.
* Choose a predefined time range, such as the last 24 hours, 7 days, 30 days, or 90 days.
* View the total number of mapping requests.
* Monitor your success rate.
* Review the average mapping duration.
* See the number of requests used during the selected period.
The page also includes charts showing request activity and usage trends over time, helping you monitor mapping performance and request consumption.
## Next Step
You're now familiar with the complete Map dashboard workflow, including creating mapping requests, reviewing discovered URLs, and managing previous jobs.
# Search Overview (/docs/scraper-api/dashboard-guides/search/00_search_overview)
The Search dashboard provides a workspace for submitting searches, reviewing recent search jobs, and monitoring search usage.
## Account Overview
At the top of the Search dashboard, you can view your current request usage.
The account overview shows:
* **Available Requests** — Number of requests currently available.
* **Free requests** — Number of free requests included with your account.
* **Renewal date** — When your free requests renew.
* **Used last 24 hours** — Number of requests used during the last 24 hours.
You can also use the **See pricing** option to view available plans.
## Search Workspace
The main Search workspace is where you enter your search query.
Enter your query in the search field and use:
* **Settings** to configure the search.
* **API & Integrations** to access API-related options.
* **Search** to submit the search.
For information about making a search from the dashboard, see [Search and View Results](/docs/scraper-api/dashboard-guides/search/01_search_and_view_results).
## Recent Searches
The **Recent Searches** section displays your previous search jobs.
Each search displays:
* **Status** — Current status of the search job.
* **Query** — Search query that was submitted.
* **Created** — Date and time when the search was created.
* **Execution Time** — Time taken to complete the search.
You can also:
* Search previous jobs by query.
* Filter jobs by status.
* Change the number of rows displayed per page.
* Move between result pages.
* Open an individual search to view its results.
## Search Status
Search jobs can appear with different statuses depending on their current state.
Use the **Status** filter to narrow the list of recent searches by status.
## Statistics
The Search dashboard also includes a **Statistics** tab for viewing search usage.
Open **Statistics** to view your search activity and usage information.
For more information about managing search jobs and monitoring usage, see [Manage Search Jobs](/docs/scraper-api/dashboard-guides/search/02_manage_search_jobs).
## What's Next?
You now know the main sections of the Search dashboard.
Continue to [Search and View Results](/docs/scraper-api/dashboard-guides/search/01_search_and_view_results) to learn how to submit a search and review the returned results.
# Search and View Results (/docs/scraper-api/dashboard-guides/search/01_search_and_view_results)
Use the Search dashboard to enter a search query, configure search options, and review the results directly in the dashboard.
## Configure Your Search
Enter your search query in the Search field.
You can open **Settings** to configure options such as:
* **Locale**
* **Time range**
* **Safe search**
For detailed information about the available search parameters and how they affect your request, see [Search Parameters](/docs/scraper-api/guides/search/03_search_parameters).
### Locale
The **Locale** setting controls the language or region used for the search results.
The default option is **Auto**.
### Time Range
The **Time range** setting allows you to apply a time range to the search.
The default option is **Off**.
### Safe Search
The **Safe search** setting controls the level of filtering applied to the search results.
Available options are:
* **Off**
* **Moderate**
* **Strict**
For the complete list of supported search parameters and their values, see [Search Parameters](/docs/scraper-api/guides/search/03_search_parameters).
## Submit the Search
After entering your query and configuring the available settings, select **Search** to start the search.
The search is then submitted and the results are displayed in the dashboard.
## Review Search Results
Once the search is completed, the results section displays the returned search results.
The results view includes:
* **Result count** — Number of results returned.
* **Query** — Search query used for the request.
* **Locale** — Locale used for the search.
* **Safe search** — Safe search setting used for the search.
* **Status** — Current status of the search.
Each result can include:
* Result position
* Title
* URL
* Snippet
* Displayed URL
## View Results as JSON
The results section provides two views:
* **Results** — Displays the search results in a readable format.
* **JSON** — Displays the returned results as JSON.
Select **JSON** when you need to inspect the returned search data in its JSON format.
## What's Next?
You now know how to configure a search, submit a query, and review the returned results.
Continue to [Manage Search Jobs](/docs/scraper-api/dashboard-guides/search/02_manage_search_jobs) to learn how to search, filter, and manage your previous searches.
# Manage Search Jobs (/docs/scraper-api/dashboard-guides/search/02_manage_search_jobs)
The **Search dashboard** lets you manage previous searches and monitor your search activity from one place.
## Recent Searches
Open the **Recent Searches** tab to view your previous search jobs.
The Recent Searches section provides:
* **Search by query** — Find a previous search by its query.
* **Status** — Filter searches by their current status.
* **Query** — View the query used for each search.
* **Created** — See when the search was created.
* **Execution Time** — See how long the search took to complete.
You can also use the pagination controls to move between pages of search jobs and change the number of rows displayed per page.
## Filter Search Jobs
Use **Search by query** to find a specific search.
For example, you can enter:
```text
web scraping best practice
```
to find searches containing that query.
You can also use the **Status** filter to narrow the list of searches by their current status.
## View a Previous Search
Each search job has an action button on the right side of the table.
Select the action button to open the search and review its results.
The search results open in a separate view where you can switch between:
* **Results** — View the returned search results.
* **JSON** — View the results in JSON format.
The search view also shows the search query and the status of the search.
## Statistics
The **Statistics** tab provides an overview of your Search API usage.
### Date Range
You can select a custom date range or use one of the available time periods:
* **Last 24 hours**
* **7 days**
* **30 days**
* **90 days**
### Usage Summary
The statistics view provides the following information for the selected period:
* **Total** — Total number of searches.
* **Success Rate** — Percentage of successful searches.
* **Average Duration** — Average search duration.
* **Requests Used** — Number of requests used during the selected period.
### Requests
The **Requests** chart shows search requests over the selected period.
The chart distinguishes between:
* Successful requests
* Failed requests
### Requests Usage
The **Requests usage** chart shows request usage for the selected period.
## What's Next?
You now know how to review previous searches, filter search jobs, open their results, and monitor Search API usage.
For the API-level details of search jobs, see [Search Jobs](/docs/scraper-api/guides/search/06_search_jobs).
# Understanding Batch (/docs/scraper-api/guides/batch/00_understanding_batch)
Batch extraction lets you process multiple URLs in a single asynchronous request.
Instead of sending a separate request for every URL, you submit all URLs together as one batch job. Geonode processes each URL in the background and lets you retrieve the results after the job has completed.
The Batch endpoint immediately returns a job ID instead of the extracted content. Use the job ID to monitor progress and retrieve the results when processing finishes.
## When to Use Batch
Batch extraction is useful when you need to process many webpages at once.
Common use cases include:
* Extracting product pages from an e-commerce website
* Processing blog articles or news posts
* Scraping documentation pages
* Monitoring multiple websites
* Running scheduled extraction jobs
If you only need to extract content from a single webpage, use the Extraction endpoint instead.
***
## Choosing the Right Endpoint
Geonode provides different endpoints depending on the task you want to perform.
| If you want to... | Use |
| ------------------------------------------ | ---------- |
| Extract content from a single URL | Extraction |
| Extract content from multiple URLs | Batch |
| Crawl an entire website by following links | Crawl |
***
## Why Use Batch?
Without batch extraction, every URL requires its own API request.
E1["Extract"]
U2["URL 2"] --> E2["Extract"]
U3["URL 3"] --> E3["Extract"]
U4["URL 4"] --> E4["Extract"]
end
subgraph B["With Batch"]
B1["URL 1"]
B2["URL 2"]
B3["URL 3"]
B4["URL 4"]
B1 --> API["Batch API"]
B2 --> API
B3 --> API
B4 --> API
API --> JOB["Batch Job"]
end
`}
/>
Batch extraction groups multiple URLs into a single request, making it easier to process large collections of webpages.
***
## How Batch Works
Batch extraction follows a simple asynchronous workflow.
B["Submit Batch Request"]
--> C["Batch Job Created"]
--> D["Job ID Returned"]
--> E["Processing"]
--> F["Retrieve Results"]
`}
/>
Unlike the Extraction endpoint, Batch does not return the extracted content immediately. Instead, it creates a job and returns a job ID that you can use to monitor progress and retrieve the results once processing has finished.
***
## Batch Workflow
A typical batch request follows these steps:
1. Submit one or more URLs to the Batch endpoint.
2. Geonode creates a new batch job.
3. The API immediately returns a unique job ID.
4. Each URL is processed independently.
5. Retrieve the completed results using the job ID.
***
## Batch vs Extraction
Although both endpoints extract webpage content, they are designed for different workloads.
| Feature | Extraction | Batch |
| ---------------- | --------------------------- | -------------------------------- |
| Number of URLs | One | Multiple |
| Processing | Synchronous or asynchronous | Asynchronous |
| Best for | Individual webpages | Large collections of webpages |
| Initial response | Extracted content or Job ID | Job ID |
| Final result | Single extraction | Collection of extraction results |
***
## Key Concepts
Before using the Batch endpoint, keep the following in mind:
* Each URL is processed independently.
* Some URLs may succeed while others fail within the same batch.
* Batch requests return a job ID instead of the extracted content.
* Results can be retrieved after processing has completed.
***
## Next Steps
Now that you understand how batch extraction works, continue to **Your First Batch** to submit your first batch extraction request.
# Batch Workflows (/docs/scraper-api/guides/batch/01_batch_workflows)
A batch extraction is an asynchronous workflow. Instead of waiting for every URL to finish processing, the API immediately creates a job and returns a unique job ID.
That job ID becomes the center of the workflow. You use it to monitor progress, retrieve results, or cancel the job if necessary.
## Complete Batch Workflow
The following diagram shows how the Batch API endpoints work together.
Every batch job follows this lifecycle, from creation to completion.
***
## Typical Workflow
Most applications follow these steps when working with batch extraction.
### Step 1 — Create a Batch
Start by submitting one or more URLs.
```http
POST /v1/batch
```
The API immediately returns:
* `job_id`
* `status`
* `status_url`
* `accepted_urls`
At this point, the extraction has been queued and continues in the background.
***
### Step 2 — Store the Job ID
Save the returned `job_id`.
You'll need this value for every subsequent operation, including:
* Monitoring progress
* Viewing results
* Cancelling the batch
***
### Step 3 — Monitor Progress
Use the Job Status endpoint until the batch finishes.
```http
GET /v1/batch/{job_id}
```
A batch can move through the following states:
```text
queued
processing
completed
failed
cancelled
```
Most applications poll this endpoint every few seconds until processing finishes.
***
### Step 4 — Process the Results
Once the job reaches the `completed` state, the response contains:
* Successfully processed URLs
* Failed URLs
* Output for each completed extraction
Your application can now download, store, or process the extracted content.
***
## Finding Previous Jobs
Sometimes an application loses the job ID due to a restart or network interruption.
Instead of creating another batch, retrieve the existing one.
```http
GET /v1/batch/jobs
```
This prevents duplicate processing and allows your application to continue from where it left off.
***
## Cancelling a Batch
You can cancel a running batch if it is no longer needed.
Common reasons include:
* Incorrect URLs
* Wrong proxy configuration
* Incorrect request settings
* User cancelled the operation
```http
DELETE /v1/batch/{job_id}
```
After cancellation:
* No new URLs are scheduled for processing.
* URLs that are already running may finish.
* Remaining URLs are skipped.
***
## End-to-End Example
The complete workflow usually looks like this.
| Step | Endpoint | Purpose |
| ---- | --------------------------- | ------------------------------------------------- |
| 1 | `POST /v1/batch` | Create a new batch job |
| 2 | Save `job_id` | Store the identifier for future requests |
| 3 | `GET /v1/batch/{job_id}` | Monitor progress until completion |
| 4 | `GET /v1/batch/jobs` | Find previous jobs when the job ID is unavailable |
| 5 | `DELETE /v1/batch/{job_id}` | Cancel the batch if it is no longer needed |
***
## Common Workflow Patterns
### Standard Batch Processing
This is the most common workflow for asynchronous processing.
***
### Recovering After an Application Restart
This allows applications to recover without creating duplicate jobs.
***
### Cancelling an Active Batch
***
## Best Practices
* Store the returned `job_id` immediately after creating a batch.
* Poll the Job Status endpoint instead of submitting duplicate batch requests.
* Use the List Jobs endpoint to recover lost job IDs.
* Cancel jobs that are no longer required instead of letting them continue processing.
* Wait until the batch reaches the `completed` state before processing the final results.
* Handle every possible job status (`queued`, `processing`, `completed`, `failed`, and `cancelled`) in your application.
***
## Next Steps
You now understand the complete lifecycle of a batch job.
Continue to the **Batch API Reference** for detailed request parameters, response fields, and endpoint-specific examples.
# Your First Batch (/docs/scraper-api/guides/batch/02_your_first_batch)
Batch jobs allow you to submit multiple URLs in a single API request.
Instead of sending one extraction request per page, you create a single batch job. Geonode queues the job, processes each URL asynchronously, and lets you retrieve the results later.
In this guide, you'll:
* Create your first batch job
* Understand the response
* Learn how invalid URLs are handled
* Retrieve the batch status
***
## Create Your First Batch
Send a `POST` request to the batch endpoint.
```http
POST /v1/batch
```
### Request
```bash title="request.sh"
curl -X POST "https://scraper.geonode.io/v1/batch" \
-H "X-Api-Key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"urls": [
"https://geonode.com",
"https://docs.geonode.com",
"https://example.com"
]
}'
```
### Response
```json title="response.json"
{
"job_id": "d16e56a0-dbe9-4586-a40d-028cf3c439a9",
"status": "queued",
"status_url": "/v1/batch/d16e56a0-dbe9-4586-a40d-028cf3c439a9",
"accepted_urls": 3,
"invalid_urls": []
}
```
The request creates a new batch job immediately.
Geonode validates the request, queues the job, and returns a unique job ID that you can use to monitor progress.
***
## Understanding the Response
The initial response confirms that the batch job has been accepted.
| Field | Description |
| --------------- | -------------------------------------------------------- |
| `job_id` | Unique identifier for the batch job. |
| `status` | Current processing status. A new job starts as `queued`. |
| `status_url` | Endpoint used to retrieve the latest status and results. |
| `accepted_urls` | Number of valid URLs accepted for processing. |
| `invalid_urls` | URLs that were rejected during validation. |
The actual extraction results are **not** returned immediately because batch jobs run asynchronously.
***
## What Happens Next?
Once the request is accepted, Geonode processes each URL in the background.
You can use the returned `job_id` or `status_url` to check the progress of the job at any time.
***
## Handling Invalid URLs
By default, every URL in the batch request must be valid.
If one or more URLs are invalid and `ignore_invalid_urls` is set to `false`, the entire request is rejected.
### Request
```json
{
"urls": [
"https://geonode.com",
"https://docs.geonode.com",
"not-a-valid-url"
],
"ignore_invalid_urls": false
}
```
### Response
```json
{
"code": "VALIDATION_ERROR",
"message": "Batch contains invalid URLs and ignore_invalid_urls is false",
"correlation_id": "1fd0958b-6583-4c32-9e9b-5ad7960ac5ad",
"retryable": false
}
```
No batch job is created until the request contains only valid URLs.
***
## Skip Invalid URLs
If you want Geonode to continue processing valid URLs, enable `ignore_invalid_urls`.
### Request
```json
{
"urls": [
"https://geonode.com",
"https://docs.geonode.com",
"https://example.com",
"not-a-valid-url"
],
"ignore_invalid_urls": true
}
```
### Response
```json
{
"job_id": "d16e56a0-dbe9-4586-a40d-028cf3c439a9",
"status": "queued",
"status_url": "/v1/batch/d16e56a0-dbe9-4586-a40d-028cf3c439a9",
"accepted_urls": 3,
"invalid_urls": [
"not-a-valid-url"
]
}
```
Only the valid URLs are queued for processing. Invalid URLs are skipped and returned in the `invalid_urls` array.
Use `ignore_invalid_urls` when importing URLs from user input, CSV files, or external systems where a few invalid entries shouldn't stop the entire batch.
***
## Check the Batch Status
Batch jobs are processed asynchronously.
After creating a job, use the returned `job_id` to check its status.
```http
GET /v1/batch/{job_id}
```
You'll learn how to monitor a running job and retrieve its results in the next guide.
***
## Next Steps
Now that you've created your first batch job, continue to **Working with Batch Inputs** to learn how to customize batch requests with output formats, JavaScript rendering, proxies, custom headers, and other request options.
# Configuring Batch Requests (/docs/scraper-api/guides/batch/03_configuring_batch_requests)
Every batch request starts with a list of URLs.
You can further customize how those URLs are processed by configuring output formats, JavaScript rendering, proxy settings, custom headers, and other optional request parameters.
All examples in this guide create a new batch job.
Regardless of which request options you use, the Batch endpoint immediately returns a job ID and begins processing the accepted URLs in the background.
## Request Body Overview
The following request fields are available when creating a batch job.
| Field | Required | Description |
| --------------------- | -------- | -------------------------------------------------- |
| `urls` | ✅ | List of URLs to process. |
| `ignore_invalid_urls` | No | Continue processing even if some URLs are invalid. |
| `formats` | No | Choose the output format for extracted content. |
| `render_js` | No | Render JavaScript before extraction. |
| `processing_mode` | No | Configure how URLs are processed. |
| `proxy` | No | Route requests through a proxy. |
| `headers` | No | Include custom HTTP headers. |
| `wait_config` | No | Wait for dynamic content before extraction. |
All of these options are optional except `urls`.
***
## URLs
Every batch request must include one or more valid URLs.
```json
{
"urls": [
"https://geonode.com",
"https://docs.geonode.com"
]
}
```
Each URL is processed independently as part of the same batch job.
***
## Ignoring Invalid URLs
By default, every URL must be valid.
If you want Geonode to continue processing valid URLs while skipping invalid ones, enable `ignore_invalid_urls`.
```json
{
"urls": [
"https://geonode.com",
"https://docs.geonode.com",
"not-a-valid-url"
],
"ignore_invalid_urls": true
}
```
For a complete example of how invalid URLs are handled, see **Your First Batch**.
***
## Output Formats
Choose one or more output formats for the extracted content.
```json
{
"urls": [
"https://geonode.com"
],
"formats": [
"markdown",
"html"
]
}
```
Learn more in **Output Formats**.
***
## JavaScript Rendering
Enable JavaScript rendering for websites that load content dynamically.
```json
{
"urls": [
"https://geonode.com"
],
"render_js": true
}
```
Learn more in **JavaScript Rendering**.
***
## Processing Mode
Control how the batch request is processed.
```json
{
"urls": [
"https://geonode.com"
],
"processing_mode": "parallel"
}
```
Learn more in **Processing Modes**.
***
## Proxy Configuration
Use a proxy when accessing geo-restricted or protected websites.
```json
{
"urls": [
"https://geonode.com"
],
"proxy": {
"...": "..."
}
}
```
Learn more in **Proxy & Geo-Targeting**.
***
## Custom Headers
Include custom HTTP headers with every request.
```json
{
"urls": [
"https://geonode.com"
],
"headers": {
"Authorization": "Bearer YOUR_TOKEN"
}
}
```
Learn more in **Using Custom Headers**.
***
## Waiting for Dynamic Content
Delay extraction until specific content becomes available.
```json
{
"urls": [
"https://geonode.com"
],
"wait_config": {
"...": "..."
}
}
```
Learn more in **Waiting for Dynamic Content**.
***
## Complete Example
The following example combines several request options.
```json
{
"urls": [
"https://geonode.com",
"https://docs.geonode.com"
],
"formats": [
"markdown"
],
"render_js": true,
"ignore_invalid_urls": true
}
```
The response immediately returns a queued batch job.
```json
{
"job_id": "e9379def-8986-4624-86f1-f49c35afe711",
"status": "queued",
"status_url": "/v1/batch/e9379def-8986-4624-86f1-f49c35afe711",
"accepted_urls": 2,
"invalid_urls": []
}
```
***
## Next Steps
Continue to **Batch Results** to learn how to monitor batch jobs, retrieve completed results, and understand the different job states.
# Monitoring Batch Jobs (/docs/scraper-api/guides/batch/04_monitoring_batch_jobs)
Batch jobs are processed asynchronously.
After creating a batch job, use the returned `job_id` to monitor its progress and retrieve the extraction results.
***
## Retrieve a Batch Job
Use the following endpoint to retrieve the latest status and results for a batch job.
```http
GET /v1/batch/{job_id}
```
### Request
```bash title="request.sh"
curl -X GET "https://scraper.geonode.io/v1/batch/YOUR_JOB_ID" \
-H "X-Api-Key: YOUR_API_KEY"
```
Replace `YOUR_JOB_ID` with the job ID returned when the batch was created.
***
## Example Response
```json title="response.json"
{
"job_id": "d54da3de-7afc-4127-a5ee-5f7e0378a3e6",
"status": "completed",
"created_at": "2026-07-05T16:14:22.721Z",
"completed_at": "2026-07-05T16:14:25.318Z",
"total_urls": 3,
"completed_urls": 2,
"failed_urls": 1,
"pending_urls": 0,
"cancelled_urls": 0,
"token_summary": {
"tokens_charged_total": 2,
"tokens_reserved": 0
},
"results": [
{
"input_index": 0,
"url": "https://geonode.com",
"status": "completed"
},
{
"input_index": 1,
"url": "https://docs.geonode.com",
"status": "completed"
},
{
"input_index": 2,
"url": "https://this-domain-does-not-exist-123456789.com",
"status": "failed"
}
]
}
```
***
## Job Information
The top-level response describes the overall batch job.
| Field | Description |
| -------------- | ------------------------------------ |
| `job_id` | Unique identifier for the batch job. |
| `status` | Current status of the batch job. |
| `created_at` | Time the batch job was created. |
| `completed_at` | Time processing finished. |
***
## Processing Statistics
These fields summarize the progress of the batch.
| Field | Description |
| ---------------- | -------------------------------------------- |
| `total_urls` | Total number of submitted URLs. |
| `completed_urls` | URLs processed successfully. |
| `failed_urls` | URLs that failed during processing. |
| `pending_urls` | URLs that are still waiting to be processed. |
| `cancelled_urls` | URLs cancelled before processing completed. |
***
## Token Usage
The response also includes a summary of token usage.
| Field | Description |
| ---------------------- | ----------------------------------- |
| `tokens_charged_total` | Total tokens consumed by the batch. |
| `tokens_reserved` | Tokens reserved for processing. |
***
## Individual Results
Each submitted URL appears in the `results` array.
Each result contains information about a single URL.
| Field | Description |
| --------------- | -------------------------------------------- |
| `input_index` | Position of the URL in the original request. |
| `url` | The processed URL. |
| `status` | Processing status for that URL. |
| `error_code` | Error code when processing fails. |
| `error_message` | Human-readable error description. |
| `data` | Extracted content for successful requests. |
| `metadata` | Processing metadata for the request. |
***
## Job Status
Each URL has its own processing status.
| Status | Description |
| ------------ | -------------------------- |
| `queued` | Waiting to be processed. |
| `processing` | Currently being processed. |
| `completed` | Successfully extracted. |
| `failed` | Processing failed. |
| `cancelled` | Processing was cancelled. |
***
## Partial Failures
A batch job can complete successfully even if some URLs fail.
For example, a batch containing three URLs may produce:
* 2 completed URLs
* 1 failed URL
* Overall batch status: `completed`
Failed URLs include additional error information.
```json
{
"url": "https://this-domain-does-not-exist-123456789.com",
"status": "failed",
"error_code": "INTERNAL_ERROR",
"error_message": "An unexpected error occurred on the server"
}
```
This allows you to retry only the failed URLs instead of rerunning the entire batch.
Failed URLs do not prevent other URLs in the same batch from completing successfully.
***
## Metadata
Each completed result includes processing metadata.
| Field | Description |
| ---------------- | ------------------------------------------------ |
| `http_status` | HTTP status code returned by the target website. |
| `duration_ms` | Time taken to process the URL. |
| `tokens_charged` | Tokens consumed for that extraction. |
This information can help monitor performance and troubleshoot extraction issues.
***
## Best Practices
When working with batch jobs:
* Save the returned `job_id` so you can retrieve the results later.
* Wait until the job status is `completed` before using the extracted data.
* Review the processing statistics to quickly identify failed URLs.
* Retry only failed URLs instead of rerunning the entire batch.
* Monitor token usage when processing large batches.
***
## Next Steps
Now that you know how to monitor batch jobs and retrieve their results, continue to **Batch Workflows** to learn common patterns for processing large collections of URLs.
# Listing Batch Jobs (/docs/scraper-api/guides/batch/05_listing_batch_jobs)
Listing batch jobs allows you to retrieve previously created batch requests. This is useful when you want to review completed jobs, monitor running batches, or locate a job ID for checking its status.
## Batch Jobs Endpoint
Use the following endpoint to retrieve previously created batch jobs.
```http
GET /v1/batch/jobs
```
The response returns a paginated list of batch jobs.
## List Batch Jobs
Use the following request to retrieve your recent batch jobs.
### Request
```bash title="request.sh"
curl -X GET "https://scraper.geonode.io/v1/batch/jobs" \
-H "X-Api-Key: YOUR_API_KEY"
```
### Response
```json title="response.json"
{
"jobs": [
{
"job_id": "9e60e068-b6ac-408d-aa6e-f120f41faf0d",
"status": "completed",
"accepted_urls": 3,
"completed_urls": 3,
"failed_urls": 0,
"config": {
"render_js": false,
"formats": [
"markdown"
],
"proxy": {
"country": null,
"type": "residential"
},
"headers": null,
"wait_config": null
},
"created_at": "2026-07-05T16:29:56.344150Z",
"completed_at": "2026-07-05T16:30:13.963398Z"
}
],
"page": 1,
"page_size": 10,
"page_count": 1
}
```
## Understanding the Response
Each batch job contains summary information about the processing request.
| Field | Description |
| ---------------- | --------------------------------------------------------- |
| `job_id` | Unique identifier for the batch job. |
| `status` | Current status of the batch job. |
| `accepted_urls` | Number of valid URLs accepted when the batch was created. |
| `completed_urls` | Number of URLs processed successfully. |
| `failed_urls` | Number of URLs that failed during processing. |
| `config` | Configuration used when the batch job was submitted. |
| `created_at` | Time the batch job was created. |
| `completed_at` | Time the batch job finished processing. |
The response also includes pagination information.
| Field | Description |
| ------------ | --------------------------------- |
| `page` | Current page number. |
| `page_size` | Number of jobs returned per page. |
| `page_count` | Total number of available pages. |
## Filtering Batch Jobs
You can filter the returned jobs using query parameters.
| Query Parameter | Description |
| --------------- | ---------------------------------------------------- |
| `status` | Return only jobs with a specific status. |
| `start_date` | Return jobs created on or after the specified date. |
| `end_date` | Return jobs created on or before the specified date. |
| `page` | Page number to retrieve. |
| `page_size` | Number of jobs to return per page. |
### Filter by Status
Retrieve only completed batch jobs.
```bash title="request.sh"
curl -X GET "https://scraper.geonode.io/v1/batch/jobs?status=completed" \
-H "X-Api-Key: YOUR_API_KEY"
```
Supported values:
```text
queued
processing
completed
failed
cancelled
```
### Filter by Date Range
Retrieve jobs created within a specific time period.
```bash title="request.sh"
curl -X GET "https://scraper.geonode.io/v1/batch/jobs?start_date=2026-07-01&end_date=2026-07-31" \
-H "X-Api-Key: YOUR_API_KEY"
```
### Pagination
Retrieve a specific page of results.
```bash title="request.sh"
curl -X GET "https://scraper.geonode.io/v1/batch/jobs?page=2&page_size=10" \
-H "X-Api-Key: YOUR_API_KEY"
```
## Understanding Job Progress
The job summary makes it easy to see how a batch performed without retrieving the full job details.
| Scenario | Accepted URLs | Completed URLs | Failed URLs |
| ------------------------------- | ------------: | -------------: | ----------: |
| All URLs processed successfully | 3 | 3 | 0 |
| One URL failed | 3 | 2 | 1 |
| Two URLs failed | 5 | 3 | 2 |
## When to Use This Endpoint
Use this endpoint to:
* View recent batch activity.
* Find a previous batch job.
* Monitor completed or failed batches.
* Search jobs created within a specific date range.
* Retrieve a job ID before checking detailed status.
* Build dashboards or reporting tools.
## Next Steps
Now that you know how to retrieve and filter batch jobs, continue to **Cancelling Batch Jobs** to learn how to stop a running batch job.
# Batch Limits (/docs/scraper-api/guides/batch/07_batch_limits)
Before creating a batch job, it's important to understand the limits and validation rules enforced by the API.
Following these guidelines helps prevent validation errors and ensures your requests are accepted successfully.
## Batch Request Limits
The following limits apply when creating a batch job.
| Constraint | Details |
| -------------- | ----------------------------------------------------------------------------------------------------------------------- |
| Required field | The `urls` field is required. Omitting it returns a `422 Unprocessable Entity` error. |
| Minimum URLs | A batch request must contain at least **1 URL**. An empty array (`[]`) is not allowed. |
| Maximum URLs | A single batch request can contain up to **1000 URLs**. |
| Duplicate URLs | Duplicate URLs are accepted and processed as separate requests. |
| Processing | Batch requests are processed asynchronously. A successful request returns a `job_id` instead of the extraction results. |
## Required URLs Field
The `urls` field must always be included in the request body.
### Invalid Request
```json
{}
```
### Response
```json
{
"detail": [
{
"type": "missing",
"loc": [
"body",
"urls"
],
"msg": "Field required"
}
]
}
```
## Empty URL List
Providing an empty array is also considered an invalid request.
### Invalid Request
```json
{
"urls": []
}
```
### Response
```json
{
"detail": [
{
"type": "too_short",
"loc": [
"body",
"urls"
],
"msg": "List should have at least 1 item after validation, not 0"
}
]
}
```
## Maximum Batch Size
A single batch request can contain up to **1000 URLs**.
If you need to process more than 1000 URLs, split them into multiple batch requests.
## Duplicate URLs
Duplicate URLs are allowed.
Each URL in the request is treated as a separate extraction request, even if the same URL appears multiple times.
For example:
```json
{
"urls": [
"https://example.com",
"https://example.com",
"https://example.com"
]
}
```
All three URLs are accepted and processed independently.
## Asynchronous Processing
Creating a batch job does not immediately return the extracted content.
Instead, the API returns a `job_id` that can be used to monitor the job until it completes.
```json
{
"job_id": "9e60e068-b6ac-408d-aa6e-f120f41faf0d",
"status": "queued",
"status_url": "/v1/batch/9e60e068-b6ac-408d-aa6e-f120f41faf0d",
"accepted_urls": 3,
"invalid_urls": []
}
```
Batch processing is asynchronous. After creating a batch job, use the returned `job_id` to check the job status and retrieve the results. The complete workflow is covered in the next guide.
## Next Steps
Continue to **Batch Processing Workflows** to learn how to create a batch job, monitor its progress, and retrieve the final results.
# Understanding Crawl (/docs/scraper-api/guides/crawl/00_overview)
The Crawl API collects content from multiple pages across a website, starting from a single URL. It automatically discovers linked pages, extracts their content, and processes them as an asynchronous crawl job.
***
## What is Crawl?
Crawl is a website traversal and content extraction service. Starting from a single **seed URL**, it visits the page, discovers additional links, and extracts content from every eligible page it crawls.
Unlike extracting a single webpage, Crawl automatically continues exploring connected pages until the crawl is complete.
**Good to know**
Crawl combines **link discovery** and **content extraction** into a single automated workflow.
***
## At a Glance
* 🌐 Start from a single seed URL.
* 🔍 Discover linked pages automatically.
* 📄 Extract content from each page.
* ⚙️ Process pages as an asynchronous crawl job.
* 📦 Retrieve crawl results when processing completes.
***
## How Crawl Works
Every crawl starts from a single **seed URL**. As each page is processed, the crawler extracts content, discovers new links, and continues exploring eligible pages until the crawl is complete.
The crawl follows a simple cycle:
1. Start from a seed URL.
2. Visit and process the page.
3. Extract the page content.
4. Discover additional links.
5. Continue crawling eligible pages.
6. Finish when the crawl completes.
***
## Crawl Workflow
From your application's perspective, every crawl follows the same lifecycle.
Once a crawl job is created, pages are processed in the background. You can then monitor the job and retrieve its results when processing finishes.
***
## Crawl vs. Map vs. Extraction
Each Scraper API endpoint is designed for a different task.
| Feature | Extraction | Map | Crawl |
| -------------------------- | ---------- | --- | ----- |
| Extract page content | ✅ | ❌ | ✅ |
| Discover URLs | ❌ | ✅ | ✅ |
| Follow links automatically | ❌ | ❌ | ✅ |
| Process multiple pages | ❌ | ❌ | ✅ |
### When should you use each endpoint?
| Use Case | Recommended Endpoint |
| ------------------------------------- | -------------------- |
| Extract content from one page | **Extraction** |
| Discover URLs on a website | **Map** |
| Collect content across multiple pages | **Crawl** |
***
## Common Use Cases
Crawl is commonly used for:
* Documentation websites
* Knowledge bases
* Blogs and news sites
* Product catalogs
* Company websites
* AI and RAG data collection
***
## Best Practices
To build efficient crawls:
* Start with a small crawl before scaling up.
* Crawl only the sections you need.
* Review crawl results before downstream processing.
* Configure crawl requests based on your use case.
***
## Next Steps
Now that you understand how Crawl works, you're ready to create your first crawl job.
Continue to **Your First Crawl** to learn how to submit your first crawl request and understand the initial API response.
# Your First Crawl (/docs/scraper-api/guides/crawl/01_first-crawl)
In this guide, you'll create your first crawl job using the Crawl API.
The Crawl API accepts a starting URL and returns a crawl job that runs asynchronously. Once the job is created, you can monitor its progress and retrieve the results using the job ID.
***
## Before You Begin
Before creating a crawl job, make sure you have:
* A GeoNode API key.
* The Scraper API base URL.
If you haven't completed the initial setup, see **Before You Start**.
***
## Create a Crawl Job
Create a crawl by sending a `POST` request to the following endpoint.
```http
POST /v1/crawl
```
For your first crawl, you only need a starting URL and can optionally specify the output format and page limit.
### Example Request
```bash
curl --request POST \
--url https://api.geonode.com/v1/crawl \
--header "Authorization: Bearer " \
--header "Content-Type: application/json" \
--data '{
"url": "https://docs.geonode.com/",
"formats": [
"markdown"
],
"limit": 5
}'
```
You can also send the request body as JSON.
```json
{
"url": "https://docs.geonode.com/",
"formats": [
"markdown"
],
"limit": 5
}
```
***
## Example Response
If the request is accepted, the API returns a `202 Accepted` response similar to the following.
```json
{
"job_id": "8c8f8ad4-a0a0-46f8-92d5-253025c6e19f",
"url": "https://docs.geonode.com/",
"status": "queued",
"status_url": "/v1/crawl/8c8f8ad4-a0a0-46f8-92d5-253025c6e19f",
"estimated_pages": 5
}
```
***
## Understanding the Response
| Field | Description |
| ----------------- | ------------------------------------------- |
| `job_id` | Unique identifier for the crawl job. |
| `url` | The seed URL used to start the crawl. |
| `status` | Current status of the crawl job. |
| `status_url` | Endpoint for checking the crawl job status. |
| `estimated_pages` | Estimated number of pages to be processed. |
The Crawl API processes requests asynchronously. A successful request creates a crawl job rather than returning the crawl results immediately.
***
## What Happens Next?
After your crawl job is created, you can use the returned job ID to monitor its progress.
The next guide explains how to configure crawl requests with additional options, while **Managing Crawl Jobs** covers how to monitor job progress and retrieve results.
***
## Next Steps
Continue to **Configuring Crawl Requests** to learn about all available request options, including crawl depth, page limits, output formats, JavaScript rendering, proxy settings, and browser wait configuration.
# Configuring Crawl Requests (/docs/scraper-api/guides/crawl/02_configuring-crawl-requests)
The Crawl API provides several request parameters that let you control how websites are crawled. This guide explains each parameter and shows how to configure common crawl scenarios.
***
## Request Body
Every crawl request requires a starting URL. All other parameters are optional.
```json
{
"url": "https://docs.geonode.com/"
}
```
***
## Required Parameter
### `url`
The starting URL for the crawl.
| Property | Value |
| -------- | ------ |
| Type | String |
| Required | ✅ Yes |
Example:
```json
{
"url": "https://docs.geonode.com/"
}
```
***
## Crawl Depth
The `depth` parameter controls how many levels deep the crawler follows links from the starting URL.
| Property | Value |
| -------- | ------- |
| Type | Integer |
| Required | No |
| Default | `2` |
Example:
```json
{
"url": "https://docs.geonode.com/",
"depth": 2
}
```
The maximum crawl depth available depends on your subscription plan.
If the requested depth exceeds your plan limit, the API returns a validation error.
```json
{
"code": "VALIDATION_ERROR",
"message": "Crawl depth limit exceeded: requested 3, plan allows at most 2",
"correlation_id": "5d2012da-9bf9-4224-978b-b4337ab7026f",
"retryable": false
}
```
***
## Page Limit
The `limit` parameter specifies the maximum number of pages to crawl.
| Property | Value |
| -------- | ------- |
| Type | Integer |
| Required | No |
Example:
```json
{
"url": "https://docs.geonode.com/",
"limit": 5
}
```
***
## Domain Scope
### `same_domain_only`
Limits crawling to pages within the same domain as the starting URL.
| Property | Value |
| -------- | ------- |
| Type | Boolean |
| Required | No |
Example:
```json
{
"url": "https://docs.geonode.com/",
"same_domain_only": true
}
```
***
### `include_subdomains`
Controls whether subdomains are included during the crawl.
| Property | Value |
| -------- | ------- |
| Type | Boolean |
| Required | No |
Example:
```json
{
"url": "https://docs.geonode.com/",
"include_subdomains": false
}
```
***
## Additional Request Options
The following parameters can be combined to further customize crawl requests.
| Parameter | Description |
| ------------- | --------------------------------------------------- |
| `formats` | Specifies the output format for crawl results. |
| `render_js` | Enables JavaScript rendering for dynamic websites. |
| `proxy` | Configures the proxy used during crawling. |
| `wait_config` | Controls browser wait behavior before page capture. |
Example:
```json
{
"url": "https://docs.geonode.com/",
"formats": [
"markdown"
],
"render_js": true,
"same_domain_only": true,
"include_subdomains": false,
"proxy": {
"country": "us"
},
"wait_config": {
"wait_for": 3000
}
}
```
> For detailed information about these parameters, see the dedicated guides for output formats, JavaScript rendering, proxy configuration, and browser wait configuration.
***
## Complete Example
The following request combines multiple crawl configuration options.
```json
{
"url": "https://docs.geonode.com/",
"depth": 2,
"limit": 5,
"formats": [
"markdown"
],
"render_js": true,
"same_domain_only": true,
"include_subdomains": false,
"proxy": {
"country": "us"
},
"wait_config": {
"wait_for": 3000
}
}
```
If the request is accepted, the API returns a response similar to the following:
```json
{
"job_id": "7344caad-297d-4eac-99ac-198a8c8e6233",
"url": "https://docs.geonode.com/",
"status": "queued",
"status_url": "/v1/crawl/7344caad-297d-4eac-99ac-198a8c8e6233",
"estimated_pages": 5
}
```
Creating a crawl returns a job immediately. Use the returned job\_id or status\_url to monitor progress and retrieve results.
***
## Next Steps
Continue to **Understanding Crawl Results** to learn how crawl results are structured and how to interpret the returned data.
# Understanding Crawl Results (/docs/scraper-api/guides/crawl/03_understanding-crawl-results)
After creating a crawl job, use the job ID to retrieve its current status and results.
The response includes information about the crawl job, the configuration used, crawl statistics, and the extracted content for each crawled page.
***
## Get Crawl Results
Retrieve the current status and results of a crawl job.
```http
GET /v1/crawl/{job_id}
```
Once the crawl completes, the response includes the extracted pages and their associated metadata.
B["Crawl Configuration"]
A --> C["Statistics"]
A --> D["Token Summary"]
A --> E["Results"]
E --> F["Page 1"]
E --> G["Page 2"]
E --> H["Page N"]
`}
/>
***
## Job Information
The top-level response provides general information about the crawl job.
| Field | Description |
| -------------- | ------------------------------------ |
| `job_id` | Unique identifier for the crawl job. |
| `url` | Starting URL used for the crawl. |
| `status` | Current status of the crawl job. |
| `created_at` | Time the crawl job was created. |
| `completed_at` | Time the crawl job completed. |
Example:
```json
{
"job_id": "9293f045-0523-409e-a4e9-874abc663a9e",
"url": "https://docs.geonode.com/",
"status": "completed",
"created_at": "2026-07-23T16:55:51.603677Z",
"completed_at": "2026-07-23T16:56:06.735093Z"
}
```
***
## Crawl Configuration
The `crawl_config` object shows the configuration that was used when the crawl job was created.
| Field | Description |
| -------------------- | -------------------------------------------------------------- |
| `render_js` | Indicates whether JavaScript rendering was enabled. |
| `formats` | Output formats requested for extracted content. |
| `same_domain_only` | Indicates whether crawling was limited to the starting domain. |
| `include_subdomains` | Indicates whether subdomains were included. |
| `proxy` | Proxy configuration used during the crawl. |
| `wait_config` | Browser wait configuration, if specified. |
Example:
```json
{
"crawl_config": {
"render_js": false,
"formats": [
"markdown"
],
"same_domain_only": true,
"include_subdomains": false,
"proxy": {
"country": null,
"type": "residential"
},
"wait_config": null
}
}
```
The returned configuration reflects the settings used for the crawl job.
***
## Crawl Statistics
The response includes statistics that summarize the crawl.
| Field | Description |
| ----------------- | -------------------------------------------- |
| `total_pages` | Total number of pages included in the crawl. |
| `completed_pages` | Number of pages successfully processed. |
| `failed_pages` | Number of pages that failed to process. |
| `cancelled_pages` | Number of pages that were cancelled. |
Example:
```json
{
"total_pages": 5,
"completed_pages": 5,
"failed_pages": 0,
"cancelled_pages": 0
}
```
***
## Token Summary
The `token_summary` object provides information about token usage for the crawl job.
| Field | Description |
| ---------------------- | ----------------------------------- |
| `tokens_charged_total` | Total tokens charged for the crawl. |
| `tokens_reserved` | Tokens reserved for the crawl job. |
Example:
```json
{
"token_summary": {
"tokens_charged_total": 5,
"tokens_reserved": 0
}
}
```
***
## Understanding the Results Array
The `results` array contains one object for each crawled page.
B["Page Result"]
B --> C["Page Information"]
B --> D["Extracted Data"]
B --> E["Metadata"]
B --> F["Links"]
`}
/>
***
## Page Information
Each object in the `results` array describes a single crawled page.
| Field | Description |
| --------------- | ----------------------------------- |
| `url` | URL of the crawled page. |
| `parent_url` | URL where this page was discovered. |
| `depth` | Crawl depth of the page. |
| `status` | Crawl status for the page. |
| `error_code` | Error code if the page failed. |
| `error_message` | Error details if the page failed. |
Example:
```json
{
"url": "https://docs.geonode.com/",
"parent_url": null,
"depth": 0,
"status": "completed",
"error_code": null,
"error_message": null
}
```
***
## Extracted Data
The `data` object contains the extracted content for the page.
Depending on the requested output formats, it may include Markdown, HTML, or both.
Example:
```json
{
"data": {
"markdown": "# Welcome to Geonode\n\nFind exactly what you need...",
"html": null
}
}
```
***
## Page Metadata
Each page includes metadata describing the crawl operation.
| Field | Description |
| ---------------- | ------------------------------------------- |
| `http_status` | HTTP response status received for the page. |
| `duration_ms` | Time taken to process the page. |
| `tokens_charged` | Tokens charged for processing the page. |
Example:
```json
{
"metadata": {
"http_status": 200,
"duration_ms": 957,
"tokens_charged": 1
}
}
```
***
## Discovered Links
The `links` array contains the URLs discovered on the crawled page.
Example:
```json
{
"links": [
"https://docs.geonode.com/",
"https://docs.geonode.com/docs/api-reference",
"https://docs.geonode.com/docs/guides",
"https://docs.geonode.com/docs/scraper-api",
"https://docs.geonode.com/docs/changelog",
"https://geonode.com/contact"
]
}
```
***
## Complete Response Example
The following example shows a completed crawl job response.
```json
{
"job_id": "9293f045-0523-409e-a4e9-874abc663a9e",
"url": "https://docs.geonode.com/",
"status": "completed",
"crawl_config": {
"render_js": false,
"formats": [
"markdown"
],
"same_domain_only": true,
"include_subdomains": false,
"proxy": {
"country": null,
"type": "residential"
},
"wait_config": null
},
"token_summary": {
"tokens_charged_total": 5,
"tokens_reserved": 0
},
"total_pages": 5,
"completed_pages": 5,
"failed_pages": 0,
"cancelled_pages": 0,
"created_at": "2026-07-23T16:55:51.603677Z",
"completed_at": "2026-07-23T16:56:06.735093Z",
"results": [
{
"url": "https://docs.geonode.com/",
"parent_url": null,
"depth": 0,
"status": "completed"
}
]
}
```
***
## Next Steps
Now that you understand the structure of crawl results, continue to **Managing Crawl Jobs** to learn how to list crawl jobs, monitor their progress, and retrieve specific jobs.
# Managing Crawl Jobs (/docs/scraper-api/guides/crawl/04_managing-crawl-jobs)
Crawl jobs are processed asynchronously. After creating a crawl job, you can retrieve its status, monitor its progress, or list previous crawl jobs.
This guide explains how to work with crawl jobs throughout their lifecycle.
***
## List Crawl Jobs
Retrieve a list of your crawl jobs.
```http
GET /v1/crawl/jobs
```
The response includes your crawl jobs along with their current status, crawl statistics, configuration, and pagination information.
Example:
```json
{
"jobs": [
{
"job_id": "453bd7d7-5355-4d6d-a38e-d9e7eb218c3f",
"url": "string",
"status": "queued",
"total_pages": 0,
"completed_pages": 0,
"failed_pages": 0,
"config": {
"render_js": true,
"formats": [
"markdown"
],
"same_domain_only": true,
"include_subdomains": true,
"proxy": {
"country": "string",
"type": "datacenter"
},
"wait_config": {
"wait_until": "commit",
"wait_for": "string",
"wait_timeout": 30000
}
},
"created_at": "2019-08-24T14:15:22Z",
"completed_at": "2019-08-24T14:15:22Z"
}
],
"page": 0,
"page_size": 0,
"page_count": 0
}
```
***
## Retrieve a Crawl Job
Retrieve the latest status and results for a specific crawl job.
```http
GET /v1/crawl/{job_id}
```
Once the crawl completes, this endpoint also returns the extracted page results.
B["Queued"]
B --> C["Running"]
C --> D["Completed"]
`}
/>
***
## Job Status
Each crawl job includes a `status` field that indicates its current state.
| Status | Description |
| ----------- | -------------------------------------------------------- |
| `queued` | The crawl job has been accepted and is waiting to start. |
| `running` | The crawler is actively processing pages. |
| `completed` | The crawl has finished and the results are available. |
Example:
```json
{
"job_id": "453bd7d7-5355-4d6d-a38e-d9e7eb218c3f",
"status": "queued"
}
```
***
## Monitor Crawl Progress
Each job includes progress information that can be used to track the crawl.
| Field | Description |
| ----------------- | -------------------------------------------- |
| `total_pages` | Total number of pages included in the crawl. |
| `completed_pages` | Number of pages successfully processed. |
| `failed_pages` | Number of pages that failed during crawling. |
Example:
```json
{
"total_pages": 20,
"completed_pages": 15,
"failed_pages": 1
}
```
Retrieve the job periodically to monitor its progress until the status becomes `completed`.
***
## Crawl Configuration
Each listed job includes the configuration used when the crawl was created.
| Field | Description |
| -------------------- | -------------------------------------------------------------- |
| `render_js` | Indicates whether JavaScript rendering was enabled. |
| `formats` | Requested output formats. |
| `same_domain_only` | Indicates whether crawling was limited to the starting domain. |
| `include_subdomains` | Indicates whether subdomains were included. |
| `proxy` | Proxy configuration used for the crawl. |
| `wait_config` | Browser wait configuration used during extraction. |
***
## Pagination
The response includes pagination information for the job list.
| Field | Description |
| ------------ | --------------------------------- |
| `page` | Current page number. |
| `page_size` | Number of jobs returned per page. |
| `page_count` | Total number of available pages. |
Example:
```json
{
"page": 0,
"page_size": 10,
"page_count": 3
}
```
***
## Typical Workflow
B
B --> C
C --> D
D --> E
E --> F
`}
/>
***
## Next Steps
If you no longer need an active crawl job, continue to **Cancelling Crawl Jobs** to learn how to stop a running crawl.
# Cancelling Crawl Jobs (/docs/scraper-api/guides/crawl/05_cancelling-crawl-jobs)
If you no longer need a crawl job to continue processing, you can cancel it using its job ID.
Cancelling a crawl job stops scheduling new crawl pages while allowing any pages already being processed to finish.
***
## Cancel a Crawl Job
Cancel an existing crawl job.
```http
DELETE /v1/crawl/{job_id}
```
If the cancellation request is accepted, the API returns information about the crawl job, including its current status, a URL for checking job status, and the number of pages that are still being processed.
Example:
```json
{
"job_id": "453bd7d7-5355-4d6d-a38e-d9e7eb218c3f",
"status": "queued",
"status_url": "string",
"in_flight_pages": 0
}
```
***
## Response Fields
| Field | Description |
| ----------------- | ---------------------------------------------------------------------------------------------- |
| `job_id` | Unique identifier of the crawl job. |
| `status` | Current status of the crawl job. |
| `status_url` | Endpoint that can be used to retrieve the latest status of the crawl job. |
| `in_flight_pages` | Number of queued or processing pages that are still completing after the cancellation request. |
***
## What Happens After Cancellation?
B
B --> C
C --> D
D --> E
`}
/>
After the cancellation request is accepted, you can continue monitoring the crawl job by retrieving its latest status using the `status_url` or the job ID.
***
## Unable to Cancel a Job
If a crawl job cannot be cancelled in its current state, the API returns a `409 Conflict` response.
For example, attempting to cancel a job that has already completed returns:
```json
{
"code": "INVALID_STATE",
"message": "Crawl job 9293f045-0523-409e-a4e9-874abc663a9e is already completed and cannot be cancelled",
"retryable": false
}
```
### Error Fields
| Field | Description |
| ----------- | -------------------------------------------------------- |
| `code` | Machine-readable error code. |
| `message` | Description of why the cancellation request failed. |
| `retryable` | Indicates whether retrying the same request may succeed. |
***
## Possible Responses
| Status Code | Description |
| ---------------------- | ------------------------------------------------------- |
| `202 Accepted` | Cancellation request accepted successfully. |
| `401 Unauthorized` | Authentication failed. |
| `404 Not Found` | The specified crawl job was not found. |
| `409 Conflict` | The crawl job cannot be cancelled in its current state. |
| `422 Validation Error` | The request contains invalid parameters. |
***
# Handling Crawl API Errors (/docs/scraper-api/guides/crawl/06_error-handling)
The Crawl API uses standard HTTP status codes together with structured JSON error responses.
The exact response depends on the endpoint and the reason the request could not be completed.
***
## Standard Error Response
Most Crawl endpoints return the following error format:
```json
{
"code": "string",
"message": "string",
"correlation_id": "string",
"retryable": false
}
```
### Fields
| Field | Description |
| ---------------- | --------------------------------------------------- |
| `code` | Machine-readable error code. |
| `message` | Human-readable description of the error. |
| `correlation_id` | Request correlation identifier when available. |
| `retryable` | Indicates whether retrying the request may succeed. |
***
## HTTP Status Codes
### `POST /v1/crawl`
| Status Code | Description |
| ------------------------- | ----------------------------------- |
| `202 Accepted` | Crawl job accepted successfully. |
| `401 Unauthorized` | Authentication failed. |
| `402 Payment Required` | Insufficient balance. |
| `422 Validation Error` | The request failed validation. |
| `429 Too Many Requests` | Request rate limit exceeded. |
| `503 Service Unavailable` | Service is temporarily unavailable. |
***
### `GET /v1/crawl/jobs`
| Status Code | Description |
| ---------------------- | ---------------------------------- |
| `200 OK` | Crawl jobs retrieved successfully. |
| `401 Unauthorized` | Authentication failed. |
| `422 Validation Error` | Invalid query parameters. |
***
### `GET /v1/crawl/{job_id}`
| Status Code | Description |
| ---------------------- | --------------------------------- |
| `200 OK` | Crawl job retrieved successfully. |
| `401 Unauthorized` | Authentication failed. |
| `404 Not Found` | Crawl job was not found. |
| `422 Validation Error` | Invalid request parameters. |
***
### `DELETE /v1/crawl/{job_id}`
| Status Code | Description |
| ---------------------- | ------------------------------------------------------- |
| `202 Accepted` | Cancellation request accepted. |
| `401 Unauthorized` | Authentication failed. |
| `404 Not Found` | Crawl job was not found. |
| `409 Conflict` | The crawl job cannot be cancelled in its current state. |
| `422 Validation Error` | Invalid request parameters. |
***
## Validation Errors
Requests that fail validation return an `HTTPValidationError` response.
The response contains a `detail` array describing one or more validation issues.
```json
{
"detail": [
{
"loc": [
"string"
],
"msg": "string",
"type": "string"
}
]
}
```
### Validation Fields
| Field | Description |
| ------ | ------------------------------------ |
| `loc` | Location of the validation error. |
| `msg` | Description of the validation error. |
| `type` | Validation error type. |
***
## Example: Invalid Job State
If a crawl job cannot be cancelled because of its current state, the API returns a `409 Conflict` response.
Example:
```json
{
"code": "INVALID_STATE",
"message": "Crawl job 9293f045-0523-409e-a4e9-874abc663a9e is already completed and cannot be cancelled",
"retryable": false
}
```
### Error Fields
| Field | Description |
| ----------- | --------------------------------------------------- |
| `code` | Machine-readable error code. |
| `message` | Explains why the request could not be completed. |
| `retryable` | Indicates whether retrying the request may succeed. |
***
## Summary
The Crawl API returns structured error responses together with standard HTTP status codes. Validation errors include additional details about invalid request parameters, while endpoint-specific errors provide information about why an operation could not be completed. Reviewing both the HTTP status code and the response body can help identify and resolve issues more efficiently.
# Understanding Extraction (/docs/scraper-api/guides/extraction/01_understanding_extraction)
import { Step, Steps } from "fumadocs-ui/components/steps";
import { Tabs, Tab } from "fumadocs-ui/components/tabs";
import { FileText, CodeXml } from "lucide-react";
The Extraction API converts webpages into clean, structured content that can be consumed by applications, AI systems, search pipelines, and automation workflows.
Instead of downloading a webpage and manually parsing raw HTML, you can send a URL and receive the extracted content in a format that is easier to process.
## How It Works
#### Submit a URL
Send the URL of the webpage you want to extract.
#### The Page Is Processed
Geonode fetches the webpage and extracts the primary content.
#### Output Is Generated
The extracted content is returned in one or more supported output formats.
#### Use the Result
Store, analyze, search, or process the extracted content in your application.
## Extraction Endpoints
The Extraction API consists of three endpoints.
| Endpoint | Purpose |
| -------------------------- | ----------------------------------------------------------- |
| `POST /v1/extract` | Extract Markdown and/or HTML from a webpage. |
| `GET /v1/extract/jobs` | List and filter previous extraction jobs. |
| `GET /v1/extract/{job_id}` | Retrieve the status or result of a specific extraction job. |
Most extraction workflows begin with `POST /v1/extract`. The remaining endpoints are primarily used to monitor and retrieve asynchronous extraction jobs.
## Processing Modes
The Extraction API supports both synchronous and asynchronous processing.
The request remains open until extraction is complete.
The extracted content is returned directly in the response.
The request immediately returns a `job_id`.
The extraction continues in the background and the result can be retrieved later using `GET /v1/extract/{job_id}`.
## Extraction Workflow
### Synchronous
```text
POST /v1/extract
↓
Extraction completes
↓
Content returned
```
### Asynchronous
```text
POST /v1/extract
↓
job_id returned
↓
GET /v1/extract/{job_id}
↓
Content returned
```
## Output Formats
The Extraction API supports Markdown, HTML, or both formats in a single request.
Markdown
HTML
Both
The extracted content is returned in the `data.markdown` field.
```json
{
"formats": ["markdown"]
}
```
Markdown returns the extracted content as plain text with lightweight formatting.
The extracted content is returned in the `data.html` field.
```json
{
"formats": ["html"]
}
```
HTML returns the extracted content with a structure closer to the original webpage.
The extracted content is returned in both the `data.markdown` and `data.html` fields.
```json
{
"formats": ["markdown", "html"]
}
```
Both Markdown and HTML are returned in the same response.
## Next Steps
Continue to **Your First Extraction** to send your first extraction request and retrieve content from a webpage.
# Your First Extraction (/docs/scraper-api/guides/extraction/02_your_first_extraction)
In this guide, you'll send your first extraction request and retrieve the content of a webpage.
## What You'll Build
By the end of this guide, you'll be able to:
* Send an extraction request
* Extract content from a webpage
* Understand the response structure
* Access the extracted content
## Send Your First Request
Use the following request:
```bash title="request.sh"
curl -X POST "https://scraper.geonode.io/v1/extract" \
-H "X-Api-Key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "https://example.com"
}'
```
## Request Breakdown
| Field | Description |
| ----- | ------------------------------------ |
| `url` | The webpage to extract content from. |
Since no `formats` field is provided, the API returns HTML by default.
## Understanding the Response
A successful request returns the extracted content and metadata.
```json title="request.json"
{
"data": {
"html": "..."
},
"metadata": {
"url": "http://example.com/",
"render_js": false,
"http_status": 200,
"formats": ["html"],
"processing_mode": "sync"
},
"tokens_charged": 1
}
```
## Access the Extracted Content
The extracted page content is available in:
```text
data.html
```
The metadata section contains additional information about the extraction, including:
* Target URL
* HTTP status
* Output format
* Processing mode
* Extraction duration
## Success
If you received a response similar to the example above, your first extraction was successful.
Your API key is working, the Extraction API is accessible, and you're ready to start working with different output formats and extraction options.
## Next Steps
Continue to **Working With Output Formats** to learn how to return Markdown, HTML, or both formats in a single request.
# Link Extraction (/docs/scraper-api/guides/extraction/03_extracting_links)
By default, the Extraction API returns the extracted page content.
If you also need links found on the page, enable link extraction using the `extract_links` option.
## Enable Link Extraction
Set `extract_links` to `true` in the extraction request.
### Request
```bash title="request.sh"
curl -X POST "https://scraper.geonode.io/v1/extract" \
-H "X-Api-Key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "https://quotes.toscrape.com/",
"formats": ["markdown"],
"extract_links": true
}'
```
## Response
When link extraction is enabled, the response includes a `links` field inside `data`.
```json title="response.json"
{
"data": {
"markdown": "...",
"html": null,
"links": [
"https://quotes.toscrape.com/login",
"https://quotes.toscrape.com/author/Albert-Einstein",
"https://www.goodreads.com/quotes"
]
}
}
```
## Access Extracted Links
The extracted links are available in:
```text title="response-path.txt"
data.links
```
Each item in the array contains a URL discovered on the extracted page.
## When to Use Link Extraction
Enable `extract_links` when you need to:
* Collect links from a webpage
* Discover related pages referenced by the content
* Build URL lists for further processing
* Analyze page relationships
## Important
`extract_links` returns links found on the extracted page.
It does not crawl those links or recursively discover additional pages.
For example:
```text
Page A
├─ Link B
├─ Link C
└─ Link D
```
The API returns links B, C, and D.
It does not visit those pages automatically.
{/* ## Link Extraction vs Map API
| Task | Recommended API |
|--------|--------|
| Extract content and links from a page | Extraction API |
| Discover URLs across a website | Map API |
| Crawl multiple pages | Map API |
| Build a site inventory | Map API |
Use the Map API when you need large-scale URL discovery before deciding which pages to extract. */}
## Success
You now know how to return links alongside extracted content in a single extraction request.
## Next Steps
Continue to **Extraction Workflows** to learn how different extraction options can be combined in real-world scenarios.
# Checking Job Status (/docs/scraper-api/guides/extraction/04_checking_job_status)
When you run an extraction in asynchronous mode, the API immediately returns a job ID instead of waiting for the extraction to finish.
You can use that job ID to check the current status of the extraction and retrieve the result once processing is complete.
## How It Works
The asynchronous extraction workflow follows these steps:
1. Start an extraction using `processing_mode: "async"`.
2. Receive a `job_id`.
3. Poll the job status endpoint.
4. Retrieve the extracted content when the job is completed.
## Get an Extraction Job
Use the following endpoint to retrieve the current status of an extraction job.
```http
GET /v1/extract/{job_id}
```
Replace `{job_id}` with the value returned by your extraction request.
## Example Request
```bash
curl -X GET "https://scraper.geonode.io/v1/extract/4844831a-a222-4cac-b5e6-7e3f2dd07b48" \
-H "X-Api-Key: YOUR_API_KEY"
```
## Response While Processing
A job may still be running when you check its status.
During this stage, the extracted content is not yet available.
```json
{
"job_id": "4844831a-a222-4cac-b5e6-7e3f2dd07b48",
"status": "processing",
"created_at": "2026-05-26T10:30:00Z",
"completed_at": null,
"data": null,
"metadata": null,
"error": null,
"tokens_charged": null
}
```
### What This Means
| Field | Description |
| ---------------- | ---------------------------------------------- |
| `status` | Current state of the extraction job |
| `created_at` | Time the job was created |
| `completed_at` | `null` until processing finishes |
| `data` | Extracted content, available after completion |
| `metadata` | Extraction details, available after completion |
| `error` | Error information if the job fails |
| `tokens_charged` | Token usage after processing completes |
## Response After Completion
Once the extraction finishes successfully, the response includes the extracted content and metadata.
```json
{
"job_id": "4844831a-a222-4cac-b5e6-7e3f2dd07b48",
"status": "completed",
"created_at": "2026-05-26T10:30:00Z",
"completed_at": "2026-05-26T10:30:04Z",
"data": {
"markdown": "# Example Page Content"
},
"metadata": {
"url": "https://docs.python.org/3/library/json.html",
"render_js": false,
"http_status": 200,
"duration_ms": 631,
"formats": ["markdown"],
"processing_mode": "async"
},
"error": null,
"tokens_charged": 1
}
```
## Job Status Values
The API can return the following job statuses.
| Status | Description |
| ------------ | ------------------------------------------------- |
| `queued` | The job has been accepted and is waiting to start |
| `processing` | The extraction is currently running |
| `completed` | The extraction finished successfully |
| `failed` | The extraction could not be completed |
| `cancelled` | The job was cancelled before completion |
## Starting an Async Extraction
To use this endpoint, first create an asynchronous extraction job.
```bash
curl -X POST "https://scraper.geonode.io/v1/extract" \
-H "X-Api-Key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "https://docs.python.org/3/library/json.html",
"formats": ["markdown"],
"processing_mode": "async"
}'
```
## Async Job Response
The extraction endpoint returns a job ID that can be used for polling.
```json
{
"job_id": "4844831a-a222-4cac-b5e6-7e3f2dd07b48",
"status": "queued",
"status_url": "/v1/extract/4844831a-a222-4cac-b5e6-7e3f2dd07b48",
"estimated_tokens": 1
}
```
## Polling for Results
A common pattern is to periodically check the job status until the extraction completes.
```text
POST /v1/extract
↓
Receive job_id
↓
GET /v1/extract/{job_id}
↓
status = processing
↓
GET /v1/extract/{job_id}
↓
status = completed
↓
Read extracted content
```
## Next Step
Now that you can monitor individual extraction jobs, the next guide explains how to view and filter multiple extraction jobs using the Jobs endpoint.
# Listing Extraction Jobs (/docs/scraper-api/guides/extraction/05_listing_extraction_jobs)
Use the Jobs endpoint to view extraction jobs associated with your account.
This endpoint is useful when you need to find a previous job, retrieve a job ID, monitor running jobs, or review extraction history.
## List Extraction Jobs
Use the following endpoint to retrieve extraction jobs.
```http
GET /v1/extract/jobs
```
### Request
```bash title="request.sh"
curl -X GET "https://scraper.geonode.io/v1/extract/jobs" \
-H "X-Api-Key: YOUR_API_KEY"
```
### Response
```json title="response.json"
{
"jobs": [
{
"job_id": "0131f784-9037-4fdd-af9e-b8d445fe2d5f",
"status": "completed",
"url": "https://geonode.com/",
"created_at": "2026-06-13T18:14:54.926869Z",
"execution_time": 2297,
"output": [
"html"
]
}
],
"page": 1,
"page_size": 100,
"page_count": 1
}
```
## Job Fields
Each job contains summary information about an extraction request.
| Field | Description |
| ---------------- | ----------------------------------------- |
| `job_id` | Unique identifier for the extraction job |
| `status` | Current status of the extraction |
| `url` | URL that was extracted |
| `created_at` | Time the job was created |
| `execution_time` | Processing time in milliseconds |
| `output` | Output formats returned by the extraction |
## Filter Jobs
### Filter by Status
Retrieve jobs with a specific status.
```bash title="request.sh"
curl -X GET "https://scraper.geonode.io/v1/extract/jobs?status=completed" \
-H "X-Api-Key: YOUR_API_KEY"
```
Supported values:
```text
queued
processing
completed
failed
cancelled
```
### Filter by Output Format
Retrieve jobs that generated a specific output format.
```bash title="request.sh"
curl -X GET "https://scraper.geonode.io/v1/extract/jobs?output=html" \
-H "X-Api-Key: YOUR_API_KEY"
```
Supported values:
```text
html
markdown
```
### Filter by URL
Retrieve jobs created for a specific URL.
```bash title="request.sh"
curl -X GET "https://scraper.geonode.io/v1/extract/jobs?url=https://geonode.com/" \
-H "X-Api-Key: YOUR_API_KEY"
```
### Filter by Job ID
Retrieve a specific job from the results.
```bash title="request.sh"
curl -X GET "https://scraper.geonode.io/v1/extract/jobs?job_id=0131f784-9037-4fdd-af9e-b8d445fe2d5f" \
-H "X-Api-Key: YOUR_API_KEY"
```
### Filter by Date Range
Retrieve jobs created within a specific time period.
```bash title="request.sh"
curl -X GET "https://scraper.geonode.io/v1/extract/jobs?start_date=2026-06-01&end_date=2026-06-30" \
-H "X-Api-Key: YOUR_API_KEY"
```
## Pagination
Use pagination when working with a large number of extraction jobs.
### Request
```bash title="request.sh"
curl -X GET "https://scraper.geonode.io/v1/extract/jobs?page=1&page_size=5" \
-H "X-Api-Key: YOUR_API_KEY"
```
### Pagination Fields
| Field | Description |
| ------------ | -------------------------------- |
| `page` | Current page number |
| `page_size` | Number of jobs returned per page |
| `page_count` | Total number of available pages |
## Retrieve Full Job Details
The Jobs endpoint returns summary information.
To retrieve the complete extraction result, use:
```http
GET /v1/extract/{job_id}
```
Example:
```bash title="request.sh"
curl -X GET "https://scraper.geonode.io/v1/extract/0131f784-9037-4fdd-af9e-b8d445fe2d5f" \
-H "X-Api-Key: YOUR_API_KEY"
```
## Common Use Cases
Use this endpoint to:
* Find a lost job ID
* Review extraction history
* Monitor running jobs
* View completed extractions
* Search jobs by URL
* Retrieve recent extraction activity
## Success
You now know how to list, search, and filter extraction jobs.
## Next Steps
Continue to **Extraction Workflows** to learn how extraction features can be combined in real-world scenarios.
# Extraction Workflows (/docs/scraper-api/guides/extraction/06_extraction_workflows)
The Extraction API supports a variety of options that can be combined depending on your use case.
This guide demonstrates common extraction workflows and the recommended settings for each scenario.
## Standard Website Extraction
Use the default extraction settings when working with traditional websites that do not rely heavily on JavaScript.
### Request
```bash title="request.sh"
curl -X POST "https://scraper.geonode.io/v1/extract" \
-H "X-Api-Key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "https://example.com"
}'
```
### Best For
* Blogs
* News websites
* Documentation sites
* Static webpages
***
## AI and RAG Workflows
Markdown is often the preferred format when content will be processed by AI systems.
### Request
```bash title="request.sh"
curl -X POST "https://scraper.geonode.io/v1/extract" \
-H "X-Api-Key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "https://docs.python.org/3/library/json.html",
"formats": ["markdown"]
}'
```
### Best For
* Vector databases
* RAG pipelines
* Embedding generation
* Knowledge bases
* LLM applications
***
## JavaScript-Powered Websites
Some websites render content in the browser and require JavaScript execution before extraction.
### Request
```bash title="request.sh"
curl -X POST "https://scraper.geonode.io/v1/extract" \
-H "X-Api-Key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "https://geonode.com/",
"render_js": true
}'
```
### Best For
* React applications
* Next.js websites
* Vue applications
* Single-page applications (SPAs)
***
## Geo-Targeted Extraction
Use a proxy configuration when content changes based on visitor location.
### Request
```bash title="request.sh"
curl -X POST "https://scraper.geonode.io/v1/extract" \
-H "X-Api-Key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "https://example.com",
"proxy": {
"country": "DE",
"type": "residential"
}
}'
```
### Best For
* Local search results
* Country-specific pricing
* Regional content
* Localized webpages
***
## Content and Link Extraction
Extract page content and collect links from the same request.
### Request
```bash title="request.sh"
curl -X POST "https://scraper.geonode.io/v1/extract" \
-H "X-Api-Key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "https://quotes.toscrape.com/",
"formats": ["markdown"],
"extract_links": true
}'
```
### Best For
* URL discovery
* Content analysis
* Link collection
* Website research
***
## Large Async Extractions
Use asynchronous processing for pages that may take longer to extract.
### Start the Extraction
```bash title="request.sh"
curl -X POST "https://scraper.geonode.io/v1/extract" \
-H "X-Api-Key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "https://geonode.com/",
"render_js": true,
"processing_mode": "async"
}'
```
### Response
```json title="response.json"
{
"job_id": "4844831a-a222-4cac-b5e6-7e3f2dd07b48",
"status": "queued",
"status_url": "/v1/extract/4844831a-a222-4cac-b5e6-7e3f2dd07b48"
}
```
### Check Job Status
```bash title="request.sh"
curl -X GET "https://scraper.geonode.io/v1/extract/4844831a-a222-4cac-b5e6-7e3f2dd07b48" \
-H "X-Api-Key: YOUR_API_KEY"
```
### Best For
* Large pages
* Slow websites
* Background processing
* High-volume workflows
***
## Recommended Settings
| Scenario | Recommended Configuration |
| ------------------------ | -------------------------- |
| Standard webpage | Default request |
| AI and RAG workflows | `formats: ["markdown"]` |
| JavaScript websites | `render_js: true` |
| Country-specific content | `proxy` |
| Link discovery | `extract_links: true` |
| Long-running extractions | `processing_mode: "async"` |
## Success
You now know how to combine extraction features for common real-world workflows.
## Next Steps
Continue to **Best Practices** to learn how to improve extraction performance, reliability, and efficiency.
# Best Practices (/docs/scraper-api/guides/extraction/07_best_practices)
The following recommendations can help improve extraction results and reduce unnecessary processing.
## Choose the Right Output Format
Use the output format that matches your use case.
| Format | Best For |
| ---------- | ---------------------------------------------------------- |
| `markdown` | AI workflows, RAG pipelines, indexing, and text processing |
| `html` | Preserving page structure and rendering content |
| Both | Applications that require both formats |
Requesting only the formats you need can reduce response size.
## Enable JavaScript Rendering Only When Needed
JavaScript rendering increases extraction time because the page must be rendered before content can be extracted.
Use:
```json title="request.json"
{
"render_js": true
}
```
only for websites that depend on client-side rendering.
Common examples include:
* React
* Next.js
* Vue
* Single-page applications (SPAs)
## Use Asynchronous Processing for Large Workloads
For long-running extractions, use asynchronous processing.
```json title="request.json"
{
"processing_mode": "async"
}
```
This prevents request timeouts and allows your application to continue processing while extraction runs in the background.
## Use Geo-Targeting Only When Required
Proxy routing may increase processing time.
Only specify a country when content differs by location.
```json title="request.json"
{
"proxy": {
"country": "DE",
"type": "residential"
}
}
```
## Reuse Job IDs
When using asynchronous extraction:
1. Create the extraction job once.
2. Store the returned `job_id`.
3. Poll the job status endpoint.
Avoid creating duplicate extraction jobs for the same request.
## Use Link Extraction Only When Needed
```json title="request.json"
{
"extract_links": true
}
```
Enable link extraction only when you need URLs from the page.
This keeps responses smaller and easier to process.
## Monitor Job Status
Before retrieving extraction results, check the job status.
```http
GET /v1/extract/{job_id}
```
Wait until the status becomes:
```text
completed
```
before processing the result.
## Store Extracted Content
If content does not change frequently, consider storing extraction results instead of repeatedly extracting the same page.
This can reduce costs and improve performance.
## Success
You now know the recommended practices for building reliable extraction workflows.
## Next Steps
Continue to **Common Errors** to learn how to troubleshoot common extraction issues.
# Common Errors (/docs/scraper-api/guides/extraction/08_common_errors)
The following issues are commonly encountered when working with the Extraction API.
## Missing API Key
Requests must include the `X-Api-Key` header.
### Incorrect
```bash title="request.sh"
curl -X POST "https://scraper.geonode.io/v1/extract"
```
### Correct
```bash title="request.sh"
curl -X POST "https://scraper.geonode.io/v1/extract" \
-H "X-Api-Key: YOUR_API_KEY"
```
## Invalid URL
The `url` field must contain a valid URL.
### Incorrect
```json title="request.json"
{
"url": "example"
}
```
### Correct
```json title="request.json"
{
"url": "https://example.com"
}
```
## Job Still Processing
When using asynchronous extraction, the result may not be available immediately.
### Response
```json title="response.json"
{
"status": "processing",
"data": null
}
```
Wait until the job status becomes:
```text
completed
```
before attempting to use the extracted content.
## Job Not Found
A job may not exist or may belong to a different account.
### Request
```http
GET /v1/extract/{job_id}
```
Verify that the job ID is correct and that it was created using the same API key.
## Empty or Unexpected Results
Some websites require JavaScript rendering before content becomes available.
Try enabling:
```json title="request.json"
{
"render_js": true
}
```
This is common for:
* React applications
* Next.js websites
* Vue applications
* Single-page applications
## Geo-Targeted Content Is Different Than Expected
Some websites return different content based on location.
Specify a country explicitly:
```json title="request.json"
{
"proxy": {
"country": "US",
"type": "residential"
}
}
```
## Custom Headers Not Applied
Headers must be passed inside the `headers` object.
### Correct
```json title="request.json"
{
"headers": {
"Accept-Language": "en-US,en;q=0.9"
}
}
```
Do not place your Geonode API key inside the `headers` object.
Authentication must use:
```text
X-Api-Key
```
## Link Extraction Does Not Crawl Websites
`extract_links` only returns links found on the extracted page.
```json title="request.json"
{
"extract_links": true
}
```
It does not visit discovered links or recursively crawl a website.
For large-scale URL discovery, use the Map API.
## Using JavaScript Rendering Unnecessarily
JavaScript rendering increases extraction time.
```json title="request.json"
{
"render_js": true
}
```
Enable it only when a website requires client-side rendering.
## Need More Help?
If you continue to experience issues:
* Verify the request payload
* Verify the target URL
* Check job status for asynchronous requests
* Review response metadata
* Contact the Geonode support team
## Success
You now understand the most common Extraction API issues and how to resolve them.
## What's Next
You have completed the Extraction guides.
Continue to the next API section to learn about additional scraping and data collection capabilities.
# Listing Jobs (/docs/scraper-api/guides/jobs/01_listing_jobs)
The Scraper API allows you to retrieve previously created jobs.
Listing jobs is useful for monitoring activity, reviewing completed requests, or locating a specific job ID.
## List Jobs
Use the list jobs endpoint provided by the API.
For example:
```http
GET /v1/extract/jobs
```
or
```http
GET /v1/batch/jobs
```
or
```http
GET /v1/crawl/jobs
```
The response returns a paginated collection of jobs.
## Filter Results
Most job endpoints support filtering and pagination.
Common query parameters include:
| Parameter | Description |
| ------------ | ------------------------------------------- |
| `status` | Return only jobs with a specific status. |
| `start_date` | Return jobs created after a specific date. |
| `end_date` | Return jobs created before a specific date. |
| `page` | Page number to retrieve. |
| `page_size` | Number of jobs returned per page. |
Use these filters to quickly locate the jobs you need.
## Pagination
Large job histories are split into pages.
Increase or decrease `page_size` depending on how many results you want returned in each request.
## Common Use Cases
Listing jobs is useful for:
* Viewing recent activity
* Finding completed jobs
* Monitoring failed jobs
* Reviewing processing history
* Building dashboards or reporting tools
# Checking Job Status (/docs/scraper-api/guides/jobs/02_checking_job_status)
Some Scraper API endpoints process requests asynchronously.
Instead of returning the final result immediately, the API creates a job and returns a unique job ID. You can use that job ID to check the progress of the request.
Common job statuses include:
| Status | Description |
| ------------ | -------------------------------------------------------- |
| `queued` | The job is waiting to be processed. |
| `processing` | The request is currently being processed. |
| `completed` | The job finished successfully and results are available. |
| `failed` | The request could not be completed. |
## Check a Job
Use the endpoint associated with your API to retrieve the latest status of a job.
For example:
```http
GET /v1/extract/jobs/{job_id}
```
or
```http
GET /v1/batch/jobs/{job_id}
```
or
```http
GET /v1/crawl/jobs/{job_id}
```
The response contains the current status and any available results.
## Typical Workflow
```text
Create Request
│
▼
Receive Job ID
│
▼
Check Job Status
│
▼
Completed
│
▼
Read Results
```
Continue checking the job until its status changes to `completed` or `failed`.
## When to Use
Checking job status is useful when:
* Processing large requests
* Crawling multiple pages
* Running batch operations
* Using asynchronous processing modes
# Output Formats (/docs/scraper-api/guides/making-requests/02_output_formats)
import { Tabs, Tab } from "fumadocs-ui/components/tabs";
import { FileText, CodeXml } from "lucide-react";
The Extraction API can return content in Markdown, HTML, or both formats in a single request.
All examples in this guide use the following endpoint:
```http
POST /v1/extract
```
The `formats` field controls which output formats are returned by the extraction request.
```json
{
"formats": ["markdown"]
}
```
If the `formats` field is omitted, the API returns HTML by default.
## Output Formats
Markdown
HTML
Both
Markdown returns the extracted content as clean, readable text.
### Request
```bash title="request.sh"
curl -X POST "https://scraper.geonode.io/v1/extract" \
-H "X-Api-Key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "https://example.com",
"formats": ["markdown"]
}'
```
### Response
```json title="response.json"
{
"data": {
"markdown": "# Example Domain..."
}
}
```
The extracted content is available in the `data.markdown` field.
Common use cases:
* AI and LLM workflows
* Search indexing
* Knowledge bases
* Text processing pipelines
HTML returns the extracted content with a structure closer to the original webpage.
### Request
```bash
curl -X POST "https://scraper.geonode.io/v1/extract" \
-H "X-Api-Key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "https://example.com",
"formats": ["html"]
}'
```
### Response
```json
{
"data": {
"html": "..."
}
}
```
The extracted content is available in the `data.html` field.
Common use cases:
* Rendering content in applications
* Preserving page structure
* Working with HTML elements
* Content transformation workflows
Request both Markdown and HTML when your application needs both formats from the same extraction.
### Request
```bash
curl -X POST "https://scraper.geonode.io/v1/extract" \
-H "X-Api-Key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "https://example.com",
"formats": ["markdown", "html"]
}'
```
### Response
```json
{
"data": {
"markdown": "# Example Domain...",
"html": "..."
}
}
```
The extracted content is available in both the `data.markdown` and `data.html` fields.
## Choosing the Right Format
| Format | Best For |
| -------- | -------------------------------------------------------------------- |
| Markdown | AI workflows, search indexing, knowledge bases, and text processing |
| HTML | Preserving page structure and rendering content |
| Both | Applications that need both representations from a single extraction |
## Success
You now know how to control the format returned by the Extraction API.
Whether you need Markdown, HTML, or both, you can choose the format that best fits your workflow.
## Next Steps
Continue to **Extracting JavaScript Websites** to learn how to extract content from pages that rely on client-side rendering.
# JavaScript Rendering (/docs/scraper-api/guides/making-requests/03_javascript-rendering)
Many modern websites load content after the initial page request using JavaScript.
When this happens, a standard extraction may return incomplete content because the page has not finished rendering.
To handle these websites, enable JavaScript rendering with the `render_js` option.
## When to Use JavaScript Rendering
Enable JavaScript rendering when:
* Important content is missing from the extraction result
* Content appears after the page loads
* The website relies on client-side rendering
## Enable JavaScript Rendering
Set `render_js` to `true` in your extraction request.
```bash title="request.sh"
curl -X POST "https://scraper.geonode.io/v1/extract" \
-H "X-Api-Key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "https://geonode.com/",
"render_js": true
}'
```
### Request Breakdown
| Field | Description |
| ----------- | -------------------------------------------------------- |
| `url` | The webpage to extract content from. |
| `render_js` | Renders the page in a browser before extracting content. |
## Response
A successful response includes the rendered page content and extraction metadata.
```json title="response.json"
{
"data": {
"html": "..."
},
"metadata": {
"url": "https://geonode.com/",
"render_js": true,
"http_status": 200
}
}
```
Notice that `metadata.render_js` is set to `true`, confirming that browser rendering was used during extraction.
## Things to Consider
JavaScript rendering provides more complete extraction results for dynamic websites, but it may:
* Take longer than a standard extraction
* Use additional browser resources
* Be unnecessary for static websites
For static websites, leave `render_js` disabled for faster extraction.
## Success
You can now extract content from websites that rely on JavaScript to display their content.
## Next Steps
Continue to **Processing Modes** to learn the difference between synchronous and asynchronous extraction requests.
# Waiting for Dynamic Content (/docs/scraper-api/guides/making-requests/04_waiting_for_dynamic_content)
Modern websites often load content after the initial page load using JavaScript. If content appears a few seconds later, extracting immediately may return incomplete results.
The `wait_config` parameter gives you control over when the extraction should begin.
## How wait\_config Works
When a `wait_config` is provided, the extraction process follows this order:
```text
wait_until
↓
wait_for
↓
wait_timeout
↓
Extract Content
```
This allows you to wait for page events, specific elements, or additional delays before extraction starts.
## wait\_until
The `wait_until` option controls which browser lifecycle event must complete before moving to the next step.
### commit
Wait until the browser receives the response headers and commits the navigation.
```json
{
"url": "https://example.com",
"wait_config": {
"wait_until": "commit"
}
}
```
Best for:
* Very fast extractions
* Cases where you only need the initial response
### domcontentloaded
Wait until the HTML is parsed and the DOM is ready.
```json
{
"url": "https://example.com",
"wait_config": {
"wait_until": "domcontentloaded"
}
}
```
Best for:
* Most websites
* Pages where content is already present in the HTML
> This is the default value when `wait_until` is not provided.
### load
Wait until the page and all resources are fully loaded.
```json
{
"url": "https://example.com",
"wait_config": {
"wait_until": "load"
}
}
```
Best for:
* Pages that depend on images or external scripts
* Slower websites that require additional loading time
### networkidle
Wait until there is no network activity for 500 milliseconds.
```json
{
"url": "https://example.com",
"wait_config": {
"wait_until": "networkidle"
}
}
```
Best for:
* Single-page applications (SPA)
* React, Vue, Angular, and Next.js websites
* Dynamic content loaded through API requests
## wait\_for
The `wait_for` option waits until a specific element appears on the page before extraction starts.
### Using a CSS Selector
```json
{
"url": "https://example.com",
"wait_config": {
"wait_for": ".product-grid"
}
}
```
Extraction begins only after an element matching `.product-grid` is found.
### Using XPath
```json
{
"url": "https://example.com",
"wait_config": {
"wait_for": "//div[@class='product-grid']"
}
}
```
XPath expressions must start with:
```text
//
```
or
```text
xpath=
```
Otherwise, the value is treated as a CSS selector.
### Common Examples
| Use Case | Selector |
| ------------------ | ----------------- |
| Product listing | `.product-grid` |
| Search results | `.search-results` |
| Article content | `article` |
| Table data | `table` |
| Loading completion | `.loaded` |
## wait\_timeout
The `wait_timeout` option adds an additional delay after all previous waits complete.
```json
{
"url": "https://example.com",
"wait_config": {
"wait_for": ".product-grid",
"wait_timeout": 5000
}
}
```
In this example:
1. Wait for `.product-grid`
2. Wait an additional 5 seconds
3. Extract content
### When to Use wait\_timeout
Use this option when content continues updating after the target element appears.
Common examples include:
* Infinite scrolling pages
* Late-loading advertisements
* Client-side rendering delays
* Dynamic dashboards
## Complete Example
```json
{
"url": "https://example.com",
"formats": ["markdown"],
"wait_config": {
"wait_until": "networkidle",
"wait_for": ".product-grid",
"wait_timeout": 3000
}
}
```
This configuration:
1. Waits until network activity stops
2. Waits for `.product-grid` to appear
3. Waits an additional 3 seconds
4. Extracts the content
## Browser Rendering Behavior
When a non-null `wait_config` is provided, browser rendering is automatically enabled if `render_js` is not explicitly specified.
Example:
```json
{
"url": "https://example.com",
"wait_config": {
"wait_until": "networkidle"
}
}
```
The extraction will automatically use browser rendering.
## Invalid Configuration
The following request is rejected:
```json
{
"url": "https://example.com",
"render_js": false,
"wait_config": {
"wait_until": "networkidle"
}
}
```
This is considered ambiguous because `wait_config` requires browser rendering while `render_js` explicitly disables it.
## Best Practices
* Use `domcontentloaded` for most websites.
* Use `networkidle` for JavaScript-heavy applications.
* Use `wait_for` when a specific element contains the data you need.
* Use `wait_timeout` only when additional rendering time is required.
* Avoid excessive delays, as they increase extraction time and cost.
## Next Steps
Now that you understand how to wait for dynamic content, learn how to work with synchronous and asynchronous extraction workflows.
# Processing Modes (/docs/scraper-api/guides/making-requests/05_processing_modes)
import { Tabs, Tab } from "fumadocs-ui/components/tabs";
The Extraction API supports two processing modes:
* `sync` for immediate results
* `async` for background processing
Use the `processing_mode` field to control how extraction requests are handled.
```json
{
"processing_mode": "sync"
}
```
If `processing_mode` is omitted, the API uses `sync` mode by default.
## Processing Modes
Synchronous mode waits for extraction to complete before returning a response.
### Request
```bash title="request.sh"
curl -X POST "https://scraper.geonode.io/v1/extract" \
-H "X-Api-Key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "https://example.com",
"processing_mode": "sync"
}'
```
### Response
```json title="response.json"
{
"data": {
"html": "...",
"markdown": null
},
"metadata": {
"url": "http://example.com/",
"http_status": 200,
"processing_mode": "sync"
},
"tokens_charged": 1
}
```
The request remains open until extraction is complete and the content is returned in the response.
Asynchronous mode immediately creates an extraction job and returns a job ID.
### Request
```bash title="request.sh"
curl -X POST "https://scraper.geonode.io/v1/extract" \
-H "X-Api-Key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "https://geonode.com/",
"processing_mode": "async"
}'
```
### Response
```json title="response.json"
{
"job_id": "32561cfc-4d87-4a46-af4a-a10e5f3168b9",
"status": "queued",
"status_url": "/v1/extract/32561cfc-4d87-4a46-af4a-a10e5f3168b9",
"estimated_tokens": 1
}
```
The extraction continues in the background while your application continues running.
Use the returned `job_id` or `status_url` to retrieve the extraction result later.
## Choosing a Processing Mode
| Mode | Best For |
| ----- | -------------------------------------------------------------------- |
| Sync | Interactive applications, quick extractions, and immediate results |
| Async | Background processing, large workloads, and long-running extractions |
## Success
You now know how to choose between synchronous and asynchronous extraction requests.
Synchronous mode returns the extracted content immediately, while asynchronous mode returns a job ID that can be used to retrieve the result later.
## Next Steps
Continue to **Checking Job Status** to learn how to retrieve the status and results of an asynchronous extraction job using `GET /v1/extract/{job_id}`.
# Proxy and Geo-Targeting (/docs/scraper-api/guides/making-requests/06_proxy_and_geo_targeting)
The Scraper API can route extraction requests through Geonode proxies.
If you do not provide a `proxy` object, the API uses residential proxies by default and automatically determines the most appropriate routing location when possible.
Geo-targeting is useful when a website returns different content, language, pricing, availability, or search results based on the visitor's location.
## Configure a Proxy
Use the `proxy` object to control the country and proxy type used for the extraction request.
### Request
```bash title="request.sh"
curl -X POST "https://scraper.geonode.io/v1/extract" \
-H "X-Api-Key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "https://example.com",
"proxy": {
"country": "US",
"type": "residential"
}
}'
```
## Country Codes
The `proxy.country` field uses standard ISO 3166-1 alpha-2 country codes.
| Country | Code |
| -------------- | ---- |
| United States | `US` |
| United Kingdom | `GB` |
| Germany | `DE` |
| Pakistan | `PK` |
| Canada | `CA` |
| France | `FR` |
For the complete list of supported country codes:
[https://en.wikipedia.org/wiki/ISO\_3166-1\_alpha-2](https://en.wikipedia.org/wiki/ISO_3166-1_alpha-2)
## Proxy Types
### Residential
Routes requests through residential IP addresses.
```json title="request.json"
{
"proxy": {
"type": "residential"
}
}
```
### Datacenter
Routes requests through datacenter IP addresses.
```json title="request.json"
{
"proxy": {
"type": "datacenter"
}
}
```
### Mix
Allows the API to use a combination of available proxy networks.
```json title="request.json"
{
"proxy": {
"type": "mix"
}
}
```
## Proxy Configuration Reference
| Field | Type | Description |
| --------------- | -------------- | --------------------------------------------------------------------------------------------------- |
| `proxy.country` | string or null | Two-letter ISO country code such as `US`, `GB`, `DE`, or `PK`. |
| `proxy.type` | string | Proxy type. Supported values are `residential`, `datacenter`, and `mix`. Defaults to `residential`. |
## Verify the Applied Proxy
The proxy configuration used for the request is returned in the response metadata.
### Response
```json title="response.json"
{
"data": {
"html": "..."
},
"metadata": {
"url": "http://example.com/",
"proxy": {
"country": "US",
"type": "residential"
},
"processing_mode": "sync"
},
"tokens_charged": 1
}
```
You can inspect `metadata.proxy.country` and `metadata.proxy.type` to verify which proxy configuration was applied during extraction.
## Success
You now know how to control the country and proxy network used for extraction requests.
## Next Steps
Continue to **Using Custom Headers** to learn how to send additional HTTP headers with extraction requests.
# Using Custom Headers (/docs/scraper-api/guides/making-requests/07_using_custom_headers)
Use the `headers` object to send custom HTTP headers with the extraction request.
## Important
```json title="request.json"
{
"url": "https://example.com",
"headers": {
"Accept-Language": "en-US,en;q=0.9"
}
}
```
Do not put your Geonode API key inside `headers`.
Authentication belongs in the `X-Api-Key` request header sent to the Scraper API.
## Send Custom Headers
Use the `headers` object to include one or more HTTP headers in your extraction request.
### Request
```bash title="request.sh"
curl -X POST "https://scraper.geonode.io/v1/extract" \
-H "X-Api-Key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "https://geonode.com/",
"headers": {
"Accept-Language": "en-US,en;q=0.9",
"User-Agent": "Mozilla/5.0"
}
}'
```
## Common Headers
| Header | Purpose |
| ----------------- | ------------------------------------------------- |
| `Accept-Language` | Request content in a specific language |
| `User-Agent` | Identify the browser or client making the request |
| `Referer` | Indicate the page that initiated the request |
| `Cookie` | Send session or authentication cookies |
## Multiple Headers
You can send multiple headers in a single request.
```json title="request.json"
{
"url": "https://geonode.com/",
"headers": {
"Accept-Language": "en-US,en;q=0.9",
"User-Agent": "Mozilla/5.0"
}
}
```
## Success
You now know how to send custom HTTP headers during extraction requests.
## Next Steps
Continue to **Link Extraction** to learn how to return links found on a webpage along with its extracted content.
# Understanding Map (/docs/scraper-api/guides/map/00_understanding_map)
The Map API helps you discover URLs across a website without extracting the content of each page.
Instead of downloading and processing every page, the Map API discovers URLs from a website and returns them as a structured list. This makes it useful for exploring websites, planning scraping workflows, and identifying the pages you want to process later.
***
## How the Map API Works
The Map API starts with a single website URL and discovers additional URLs from the website.
During discovery, the Map API can collect URLs from:
* Website sitemaps.
* Links discovered in HTML pages.
Each discovered URL includes information about how it was found, such as `sitemap` or `html`.
***
## Map vs Crawl vs Extraction
Although these APIs work together, they solve different problems.
| Feature | Map | Crawl | Extraction |
| -------------------------------- | :-: | :---: | :--------: |
| Discover website URLs | ✓ | ✓ | — |
| Extract page content | — | ✓ | ✓ |
| Process a single page | — | — | ✓ |
| Process multiple pages | ✓ | ✓ | — |
| Return a list of discovered URLs | ✓ | — | — |
Use the Map API to discover pages first. Once you've identified the pages you need, use the Extraction API to extract individual pages or the Crawl API to process larger sections of the website.
***
## Common Use Cases
The Map API is useful for:
* Building a list of URLs from a website.
* Discovering documentation, blog, or support pages.
* Preparing URLs before extraction or crawling.
* Exploring the content available on a website.
* Identifying pages for further processing or analysis.
***
## When Should You Use the Map API?
Choose the Map API when your goal is to discover **where content exists**, rather than extracting the content itself.
| Goal | Recommended API |
| ----------------------------------------- | --------------- |
| Discover available pages | Map API |
| Extract content from one page | Extraction API |
| Extract content from many connected pages | Crawl API |
***
## Next Steps
Now that you understand what the Map API does, continue to **Map Workflows** to learn how to use the Map endpoint as part of a complete URL discovery workflow.
# Map Workflows (/docs/scraper-api/guides/map/01_map_workflows)
The Map API is typically the first step in a scraping workflow. It helps you discover URLs from a website so you can decide which pages to process next.
Instead of extracting page content, the Map API returns a list of discovered URLs that can be used with other Scraper API endpoints.
***
## Typical Workflow
Most applications use the Map API as part of the following workflow.
B[POST /v1/map]
B-->C[Review Discovered URLs]
C-->D[Extract Selected Pages]
C-->E[Crawl Website]
C-->F[Export URL List]
`}
/>
The workflow starts with a website URL. After reviewing the discovered URLs, you can decide whether to extract specific pages, crawl larger sections of the website, or export the URL list for further processing.
***
## Step 1 — Submit a Website
Start by sending a request to the Map endpoint.
```http
POST /v1/map
```
Provide the website URL that you want to discover.
The Map API returns a response containing the discovered URLs.
***
## Step 2 — Review the Results
The response includes a list of discovered URLs.
Each URL also indicates how it was discovered, such as:
* `html`
* `sitemap`
Review the results before deciding which pages should be processed further.
***
## Step 3 — Choose Your Next Step
After reviewing the discovered URLs, choose the workflow that best fits your application.
B[Extract Individual Pages]
A-->C[Crawl Multiple Pages]
A-->D[Export URL Inventory]
`}
/>
Each option serves a different purpose:
| Next Step | When to Use |
| -------------- | ---------------------------------------------------------------------- |
| Extraction API | Process individual pages and extract structured content. |
| Crawl API | Process larger sections of a website. |
| Export URLs | Save the discovered URLs for reporting, analysis, or another workflow. |
***
## Example Workflow
The following example shows how the Map API fits into a typical scraping pipeline.
B[Map API]
-->C[Discovered URLs]
-->D[Filter Documentation Pages]
D-->E[Extraction API]
`}
/>
Rather than processing every page, the application first discovers all available URLs, filters the pages it needs, and then extracts content only from the relevant pages.
***
## Building a Processing Pipeline
Many applications use the Map API as the first stage of a larger workflow.
B[Discovered URLs]
B-->C[Filter URLs]
C-->D[Extraction API]
C-->E[Crawl API]
C-->F[Store Results]
`}
/>
Separating URL discovery from content extraction gives you greater control over what your application processes.
***
## Best Practices
* Start with the Map API before extracting or crawling a large website.
* Review the discovered URLs before processing them.
* Use the `search` parameter to narrow the discovered URLs when appropriate.
* Use the Extraction API for individual pages.
* Use the Crawl API when processing larger sections of a website.
* Export or store discovered URLs if they will be reused later.
***
## End-to-End Workflow
| Step | Action | Outcome |
| ---- | ----------------------------------- | ----------------------------- |
| 1 | Submit a website to the Map API | Discover available URLs |
| 2 | Review the discovered URLs | Identify relevant pages |
| 3 | Filter the results if needed | Reduce unnecessary processing |
| 4 | Extract or crawl the selected pages | Process the website content |
***
## Next Steps
Now that you understand how the Map API fits into a complete workflow, continue to **Your First Map** to create your first request and explore the request and response in detail.
# Your First Map (/docs/scraper-api/guides/map/02_your_first_map)
In this guide, you'll use the Map endpoint to discover URLs under a website.
By the end of this guide, you'll know how to:
* Send your first Map request.
* Understand the request body.
* Read the response.
* Work with the discovered URLs.
***
## Before You Begin
Before using the Map API, make sure you have:
* A valid Geonode API key.
* The Scraper API base URL.
* A website URL to map.
The Map API discovers URLs by combining sitemap parsing with HTML link extraction from the provided website.
***
## Step 1 — Send a Map Request
Send a POST request to the Map endpoint.
```http
POST /v1/map
```
The minimum request only requires a website URL.
```json
{
"url": "https://geonode.com"
}
```
B["POST /v1/map"]
-->C["Map API"]
-->D["Discovered URLs"]
`}
/>
***
## Step 2 — Understand the Request
The Map endpoint accepts the following request fields.
| Field | Required | Description |
| ------------------------- | -------- | --------------------------------------------------------------------------- |
| `url` | Yes | Base URL to discover links from. |
| `search` | No | Filters discovered URLs using a case-insensitive substring match. |
| `include_subdomains` | No | Includes common sibling subdomains during discovery. Default: `true`. |
| `ignore_query_parameters` | No | Removes query parameters when normalizing discovered URLs. Default: `true`. |
For example:
```json
{
"url": "https://geonode.com",
"include_subdomains": true,
"ignore_query_parameters": true
}
```
***
## Step 3 — Review the Response
A successful request returns a response similar to:
```json
{
"success": true,
"links": [
{
"url": "https://geonode.com/docs",
"source": "html"
},
{
"url": "https://geonode.com/blog",
"source": "sitemap"
}
]
}
```
The response contains three fields.
| Field | Description |
| --------- | ----------------------------------------------------- |
| `success` | Indicates whether the request completed successfully. |
| `links` | List of discovered URLs. |
| `warning` | Optional non-fatal advisory message. |
***
## Understanding Discovered Links
Each discovered URL contains two values.
| Field | Description |
| -------- | -------------------------------------------------------- |
| `url` | The discovered page URL. |
| `source` | Where the URL was discovered from (`html` or `sitemap`). |
B["success"]
A --> C["links"]
C --> D["url"]
C --> E["source"]
E --> F["html"]
E --> G["sitemap"]
`}
/>
***
## Filtering Results
Use the optional `search` field to return only URLs containing specific text.
For example:
```json
{
"url": "https://geonode.com",
"search": "docs"
}
```
The `search` parameter performs a case-insensitive substring match against the discovered URLs. It does not query a search engine.
***
## Working with the Results
After receiving the response, you can:
* Review the discovered website structure.
* Select specific pages for extraction.
* Pass URLs to the Crawl API.
* Build your own processing workflow.
B["Review"]
A --> C["Extract Pages"]
A --> D["Crawl Website"]
`}
/>
***
## Complete Workflow
B["POST /v1/map"]
-->C["Receive Response"]
-->D["Review Links"]
-->E["Use URLs"]
`}
/>
***
## Next Steps
Now that you've created your first Map request, continue to **Configuring Map Requests** to learn how to customize URL discovery using the available request parameters.
# Configuring Map Requests (/docs/scraper-api/guides/map/03_configuring_map_requests)
The Map API provides several optional request parameters that let you control how URLs are discovered and returned.
This guide explains when to use each option and how it affects the mapping results.
***
## Request Parameters
The Map endpoint accepts the following parameters.
| Field | Required | Description |
| ------------------------- | -------- | --------------------------------------------------------------------------- |
| `url` | Yes | The website to map. |
| `search` | No | Returns only discovered URLs that contain the specified text. |
| `include_subdomains` | No | Includes URLs from supported subdomains during discovery. Default: `true`. |
| `ignore_query_parameters` | No | Removes query parameters when normalizing discovered URLs. Default: `true`. |
***
## Basic Request
The minimum request only requires the website URL.
```json
{
"url": "https://geonode.com"
}
```
B["Map API"]
-->C["Discover URLs"]
`}
/>
***
## Filter Results with `search`
Use the `search` parameter to return only URLs containing a specific keyword.
```json
{
"url": "https://geonode.com",
"search": "docs"
}
```
The search performs a case-insensitive substring match against the discovered URLs.
For example, searching for:
```text
docs
```
may return URLs such as:
```text
https://docs.geonode.com
https://docs.geonode.com/docs/api-reference
```
If no URLs match your search term, the request still succeeds.
Example response:
```json
{
"success": true,
"links": [],
"metadata": {
"url": "https://geonode.com/",
"duration_ms": 4208,
"links_count": 0
},
"warning": "No results found. If you targeted a sub-path, try mapping the base domain for broader coverage."
}
```
An empty `links` array does not indicate an error. It simply means that no discovered URLs matched the search value.
***
## Include Subdomains
By default, the Map API includes supported subdomains during discovery.
You can control this behavior using `include_subdomains`.
```json
{
"url": "https://geonode.com",
"include_subdomains": true
}
```
Set the value to `false` if you only want to map the primary domain.
```json
{
"url": "https://geonode.com",
"include_subdomains": false
}
```
Use this option when:
* Discovering documentation hosted on subdomains.
* Mapping an organization's entire website.
* Limiting discovery to the primary domain only.
***
## Ignore Query Parameters
Many websites generate multiple URLs that differ only by query parameters.
For example:
```text
/products?page=1
/products?page=2
/products?sort=newest
```
When `ignore_query_parameters` is enabled, the Map API normalizes these URLs to reduce duplicates.
```json
{
"url": "https://geonode.com",
"ignore_query_parameters": true
}
```
Disable this option if query parameters represent unique pages that you want to keep.
```json
{
"url": "https://geonode.com",
"ignore_query_parameters": false
}
```
***
## Combining Parameters
You can combine multiple options in the same request.
```json
{
"url": "https://geonode.com",
"search": "docs",
"include_subdomains": true,
"ignore_query_parameters": true
}
```
B["Include Subdomains"]
B --> C["Ignore Query Parameters"]
C --> D["Apply Search Filter"]
D --> E["Return Matching URLs"]
`}
/>
***
## Best Practices
* Always provide the base website URL.
* Use `search` to reduce the number of returned URLs.
* Enable `include_subdomains` when mapping websites that host content across multiple subdomains.
* Leave `ignore_query_parameters` enabled unless query parameters represent unique content.
* Combine parameters to return only the URLs that are relevant to your application.
***
## Next Steps
Now that you know how to configure mapping requests, continue to **Understanding Map Results** to learn how to interpret the response returned by the Map API.
# Understanding Map Results (/docs/scraper-api/guides/map/04_understanding_map_results)
After a successful mapping request, the Map API returns a structured response containing the discovered URLs and information about the mapping operation.
Understanding this response helps you decide which URLs to extract, crawl, or process further.
***
## Response Structure
A successful response contains the following top-level fields.
B["success"]
A --> C["links"]
C --> D["url"]
C --> E["source"]
A --> F["metadata"]
F --> G["url"]
F --> H["duration_ms"]
F --> I["links_count"]
A --> J["warning (optional)"]
`}
/>
***
## Example Response
```json
{
"success": true,
"links": [
{
"url": "https://docs.geonode.com",
"source": "sitemap"
},
{
"url": "https://docs.geonode.com/docs/api-reference",
"source": "sitemap"
}
],
"metadata": {
"url": "https://geonode.com/",
"duration_ms": 0,
"links_count": 202
}
}
```
***
## Response Fields
| Field | Description |
| ---------- | -------------------------------------------------------------------------------- |
| `success` | Indicates whether the mapping request completed successfully. |
| `links` | List of discovered URLs. |
| `metadata` | Information about the completed mapping request. |
| `warning` | Optional message returned when the request succeeds but requires your attention. |
***
## Understanding `success`
The `success` field indicates whether the request completed successfully.
```json
{
"success": true
}
```
A value of `true` means the Map API successfully processed the request and returned a response.
***
## Understanding `links`
The `links` array contains every URL discovered during the mapping process.
Each item contains:
| Field | Description |
| -------- | --------------------------- |
| `url` | The discovered page URL. |
| `source` | How the URL was discovered. |
Example:
```json
{
"url": "https://docs.geonode.com/docs/api-reference",
"source": "sitemap"
}
```
The `source` field helps you understand where each URL originated.
Possible values include:
| Value | Description |
| --------- | -------------------------------------------------------- |
| `sitemap` | The URL was discovered from a website sitemap. |
| `html` | The URL was discovered by following links in HTML pages. |
The same response may contain URLs discovered from both `sitemap` and `html` sources.
***
## Understanding `metadata`
The `metadata` object provides information about the mapping request itself.
```json
{
"metadata": {
"url": "https://geonode.com/",
"duration_ms": 0,
"links_count": 202
}
}
```
| Field | Description |
| ------------- | ------------------------------------------------------------ |
| `url` | The website that was mapped. |
| `duration_ms` | Time taken to complete the mapping request, in milliseconds. |
| `links_count` | Total number of discovered URLs returned. |
This information can be useful for monitoring request performance and understanding the size of the mapping result.
***
## Understanding `warning`
The `warning` field is optional.
It appears when the request succeeds but there is additional information that may help you improve the results.
For example:
```json
{
"success": true,
"links": [],
"metadata": {
"url": "https://geonode.com/",
"duration_ms": 4208,
"links_count": 0
},
"warning": "No results found. If you targeted a sub-path, try mapping the base domain for broader coverage."
}
```
In this example:
* The request completed successfully.
* No matching URLs were found.
* The warning suggests mapping the base domain instead of a sub-path.
A `warning` does not indicate that the request failed. Always check the `success` field before determining whether the request was successful.
***
## Empty Results
Sometimes a successful request may not return any discovered URLs.
Example:
```json
{
"success": true,
"links": [],
"metadata": {
"links_count": 0
}
}
```
This usually means:
* No URLs matched the request.
* The `search` filter did not match any discovered URLs.
* The mapped location did not contain discoverable pages.
***
## What Can You Do with the Results?
Once you've received the response, you can use the discovered URLs in several ways.
B["Review URLs"]
B --> C["Extraction API"]
B --> D["Crawl API"]
B --> E["Export Results"]
B --> F["Store in Database"]
`}
/>
Typical next steps include:
* Extract structured content from selected pages.
* Crawl larger sections of the website.
* Export the discovered URLs for reporting or analysis.
* Store the URLs for future processing.
***
## Best Practices
* Always check the `success` field before processing the response.
* Handle an empty `links` array as a valid response.
* Review the `warning` field when it is present.
* Use `links_count` to understand the size of the mapping result.
* Use the `source` field to understand how each URL was discovered.
***
## Next Steps
Now that you understand the Map response, continue to **Map Best Practices** to learn recommended approaches for building efficient URL discovery workflows.
# Retrieving Map Jobs (/docs/scraper-api/guides/map/05_retrieving_map_jobs)
After creating a mapping job, you can retrieve it later instead of creating a new request.
The Map API provides two endpoints for this:
* `GET /v1/map/jobs` — List your previous mapping jobs.
* `GET /v1/map/{job_id}` — Retrieve the complete details of a specific mapping job.
These endpoints are useful for reviewing previous jobs, recovering after an application restart, and accessing discovered URLs.
***
## Retrieval Workflow
The following diagram shows how both endpoints work together.
B["Receive job_id"]
B-->C["GET /v1/map/jobs"]
C-->D["Select a Job"]
D-->E["GET /v1/map/:job_id"]
E-->F["Review Discovered URLs"]
`}
/>
***
## When to Use Each Endpoint
| Endpoint | Use Case |
| ---------------------- | ------------------------------------------------- |
| `GET /v1/map/jobs` | View all of your previous mapping jobs. |
| `GET /v1/map/{job_id}` | Retrieve the complete details of one mapping job. |
***
## List Previous Mapping Jobs
Retrieve a paginated list of your mapping jobs.
```http
GET /v1/map/jobs
```
A successful response contains a list of jobs together with pagination information.
Each job includes:
| Field | Description |
| -------------- | ---------------------------------------------------- |
| `job_id` | Unique identifier of the mapping job. |
| `url` | Website that was mapped. |
| `status` | Current job status. |
| `links_count` | Number of discovered URLs. |
| `duration_ms` | Time required to complete the job. |
| `search` | Search filter used when the job was created, if any. |
| `error_code` | Error code if the job failed. Otherwise `null`. |
| `created_at` | Time when the job started. |
| `completed_at` | Time when the job finished. |
The response also includes:
| Field | Description |
| ------------ | --------------------------------- |
| `page` | Current page number. |
| `page_size` | Number of jobs returned per page. |
| `page_count` | Total number of available pages. |
***
## Finding the Right Job
Most applications identify a job by checking:
* The website URL.
* The job status.
* The creation time.
* The optional search filter.
B["Review URL"]
-->C["Check Status"]
-->D["Select job_id"]
-->E["Retrieve Details"]
`}
/>
***
## Retrieve a Specific Mapping Job
Once you have a `job_id`, retrieve the complete job information.
```http
GET /v1/map/{job_id}
```
This endpoint returns:
* Job information.
* Mapping configuration.
* Statistics.
* Every discovered URL.
***
## Understanding the Response
A completed mapping job contains several sections.
B["Job Information"]
A-->C["Configuration"]
A-->D["Statistics"]
A-->E["Discovered Links"]
`}
/>
### Job Information
These fields describe the mapping job itself.
| Field | Description |
| -------------- | ---------------------------------- |
| `job_id` | Unique job identifier. |
| `url` | Website that was mapped. |
| `status` | Current status of the mapping job. |
| `created_at` | Time when the job was created. |
| `completed_at` | Time when the job completed. |
***
### Mapping Configuration
These values show how the mapping job was executed.
| Field | Description |
| ------------------------- | ---------------------------------------------------------------------- |
| `search` | Search filter applied during URL discovery. |
| `include_subdomains` | Indicates whether subdomains were included. |
| `ignore_query_parameters` | Indicates whether query parameters were ignored when discovering URLs. |
***
### Statistics
The response also includes useful information about the completed job.
| Field | Description |
| ---------------- | --------------------------------------------- |
| `links_count` | Total number of discovered URLs. |
| `duration_ms` | Time taken to complete the mapping job. |
| `tokens_charged` | Number of API tokens charged for the request. |
***
### Discovered Links
The `links` array contains every URL discovered during the mapping process.
Each entry contains:
| Field | Description |
| -------- | --------------------------------------------------- |
| `url` | The discovered page URL. |
| `source` | Where the URL was discovered (`html` or `sitemap`). |
Example:
```json
{
"url": "https://docs.geonode.com",
"source": "sitemap"
}
```
***
## Recovering Previous Jobs
If your application loses the original `job_id`, you don't need to create another mapping request.
Instead:
1. Call `GET /v1/map/jobs`.
2. Find the required job.
3. Copy its `job_id`.
4. Retrieve the complete results with `GET /v1/map/{job_id}`.
B["GET /v1/map/jobs"]
-->C["Locate job_id"]
-->D["GET /v1/map/:job_id"]
-->E["Continue Processing"]
`}
/>
This approach helps avoid creating duplicate mapping jobs.
***
## Best Practices
* Save the returned `job_id` whenever you create a mapping job.
* Use `GET /v1/map/jobs` to locate previous jobs if the identifier is unavailable.
* Check the job `status` before using its results.
* Review `links_count` to understand how many URLs were discovered.
* Reuse completed mapping jobs whenever possible instead of creating duplicate requests.
***
## Next Steps
Now that you know how to retrieve previous mapping jobs and inspect their results, continue to **Best Practices** to learn recommendations for building efficient and reliable mapping workflows.
# Map Best Practices (/docs/scraper-api/guides/map/06_map_best_practices)
The Map API is often the first step in a scraping workflow. Following a few best practices can help you discover relevant URLs, reduce unnecessary processing, and build more efficient applications.
***
## Start with the Base Domain
Whenever possible, begin mapping from the root of the website.
For example:
```text
https://example.com
```
instead of:
```text
https://example.com/docs/getting-started
```
Starting from the base domain gives the Map API a broader view of the website and increases the chances of discovering all relevant pages.
B["More Discovered URLs"]
C["Sub-page"]
-->D["Limited Discovery"]
`}
/>
***
## Use `search` to Reduce Results
If you're only interested in a specific section of a website, use the `search` parameter.
For example:
```json
{
"url": "https://geonode.com",
"search": "docs"
}
```
This reduces the number of returned URLs and makes it easier to work with the results.
Typical search values include:
* `docs`
* `blog`
* `api`
* `support`
The `search` parameter performs a case-insensitive substring match against discovered URLs.
***
## Configure Subdomain Discovery Appropriately
The `include_subdomains` option controls whether supported subdomains are included during URL discovery.
Enable it when your content is distributed across multiple subdomains, such as:
```text
docs.example.com
blog.example.com
support.example.com
```
Disable it if you only want to discover pages from the primary domain.
***
## Decide How to Handle Query Parameters
Many websites generate multiple URLs that differ only by query parameters.
For example:
```text
/products?page=1
/products?page=2
/products?sort=newest
```
When `ignore_query_parameters` is enabled, these URLs are normalized during discovery.
Disable this option only if the query parameters represent unique content that should be treated as separate pages.
***
## Review Results Before Processing
The Map API discovers URLs—it does not extract page content.
Review the discovered URLs before deciding what to process next.
B["Review URLs"]
B-->C["Extract Selected Pages"]
B-->D["Crawl Website"]
B-->E["Export URL List"]
`}
/>
This approach helps reduce unnecessary extraction and crawling.
***
## Reuse Existing Mapping Jobs
If you have already mapped a website, retrieve the existing job instead of creating another mapping request.
B["GET /v1/map/jobs"]
-->C["Select job_id"]
-->D["GET /v1/map/:job_id"]
-->E["Continue Processing"]
`}
/>
Reusing existing jobs helps avoid duplicate requests and allows your application to continue working with previously discovered URLs.
***
## Handle Empty Results Gracefully
A successful request may return an empty `links` array.
For example:
```json
{
"success": true,
"links": [],
"warning": "No results found. If you targeted a sub-path, try mapping the base domain for broader coverage."
}
```
This does not indicate that the request failed.
Instead, consider:
* Mapping the base domain instead of a sub-path.
* Using a less restrictive `search` value.
* Verifying that the target website contains discoverable pages.
***
## Build a Complete Mapping Workflow
The Map API works best as the first stage of a larger scraping pipeline.
B["Map API"]
-->C["Review URLs"]
C-->D["Extraction API"]
C-->E["Crawl API"]
D-->F["Store Results"]
E-->F
`}
/>
This workflow allows you to discover URLs first and then process only the pages that are relevant to your application.
***
## Best Practices Checklist
Before deploying your application, make sure you:
* Start mapping from the base domain whenever possible.
* Use `search` to narrow large result sets.
* Enable subdomain discovery only when needed.
* Configure query parameter handling based on your use case.
* Review discovered URLs before extraction or crawling.
* Reuse existing mapping jobs whenever possible.
* Handle empty results as a valid response.
* Build your workflow around URL discovery before content extraction.
***
## Next Steps
You now understand how to use the Map API effectively.
Continue to the **Map API Reference** for detailed endpoint documentation, request schemas, response fields, and additional examples.
# Overview (/docs/scraper-api/guides/mcp/00_overview)
AI assistants are changing how developers build software. Instead of switching between documentation, browsers, terminals, and APIs, developers can now ask their AI assistant to perform tasks such as writing code, debugging applications, analyzing data, and interacting with external services.
However, AI assistants have one important limitation—they cannot access live websites, external APIs, or other tools on their own.
If you are an AI coding agent onboarding to Geonode, start with the [Agent Onboarding skill](/agent-onboarding/SKILL.md).
## What is MCP?
The **Model Context Protocol (MCP)** is an open standard that enables AI assistants to securely connect to external tools, APIs, and data sources.
Instead of relying only on the information available in a conversation, an AI assistant can use MCP to retrieve live data, execute actions, and return results in real time.
This allows AI assistants to become active participants in your workflow rather than simple conversational assistants.
## Why use MCP instead of manually calling APIs?
Without MCP, interacting with external services often requires switching between multiple tools.
You typically need to:
* Find the documentation.
* Write an API request.
* Send the request.
* Review the response.
* Copy the results back into your AI assistant.
With MCP, your AI assistant performs these tasks for you. You simply describe what you need, and the assistant communicates with the appropriate service behind the scenes.
For example, instead of manually calling an API, you can simply ask:
* Extract the content from this product page.
* Crawl the official React documentation.
* Process these 500 URLs.
* Check whether my extraction job has finished.
The AI assistant automatically selects the appropriate MCP tool to complete your request.
## What is Geonode MCP?
Geonode MCP connects your AI assistant directly to Geonode's web scraping platform.
Once connected, your assistant can retrieve live web content, extract structured data, crawl entire websites, process batch extraction jobs, and monitor long-running tasks without leaving your editor or chat.
Instead of managing API requests yourself, you simply describe the task, and your assistant uses Geonode's scraping capabilities on your behalf.
## What can you do with Geonode MCP?
With Geonode MCP, your AI assistant can:
* Extract content from individual web pages.
* Crawl websites and documentation portals.
* Process hundreds or thousands of URLs in batch jobs.
* Retrieve structured content in Markdown or HTML.
* Render JavaScript-powered websites.
* Route requests through residential proxies with geo-targeting.
* Include custom request headers when required.
* Monitor the progress of long-running extraction and crawl jobs.
## How Geonode MCP Works
```text
You
│
▼
AI Assistant
(Cursor, Claude, Windsurf, etc.)
│
▼
Geonode MCP Server
│
▼
Geonode Scraper API
│
▼
Target Website
│
▼
Structured Results
│
▼
AI Assistant Response
```
## Try without an account
You can point an MCP client at `https://scraper.geonode.io/mcp` with **no API key** and use the `extract` tool under free anonymous limits.
See [Use Without an API Key](/docs/scraper-api/guides/mcp/01_use-without-an-api-key) for the endpoint, client config, limits, and how to upgrade.
## Supported Clients
Geonode MCP supports the following AI assistants and development environments:
* [Claude Desktop](/docs/scraper-api/guides/mcp/03_claude-desktop)
* [Claude Code](/docs/scraper-api/guides/mcp/02_claude-code)
* [Cursor](/docs/scraper-api/guides/mcp/04_cursor)
* Windsurf
* Smithery
* [Docker](/docs/scraper-api/guides/mcp/05_docker-mcp)
* [Visual Studio Code](/docs/scraper-api/guides/mcp/08_visual-studio-code)
* [Codex](/docs/scraper-api/guides/mcp/09_codex)
Each client has its own setup guide with step-by-step installation instructions.
## Next Steps
To try Geonode MCP with no signup, start with [Use Without an API Key](/docs/scraper-api/guides/mcp/01_use-without-an-api-key).
For full access (all tools, async jobs, JavaScript rendering, and plan limits), continue to [Before You Start](/docs/scraper-api/guides/mcp/01_before-you-start), then follow the setup guide for your preferred client.
# Before You Start (/docs/scraper-api/guides/mcp/01_before-you-start)
Before configuring Geonode MCP with full access, make sure you have everything you need. This guide covers the API key, authentication, endpoint, and the tools your AI assistant can access once connected.
You can use Geonode MCP without an API key for small jobs. See [Use Without an API Key](/docs/scraper-api/guides/mcp/01_use-without-an-api-key). This page covers authenticated setup for the full tool set.
## Prerequisites
Before continuing, you'll need:
* A Geonode account.
* A valid Geonode API key.
* An MCP-compatible client such as Claude Desktop, Claude Code, Cursor, Windsurf, or Docker.
## Get Your API Key
Authenticated requests through the Geonode MCP Server use an API key.
To create or view your API key:
1. Sign in to your Geonode dashboard.
2. Open **API Keys**.
3. Copy an existing key or create a new one.
4. Store your API key securely.
Never share your API key or commit it to a public repository. Anyone with your API key can make requests on your behalf.
## MCP Endpoint
Configure your MCP client to connect to the following endpoint:
```text
https://scraper.geonode.io/mcp
```
## Authentication
Geonode MCP authenticates requests using your API key.
Most MCP clients send the key through the `X-Api-Key` HTTP header.
Some clients cannot configure custom headers. In those cases, you can provide the key as the `api_key` argument when calling a tool.
Using the `X-Api-Key` header is the recommended approach, and all setup guides in this documentation use this method.
Replace `YOUR_API_KEY` with your own Geonode API key in every configuration example throughout this documentation.
## Available MCP Tools
Once connected, your AI assistant can automatically use the following Geonode tools.
### Content Extraction
| Tool | Description |
| --------- | -------------------------------------------------------------------------------------------------------------------------------------------------- |
| `extract` | Extract content from a single web page in Markdown or HTML. Supports JavaScript rendering, residential proxies, geo-targeting, and custom headers. |
| `job` | Retrieve the result of a previously submitted asynchronous extraction job. |
| `jobs` | List your extraction jobs and filter them by status, URL, date, or output format. |
### Batch Processing
| Tool | Description |
| -------------- | ------------------------------------------------------------------- |
| `batch` | Extract up to 1,000 URLs in a single asynchronous batch job. |
| `batch_status` | Monitor the progress of a batch job and retrieve paginated results. |
| `cancel_batch` | Cancel a running batch job. |
### Website Crawling
| Tool | Description |
| -------------- | ---------------------------------------------------------------------------------------------------------------------------- |
| `crawl` | Crawl an entire website starting from a seed URL. Supports breadth-first crawling, configurable depth, and domain filtering. |
| `crawl_status` | Check the progress of a crawl job and retrieve discovered pages. |
| `cancel_crawl` | Cancel a running crawl job. |
### Usage Statistics
| Tool | Description |
| ------------ | ---------------------------------------------------------------------------------- |
| `statistics` | Retrieve usage statistics such as extraction count, success rate, and token usage. |
## How Your AI Uses These Tools
You don't need to manually choose which tool to use.
Instead, simply describe the task, and your AI assistant selects the appropriate Geonode tool automatically.
For example:
| You ask | AI uses |
| ----------------------------------------------- | ----------------------- |
| "Extract the content from this page." | `extract` |
| "Process these 500 URLs." | `batch` |
| "Crawl this documentation website." | `crawl` |
| "Check whether my extraction job has finished." | `job` or `batch_status` |
| "Show my extraction statistics." | `statistics` |
## Next Steps
You're now ready to configure Geonode MCP with an API key.
To try first without a key, see [Use Without an API Key](/docs/scraper-api/guides/mcp/01_use-without-an-api-key).
Choose the setup guide for your preferred client:
* [Claude Desktop](/docs/scraper-api/guides/mcp/03_claude-desktop)
* [Claude Code](/docs/scraper-api/guides/mcp/02_claude-code)
* [Cursor](/docs/scraper-api/guides/mcp/04_cursor)
* Windsurf
* Smithery
* [Docker MCP](/docs/scraper-api/guides/mcp/05_docker-mcp)
# Use Without an API Key (/docs/scraper-api/guides/mcp/01_use-without-an-api-key)
You can connect Geonode MCP to your AI assistant and start extracting web pages without creating an account or adding an API key.
Point your MCP client at the Geonode endpoint, and your assistant can use the `extract` tool right away. This path is ideal for trying the service and for small jobs. When you need more capacity or more tools, add an API key to the same configuration.
* How to connect without an API key
* What works in anonymous mode
* Free rate and volume limits
* How to upgrade to authenticated access
***
## MCP Endpoint
Configure your MCP client with this endpoint:
```text
https://scraper.geonode.io/mcp
```
The server uses MCP over streamable HTTP.
To confirm the endpoint is reachable, run:
```bash
curl https://scraper.geonode.io/mcp/health
```
A healthy response looks like this:
```json
{"status":"ok","transport":"streamable-http","endpoint":"/mcp"}
```
Always use `https://scraper.geonode.io/mcp`. Other hostnames, including preproduction ones, are not supported for public use and will fail.
***
## Connect Your Client
You do not need a key, token, or custom header for anonymous access. If your client asks for authentication, leave it empty.
### Claude Code
```bash
claude mcp add --transport http geonode https://scraper.geonode.io/mcp
```
### Cursor, Claude Desktop, and other JSON-based clients
```json
{
"mcpServers": {
"geonode": {
"url": "https://scraper.geonode.io/mcp"
}
}
}
```
After you save the configuration, restart your client if required, then open a new chat and ask your assistant to extract a page.
***
## What You Can Do Without a Key
Without an API key, Geonode MCP exposes one tool: `extract`. It fetches a URL and returns the page content.
You can ask your assistant something like:
```text
Extract the content from https://example.com
```
Behind the scenes, the assistant calls `extract` on that URL.
Anonymous access is intentionally limited compared with a keyed connection:
| Capability | Without a key | With an API key |
| -------------------- | ---------------------- | ---------------------------------------------------------------- |
| Tools | `extract` only | `extract`, `map`, `search`, `crawl`, `batch`, and job management |
| Mode | Synchronous only | Synchronous and asynchronous jobs |
| JavaScript rendering | Not available | Available with `render_js` |
| Custom headers | `Accept-Language` only | Any headers you need |
| Country selection | Not available | Available |
If you request a feature that anonymous access does not support, the server returns a clear message about which limit applied, instead of failing silently.
***
## Free Limits
Anonymous use is rate limited by network address:
| Limit | Value |
| --------------- | ------------------------- |
| Requests | 5 per minute |
| Concurrency | 1 request at a time |
| Download volume | 100 MB of content per day |
If several people share the same office or provider network, they may share a larger group allowance for that network prefix. Limits then apply to the group rather than to each person alone.
If requests keep getting refused, a short cooling-off period starts and can grow if the refusals continue. Wait for it to clear; the limits reset on their own.
Anonymous traffic is metered but not billed. No payment method is required.
Requests that are refused are not billed, with or without an API key.
***
## Upgrade to an API Key
When you need the full tool set, async jobs, JavaScript rendering, country targeting, or higher limits, create a Geonode account, get an API key, and add it to the same MCP configuration:
```json
{
"mcpServers": {
"geonode": {
"url": "https://scraper.geonode.io/mcp",
"headers": {
"X-Api-Key": "YOUR_API_KEY"
}
}
}
}
```
Replace `YOUR_API_KEY` with your Geonode API key.
As soon as the key is recognized, the available tools expand and your plan rate limits apply.
* Plans and pricing: [geonode.com/pricing](https://geonode.com/pricing)
* Full authenticated setup: [Before You Start](/docs/scraper-api/guides/mcp/01_before-you-start)
Then follow the setup guide for your client, such as [Claude Code](/docs/scraper-api/guides/mcp/02_claude-code), [Claude Desktop](/docs/scraper-api/guides/mcp/03_claude-desktop), or [Cursor](/docs/scraper-api/guides/mcp/04_cursor).
***
## Next Steps
1. Connect your client to `https://scraper.geonode.io/mcp` with no key.
2. Ask your assistant to extract a public page.
3. When you outgrow the free limits, add an API key and continue with [Before You Start](/docs/scraper-api/guides/mcp/01_before-you-start).
# Claude Code (/docs/scraper-api/guides/mcp/02_claude-code)
Claude Code has built-in support for the Model Context Protocol (MCP), allowing it to securely connect to external tools such as the Geonode Scraper API.
Once configured, Claude can automatically extract webpages, crawl websites, process multiple URLs, and retrieve structured content using natural language prompts.
Make sure you've completed the **Before You Start** guide and have:
* A Geonode API key
* The Geonode MCP endpoint
* Claude Code installed
* Claude Code authenticated with your Anthropic account
***
## Step 1 — Open your terminal
Open your terminal and verify that Claude Code is installed.
Launch Claude Code by running:
```bash
claude
```
If Claude Code starts successfully, you're ready to configure the Geonode MCP server.
***
## Step 2 — Register the Geonode MCP Server
Run the following command to register the Geonode MCP server with Claude Code.
```bash
claude mcp add \
--transport http \
geonode-scraper \
https://scraper.geonode.io/mcp \
--header "X-Api-Key:YOUR_API_KEY"
```
Replace `YOUR_API_KEY` with your Geonode API key.
Claude Code saves this configuration locally and automatically registers the MCP server.
***
## Step 3 — Verify the Connection
To verify that the server has been registered successfully, run:
```bash
claude mcp list
```
You should see output similar to:
```text
geonode-scraper: https://scraper.geonode.io/mcp (HTTP) - Connected
```
If the status is **Connected**, Claude Code can communicate with the Geonode MCP server.
***
## Step 4 — Start Claude Code
Launch Claude Code.
```bash
claude
```
Claude automatically loads all configured MCP servers when it starts.
***
## Step 5 — Use Geonode with Natural Language
You don't need to manually invoke MCP tools.
Simply describe the task, and Claude automatically chooses the appropriate Geonode MCP tool.
### Extract a web page
```text
Extract the content from https://geonode.com
```
### Summarize a page
```text
Read https://geonode.com and summarize the homepage.
```
### Crawl a documentation website
```text
Crawl https://docs.geonode.com and tell me what documentation sections exist.
```
### Process multiple pages
```text
Extract these URLs and summarize each page.
https://example.com/page1
https://example.com/page2
https://example.com/page3
```
### Check extraction statistics
```text
Show my Geonode extraction statistics.
```
Claude automatically calls the appropriate MCP tool, waits for the response, and presents the results directly in the conversation.
***
## How Claude Code Uses MCP
When you submit a prompt, Claude analyzes your request and automatically selects the correct Geonode MCP tool.
| Your Prompt | MCP Tool |
| -------------------------------- | ------------ |
| "Extract this page." | `extract` |
| "Crawl this documentation site." | `crawl` |
| "Process these URLs." | `batch` |
| "Show my extraction statistics." | `statistics` |
There is no need to manually select or invoke MCP tools. Claude Code handles tool selection, parameter mapping, execution, and result processing automatically.
***
## FAQs
Verify that:
* The server URL is `https://scraper.geonode.io/mcp`.
* Your `X-Api-Key` header contains a valid Geonode API key.
* Your internet connection is active.
Run the following command to verify the connection:
```bash
claude mcp list
```
If necessary, remove and add the server again.
Authentication failures usually occur when:
* The API key is incorrect.
* The `X-Api-Key` header is missing.
* The API key has expired or been revoked.
Generate a new API key and register the server again.
If Claude cannot reach the MCP server:
* Verify the server URL.
* Check your internet connection.
* Confirm that the Geonode service is available.
Then run:
```bash
claude mcp list
```
to confirm the server status.
Run:
```bash
claude mcp list
```
Claude displays every configured MCP server along with its current connection status.
Run:
```bash
claude mcp remove geonode-scraper
```
This removes the Geonode MCP configuration from Claude Code.
# Claude Desktop (/docs/scraper-api/guides/mcp/03_claude-desktop)
import { Accordion, Accordions } from 'fumadocs-ui/components/accordion';
Claude Desktop supports the Model Context Protocol (MCP), so you can use Geonode scraping tools directly in chat.
Geonode authenticates with a static `X-Api-Key` header. In Claude Desktop, the working setup today is **Settings → Developer → Edit Config** with `mcp-remote`.
**Add custom connector** only accepts a server URL and optional OAuth. It has no field for static headers, so it cannot authenticate to Geonode today.
Use **Developer → Edit Config** instead. Custom connector support will be improved later.
Make sure you've completed the **Before You Start** guide and have:
* A Geonode API key
* The Geonode MCP endpoint: `https://scraper.geonode.io/mcp`
* Claude Desktop installed
* **Node.js 18+** installed (required by `mcp-remote`)
***
## Step 1 — Open Settings
Open Claude Desktop.
Click your profile menu in the bottom-left corner, then select **Settings**.
***
## Step 2 — Open Developer → Edit Config
In Settings:
1. Under **Desktop app**, open **Developer**.
2. Find **Local MCP servers**.
3. Click **Edit Config**.
This opens `claude_desktop_config.json`.
You can also edit the file directly:
| OS | Path |
| ----------- | ----------------------------------------------------------------- |
| **macOS** | `~/Library/Application Support/Claude/claude_desktop_config.json` |
| **Windows** | `%APPDATA%\Claude\claude_desktop_config.json` |
***
## Step 3 — Add the Geonode MCP server
Add the Geonode server under the top-level `mcpServers` key.
If the file already has other settings, keep them and place `mcpServers` **alongside** those top-level keys — do not nest it inside another object.
```json
{
"mcpServers": {
"geonode-scraper": {
"command": "npx",
"args": [
"mcp-remote",
"https://scraper.geonode.io/mcp",
"--header",
"X-Api-Key:${GEONODE_API_KEY}"
],
"env": {
"GEONODE_API_KEY": "YOUR_API_KEY"
}
}
}
}
```
Replace `YOUR_API_KEY` with your Geonode API key.
Write `X-Api-Key:${GEONODE_API_KEY}` with **no space** after the colon.
Claude Desktop splits `args` on spaces. A space after the colon drops the key value and authentication fails with `401`.
`mcpServers` must sit at the root of `claude_desktop_config.json`, next to any existing keys such as preferences.
Do not nest `mcpServers` inside another object.
***
## Step 4 — Fully restart Claude Desktop
1. Save `claude_desktop_config.json`.
2. **Quit Claude Desktop completely** — close the process, not only the window.
3. Open Claude Desktop again.
4. Start a **new chat**.
A partial window close is not enough. MCP servers load on a full app restart.
***
## Step 5 — Verify the connection
In a new chat, ask Claude to call a Geonode tool. For example:
```text
Call the geonode scraper statistics endpoint and tell me the HTTP status and a short summary of the response.
```
If the connection works:
* Claude uses the Geonode MCP tools
* You do **not** need to pass `api_key` in the chat prompt
* A successful statistics call returns **HTTP 200**
***
## Step 6 — Start using Geonode MCP
Once connected, describe tasks in natural language. Claude selects the Geonode tool automatically.
### Extract a web page
```text
Extract the content from https://geonode.com
```
### Summarize a page
```text
Read https://geonode.com and summarize the homepage.
```
### Crawl a documentation website
```text
Crawl https://docs.geonode.com and tell me what documentation sections exist.
```
### Check extraction statistics
```text
Show my Geonode extraction statistics.
```
***
## Important setup notes
| Requirement | Why it matters |
| --------------------------------------------- | ------------------------------------------------------------- |
| Use **Edit Config**, not Add custom connector | Custom connector cannot send static `X-Api-Key` headers today |
| Keep `mcpServers` at the **top level** | Nested config is ignored |
| No space in `X-Api-Key:${GEONODE_API_KEY}` | Spaces break header parsing and cause `401` |
| **Node.js 18+** | Required by `npx mcp-remote` |
| Full app restart + new chat | MCP servers load only after a complete restart |
***
## FAQs
Not for Geonode right now. **Add custom connector** only supports a server URL and optional OAuth. It has no static header field, so it cannot send `X-Api-Key`.
Use **Developer → Edit Config** with `mcp-remote` instead.
Check these first:
* `X-Api-Key:${GEONODE_API_KEY}` has **no space** after the colon
* `GEONODE_API_KEY` in `env` is set to a valid Geonode API key
* You fully quit and restarted Claude Desktop
* You are testing in a **new chat**
Verify that:
* `mcpServers` is a top-level key in `claude_desktop_config.json`
* Node.js 18+ is installed (`node -v`)
* The JSON file is valid
* Claude Desktop was fully quit and reopened
* You opened a new chat after restarting
No. After Edit Config is set up correctly, the API key is sent through the `X-Api-Key` header by `mcp-remote`. You do not need to include the key in prompts.
| OS | Path |
| ----------- | ----------------------------------------------------------------- |
| **macOS** | `~/Library/Application Support/Claude/claude_desktop_config.json` |
| **Windows** | `%APPDATA%\Claude\claude_desktop_config.json` |
You can also open it from **Settings → Developer → Edit Config**.
# Cursor (/docs/scraper-api/guides/mcp/04_cursor)
import { Accordion, Accordions } from 'fumadocs-ui/components/accordion';
Cursor supports remote MCP servers over HTTP, allowing your AI assistant to securely connect to Geonode's scraping platform.
Once connected, Cursor can automatically extract web pages, crawl websites, process batch jobs, and retrieve structured content based on your prompts.
Make sure you've completed the **Before You Start** guide and have:
* A Geonode API key
* The Geonode MCP endpoint
* Cursor installed
## Step 1 — Open MCP Settings
Open Cursor and navigate to:
Settings → Tools & MCPs
Under Home MCP Servers, click ew MCP Server
***
## Step 2 — Create a New MCP Server
Configure the server with the following information:
| Field | Value |
| -------------- | -------------------------------- |
| **Name** | `geonode-scraper` |
| **Type** | `URL` |
| **Server URL** | `https://scraper.geonode.io/mcp` |
Under **Headers**, add:
| Key | Value |
| ----------- | -------------- |
| `X-Api-Key` | `YOUR_API_KEY` |
Replace `YOUR_API_KEY` with your Geonode API key.
***
## Step 3 — Save the Configuration
Click **Add MCP**.
Cursor creates the server configuration automatically.
If you prefer editing the configuration manually, the generated `mcp.json` looks like this:
```json
{
"mcpServers": {
"geonode-scraper": {
"url": "https://scraper.geonode.io/mcp",
"headers": {
"X-Api-Key": "YOUR_API_KEY"
}
}
}
}
```
***
## Step 4 — Verify the Connection
Return to **Settings → Tools & MCPs**.
If the connection is successful, you'll see:
* A green status indicator.
* The **geonode-scraper** server.
* All available Geonode tools.
The available tools include:
* `extract`
* `job`
* `jobs`
* `statistics`
* `batch`
* `batch_status`
* `cancel_batch`
* `crawl`
* `crawl_status`
* `cancel_crawl`
***
## Step 5 — Start Using Geonode MCP
Open a new **Agent** or **Composer** chat.
You don't need to manually select a tool.
Simply describe the task in natural language, and Cursor automatically chooses the appropriate Geonode MCP tool.
For example:
### Extract a web page
```text
Extract the content from https://geonode.com
```
### Summarize a page
```text
Read https://geonode.com and summarize the homepage.
```
### Crawl a documentation website
```text
Crawl https://docs.geonode.com and tell me what documentation sections exist.
```
### Process multiple pages
```text
Extract these URLs and summarize each page.
https://example.com/page1
https://example.com/page2
https://example.com/page3
```
### Check extraction statistics
```text
Show my Geonode extraction statistics.
```
***
## How it works
When you send a prompt, Cursor analyzes your request and automatically calls the appropriate Geonode MCP tool.
For example:
| Your Prompt | MCP Tool |
| -------------------------------- | ------------ |
| "Extract this page." | `extract` |
| "Crawl this documentation site." | `crawl` |
| "Process these URLs." | `batch` |
| "Show my extraction statistics." | `statistics` |
You never need to call these tools manually—Cursor handles tool selection and parameter mapping for you.
***
## Viewing Tool Calls
During execution, Cursor displays each MCP tool invocation in the chat.
You can expand each tool call to inspect:
* The tool that was executed.
* The request parameters.
* The response returned by the Geonode MCP Server.
This can be useful for debugging, understanding how Cursor interprets your prompts, or learning which MCP tool was used for a particular request.
***
## FAQs
Verify the following:
* The server URL is correct.
* Your Geonode API key is valid.
* The MCP server configuration has been saved.
If the server still doesn't appear, restart Cursor and try again.
Ensure your `X-Api-Key` header contains a valid Geonode API key.
If you've recently generated a new key, update your MCP configuration and reconnect the server.
Open **Settings → Tools & MCPs** and verify:
* The server has a green status indicator.
* The connection is active.
* The list of Geonode tools is displayed.
If the tools are missing, refresh the MCP configuration or reconnect the server.
# Docker MCP (/docs/scraper-api/guides/mcp/05_docker-mcp)
Docker provides two ways to use the Geonode MCP Server.
* **Docker Desktop** provides a graphical interface for managing MCP servers and AI clients.
* **Docker CLI** runs a lightweight MCP bridge directly from the terminal.
Choose the option that best fits your workflow.
Make sure you've completed the **Before You Start** guide and have:
* A Geonode API key
* The Geonode MCP endpoint
* Docker installed
***
# Option 1 — Docker Desktop
Docker Desktop includes the **MCP Toolkit**, which allows you to manage MCP servers and connect them to supported AI clients.
### Step 1 — Install Docker Desktop
Download and install Docker Desktop for your operating system.
Launch Docker Desktop and ensure the Docker Engine is running.
### Step 2 — Open the MCP Toolkit
From the Docker Desktop sidebar, open:
**Models → MCP Toolkit**
This is where Docker manages MCP servers and connected AI clients.
### Step 3 — Connect Your AI Client
Open the **Clients** tab.
Choose the AI client you want to use.
For example:
* Cursor
* Continue.dev
* Codex
* Gemini CLI
* Goose
Click **Connect** next to your client.
Docker configures the client to communicate with the MCP Toolkit.
### Step 4 — Restart Your AI Client
After connecting the client, restart it.
Once restarted, Docker automatically exposes the configured MCP servers to the client.
### Step 5 — Verify the Connection
Open your AI client's MCP settings.
You should see:
* a green status indicator
* the configured MCP server
* the available Geonode tools
***
# Option 2 — Docker CLI
If you prefer using the terminal, Docker can run a lightweight MCP bridge without any additional installation.
### Step 1 — Verify Docker
Open a terminal.
Run:
```bash
docker --version
```
If Docker is installed correctly, the installed version is displayed.
### Step 2 — Start the MCP Bridge
Run:
```bash
docker run -i --rm \
-e GEONODE_API_KEY=YOUR_API_KEY \
node:20-alpine \
npx -y mcp-remote https://scraper.geonode.io/mcp \
--header "X-Api-Key:${GEONODE_API_KEY}"
```
Replace `YOUR_API_KEY` with your Geonode API key.
This command:
* starts a temporary Docker container
* downloads `mcp-remote`
* connects to the Geonode MCP Server
* creates a local MCP bridge over STDIO
### Step 3 — Verify the Bridge
When the bridge starts successfully, you'll see output similar to:
```text
Using transport strategy: http-first
Connected to remote server using StreamableHTTPClientTransport
Local STDIO server running
Proxy established successfully
Press Ctrl+C to exit
```
The bridge is now ready to receive requests.
### Step 4 — Configure Your AI Client
Configure your AI client to use the local Docker bridge.
Once connected, the client automatically gains access to the available Geonode MCP tools.
***
# Start Using Geonode MCP
You can now ask your AI assistant to perform tasks such as:
#### Extract a web page
```text
Extract the content from https://geonode.com
```
#### Summarize a page
```text
Summarize https://geonode.com
```
#### Crawl a website
```text
Crawl https://docs.geonode.com
```
#### Process multiple pages
```text
Extract and summarize these URLs:
https://example.com/page1
https://example.com/page2
https://example.com/page3
```
#### View extraction statistics
```text
Show my Geonode extraction statistics.
```
***
# Stopping the Docker Bridge
To stop the Docker bridge, press:
```text
Ctrl + C
```
Since the container is started with the `--rm` flag, Docker automatically removes it after it exits.
***
# FAQ
Ensure Docker Desktop is installed and the Docker Engine is running before starting the bridge.
Verify that your Geonode API key is correct.
Confirm that Docker has internet access and that the Geonode MCP endpoint is reachable.
Restart your AI client after configuring the MCP bridge, then verify the MCP server is enabled in the client's settings.
Press **Ctrl +C** in the terminal. Docker automatically removes the temporary container because the command uses the `--rm` option.
# Devin AI (/docs/scraper-api/guides/mcp/07_devin-ai)
Devin.ai supports the Model Context Protocol (MCP), allowing AI agents to securely connect to external tools like the Geonode Scraper API.
Once connected, Cascade can automatically extract web pages, crawl websites, process multiple URLs, and retrieve structured content based on your prompts.
Windsurf has been acquired by Devin AI. Recent versions use the Devin branding throughout the application, while older releases may still display Windsurf. The MCP configuration process is the same for both.
Make sure you've completed the **Before You Start** guide and have:
* A Geonode API key
* The Geonode MCP endpoint
* A Devin.ai installation with MCP support.
***
## Step 1 — Open the MCP Marketplace
Open **Devin.ai**.
Navigate to:
**Settings → MCP Marketplace**
From the MCP Marketplace, click **Add custom MCP**.
***
## Step 2 — Add the Geonode MCP Server
Configure the MCP server using the following information.
| Field | Value |
| -------------- | -------------------------------- |
| **Name** | `geonode-scraper` |
| **Server URL** | `https://scraper.geonode.io/mcp` |
Under **Headers**, add the following authentication header.
| Header | Value |
| ----------- | -------------- |
| `X-Api-Key` | `YOUR_API_KEY` |
Replace `YOUR_API_KEY` with your Geonode API key.
***
## Step 3 — Connect the Server
Click **Connect**.
If the configuration is correct, Devin.ai establishes a connection to the Geonode MCP server.
The status changes from **Not Connected** to **Connected**.
***
## Step 4 — Verify the Configuration
After connecting, open the MCP server details to verify that:
* The server is connected.
* Your API key has been saved.
* The Geonode MCP server appears in the list of installed MCP servers.
If the server is shown as **Connected**, the setup is complete.
***
## Step 5 — Start Using Geonode MCP
Open a new **Cascade** conversation.
You don't need to manually choose an MCP tool.
Simply describe the task in natural language, and Devin.ai automatically selects the appropriate Geonode MCP tool.
For example:
### Extract a web page
```text
Extract the content from https://geonode.com
```
### Summarize a page
```text
Read https://geonode.com and summarize the homepage.
```
### Crawl a documentation website
```text
Crawl https://docs.geonode.com and tell me what documentation sections exist.
```
### Process multiple pages
```text
Extract these URLs and summarize each page.
https://example.com/page1
https://example.com/page2
https://example.com/page3
```
### Check extraction statistics
```text
Show my Geonode extraction statistics.
```
***
## How Devin.ai Uses MCP
When you send a prompt, Devin.ai analyzes your request and automatically calls the appropriate Geonode MCP tool.
For example:
| Your Prompt | MCP Tool |
| ------------------------------ | ------------ |
| Extract this page. | `extract` |
| Crawl this documentation site. | `crawl` |
| Process these URLs. | `batch` |
| Show my extraction statistics. | `statistics` |
You never need to manually invoke these tools—Cascade handles tool selection and parameter mapping automatically.
***
## FAQ's
Verify that:
* The server URL is `https://scraper.geonode.io/mcp`.
* Your `X-Api-Key` header contains a valid Geonode API key.
* Your internet connection is active.
After making changes, try connecting again.
Authentication failures usually occur when:
* The API key is incorrect.
* The `X-Api-Key` header is missing.
* The API key has expired or been revoked.
Generate a new API key if necessary and reconnect the server.
If the MCP server cannot be reached:
* Verify that the server URL is correct.
* Check your network connection.
* Confirm that the Geonode service is available.
Then reconnect the MCP server.
Ensure that:
* The MCP server status is **Connected**.
* The server appears in the installed MCP servers list.
* Devin.ai has finished loading the MCP configuration.
If necessary, reconnect the MCP server or restart Devin.ai.
# Visual Studio Code (/docs/scraper-api/guides/mcp/08_visual-studio-code)
Visual Studio Code supports the **Model Context Protocol (MCP)**, allowing **GitHub Copilot Agent Mode** to securely connect to external tools such as the Geonode Scraper API.
Once configured, GitHub Copilot can automatically extract web pages, crawl websites, process multiple URLs, and retrieve structured content directly from your editor.
Make sure you've completed the **Before You Start** guide and have:
* A Geonode API key.
* The Geonode MCP endpoint.
* Visual Studio Code installed.
* GitHub Copilot installed and signed in.
* Agent Mode enabled.
***
## Step 1 — Open the MCP Servers Panel
Open **Visual Studio Code** and open the **GitHub Copilot Chat** panel.
Click the **Settings (⚙️)** icon, then open **Agent Customizations** and select **MCP Servers** from the left navigation.
This page lets you add and manage MCP servers for your workspace.
***
## Step 2 — Add a New MCP Server
Click the **+** button in the upper-right corner of the **MCP Servers** page.
From the available options, select:
```text
HTTP (HTTP or Server-Sent Events)
```
This option allows Visual Studio Code to connect to a remote MCP server over HTTP.
***
## Step 3 — Enter the Geonode MCP Endpoint
When prompted, enter the Geonode MCP endpoint:
```text
https://scraper.geonode.io/mcp
```
Press **Enter** to save the server configuration.
***
## Step 4 — Verify the Connection
After adding the endpoint, Visual Studio Code automatically attempts to connect to the MCP server.
If the connection is successful, the server appears under **Workspace** with a **Running** status.
If the Geonode MCP server displays **Running**, the connection has been established successfully and GitHub Copilot can begin using the available MCP tools.
***
## Step 5 — Review the Generated Configuration
Visual Studio Code automatically creates an **mcp.json** file in your workspace.
This file stores your MCP server configuration and can be updated later if your server configuration changes.
The Geonode MCP server requires authentication using your Geonode API key. Configure authentication according to your deployment before using the server.
***
## Step 6 — Start Using Geonode MCP
Once the MCP server is connected, open the **GitHub Copilot Chat** panel and make sure **Agent** mode is selected.
You can now interact with the Geonode MCP server using natural language. GitHub Copilot automatically selects the appropriate MCP tool based on your request.
For example:
### Extract a web page
```text
Extract the content from https://geonode.com
```
### Summarize a page
```text
Read https://geonode.com and summarize the homepage.
```
### Crawl a website
```text
Crawl https://docs.geonode.com and summarize its documentation.
```
### Process multiple URLs
```text
Extract and summarize the following URLs:
https://example.com/page1
https://example.com/page2
https://example.com/page3
```
***
## How GitHub Copilot Uses MCP
When you submit a prompt in **Agent Mode**, GitHub Copilot analyzes your request and automatically determines whether an MCP tool is required.
If your request involves extracting web content, crawling websites, or processing multiple URLs, GitHub Copilot invokes the appropriate Geonode MCP tool behind the scenes and returns the results directly in the chat.
You don't need to manually choose or invoke MCP tools.
***
## FAQs
Verify that:
* The MCP endpoint is correct.
* Your Geonode API key has been configured.
* Your internet connection is active.
After making any changes, restart the MCP server or reload Visual Studio Code.
If the server does not display a **Running** status, verify that the endpoint and authentication have been configured correctly.
Ensure that:
* GitHub Copilot is signed in.
* Agent Mode is enabled.
* The Geonode MCP server is running.
Then submit your request again.
Open the **MCP Servers** panel, remove the server configuration, and save your changes.
# Codex (/docs/scraper-api/guides/mcp/09_codex)
The Codex desktop application supports the **Model Context Protocol (MCP)**, allowing it to securely connect to external tools such as the Geonode Scraper API.
Once connected, Codex can automatically extract web pages, crawl websites, process multiple URLs, and retrieve structured content using natural language prompts.
Before continuing, make sure you've completed the **Before You Start** guide and have:
* A Geonode API key.
* The Geonode MCP endpoint.
* The Codex desktop application installed.
* Internet access.
***
## Step 1 — Open MCP Settings
Open the **Codex** desktop application.
Click the **Settings** icon in the upper-right corner.
***
## Step 2 — Open the MCP Servers Page
From the Settings menu:
1. Select **Plugins** from the left sidebar.
2. Open the **MCPs** tab.
3. Click **Add server**.
This opens the configuration page for adding a custom MCP server.
***
## Step 3 — Configure the GeoNode MCP Server
Enter the following information:
### Name
You can use any descriptive name, for example:
```text
geonode-server
```
### Type
Select:
```text
Streamable HTTP
```
### URL
Enter the GeoNode MCP endpoint:
```text
https://scraper.geonode.io/mcp
```
### Authentication
Under **Headers**, add the following header:
| Key | Value |
| ----------- | -------------------- |
| `X-API-Key` | Your Geonode API key |
After completing the configuration, click **Save**.
***
## Step 4 — Verify the Connection
After saving the configuration, return to the MCP Servers page.
If the connection is successful, the GeoNode MCP server appears in the server list and is enabled.
If the GeoNode MCP server appears in the server list and is enabled, Codex is successfully connected and ready to use the available MCP tools.
***
## Step 5 — Test the MCP Server
Open a new chat in Codex and ask it to list the available GeoNode MCP tools.
For example:
```text
List all available tools from the GeoNode MCP server.
```
If the integration is configured correctly, Codex will display the available MCP tools exposed by the GeoNode server.
***
## Step 6 — Start Using GeoNode MCP
Once the MCP server is connected, you can interact with the GeoNode Scraper API using natural language.
For example:
### Extract a web page
```text
Extract the content from https://geonode.com
```
### Summarize a page
```text
Read https://geonode.com and summarize the homepage.
```
### Crawl a website
```text
Crawl https://docs.geonode.com and summarize its documentation.
```
### Process multiple URLs
```text
Extract and summarize the following URLs:
https://example.com/page1
https://example.com/page2
https://example.com/page3
```
Codex automatically selects the appropriate GeoNode MCP tool based on your request.
***
## How Codex Uses MCP
When you submit a prompt, Codex analyzes your request to determine whether an MCP tool is required.
If your request involves extracting web content, crawling websites, processing multiple URLs, or retrieving statistics, Codex automatically invokes the appropriate GeoNode MCP tool and returns the results directly in the conversation.
You don't need to manually choose or invoke MCP tools.
***
## FAQs
Verify that:
* The MCP endpoint is correct.
* The `X-API-Key` header contains a valid Geonode API key.
* Your internet connection is active.
After making changes, save the configuration and try connecting again.
Ensure that:
* The server configuration has been saved.
* The server URL is correct.
* The server is enabled from the MCP Servers page.
Verify that:
* The GeoNode MCP server is enabled.
* Your API key is valid.
* The MCP server connection is active.
Then try your prompt again.
Authentication failures usually occur when:
* The API key is incorrect.
* The `X-API-Key` header is missing.
* The API key has expired or been revoked.
Update the API key in the server configuration and save the changes.
Open **Settings → Plugins → MCPs**, select the GeoNode MCP server, and choose **Uninstall** or disable it from the server list.
# Help & FAQ (/docs/scraper-api/guides/mcp/help-and-faq)
import { Accordion, Accordions } from 'fumadocs-ui/components/accordion';
## Troubleshooting
* **No tools appear:** check that Node.js is installed (for `mcp-remote` setups), restart the client fully, and re-check the endpoint URL.
* **401 / authentication errors:** check the `X-Api-Key` value and that the key is active in your Geonode dashboard. Keyless access does not require a key; see [Use Without an API Key](/docs/scraper-api/guides/mcp/01_use-without-an-api-key).
* **Header not recognized:** with `mcp-remote`, keep the format `X-Api-Key:${VAR}` with no space after the colon.
* **Only `extract` appears:** that is expected without an API key. Add `X-Api-Key` to unlock the full tool set.
***
## FAQs
Any MCP-compatible client. Claude Desktop, Cursor, Claude Code, and Windsurf
have direct support; Smithery and Docker MCP help you deploy and manage
connections.
No, not to try. Point your client at `https://scraper.geonode.io/mcp` with no
key to use `extract` under free anonymous limits. Add an API key when you need
more tools, async jobs, JavaScript rendering, or higher limits. See
[Use Without an API Key](/docs/scraper-api/guides/mcp/01_use-without-an-api-key).
Only what you ask it to scrape. With a key, that runs under your Geonode
account and plan limits. Without a key, anonymous rate and volume limits apply.
Not for keyless trial use. For authenticated access you need a Geonode API key
and a plan that covers your usage. On the client side, some assistants require
their own paid tier to turn on MCP.
# Search Workflows (/docs/scraper-api/guides/search/01_search_overview)
The Search API lets you submit a search query and receive search results. Each search returns a unique `job_id`, which can be used to retrieve the full details of a completed search job.
## Complete Search Workflow
The following diagram shows how the Search API endpoints work together.
A search starts with a query and returns search results together with a unique job ID.
***
## Typical Workflow
Most applications follow these steps when working with Search.
### Step 1 — Submit a Search
Start by submitting a search query.
```http
POST /v1/search
```
The request requires a `query` field.
```json
{
"query": "web scraping"
}
```
You can also provide optional search settings such as:
* `locale`
* `page`
* `safe`
* `time_range`
The `query` can contain between 1 and 1,000 characters. The available values and limits for the optional fields are covered in [Search Parameters](/docs/scraper-api/guides/search/03_search_parameters).
***
### Step 2 — Receive the Search Response
A successful search returns a `SearchResponse`.
The response includes:
* `job_id`
* `query`
* `page`
* `attempts`
* `results`
* `suggestions`
* `spelling_correction`
The `job_id` uniquely identifies the search job.
***
### Step 3 — Review the Results
The `results` array contains the search results returned for the requested page.
Each result contains:
* `position`
* `title`
* `url`
It can also include:
* `snippet`
* `thumbnail`
* `displayed_url`
* `source_host`
The `position` field represents the 1-based rank of the result on the page.
***
## Finding Previous Search Jobs
Search responses include a `job_id` that identifies the search.
You can list search jobs for the authenticated user using:
```http
GET /v1/search/jobs
```
The Search Jobs endpoint supports optional filters for:
* Search query
* Job status
* Start date
* End date
* Page
* Page size
The `page_size` can be between 1 and 100 and defaults to 10.
***
## Retrieving Search Job Details
After you have a search `job_id`, retrieve the full details of a completed search job using:
```http
GET /v1/search/{job_id}
```
The response provides the search configuration and its results.
It can include:
* `job_id`
* `query`
* `locale`
* `page`
* `safe`
* `time_range`
* `status`
* `results`
* `results_count`
* `suggestions`
* `spelling_correction`
* `attempts`
* `final_url`
* `duration_ms`
* `tokens_charged`
* `block_reason`
* `error_code`
* `error_message`
* `created_at`
* `completed_at`
The endpoint description specifically defines this operation as retrieving the full details and results for a **completed search job**.
***
## Pagination
The Search API supports result pagination through the `page` request parameter.
The page number must be between **1 and 20**.
For example:
```json
{
"query": "web scraping",
"page": 2
}
```
The response returns the page that was requested in the `page` field.
***
## Search Options
The Search API provides additional request options for controlling the search.
### Locale
Use `locale` to specify a language code, language-region tag, or `all`.
Examples supported by the OpenAPI schema include:
```text
en
en-US
all
```
### Safe Search
Use `safe` to select a safe-search level.
Supported values are:
```text
off
moderate
strict
```
The default is `off`.
### Time Range
Use `time_range` to restrict results to a recency window.
Supported values are:
```text
day
week
month
year
```
### Result Page
Use `page` to request a specific result page from 1 through 20.
These options can be combined with the required `query` field in the same request.
***
## Common Workflow Patterns
### Standard Search
This is the basic Search workflow.
***
### Search with Pagination
Use the `page` parameter when you need to retrieve another result page.
***
### Find and Retrieve a Previous Search
Use this workflow when you need to find a previous search job and retrieve its details.
***
## Best Practices
* Store the `job_id` returned by the Search API if you need to retrieve the search later.
* Use `page` to request additional result pages.
* Use `locale` when you need to specify the search locale.
* Use `safe` when you need a specific safe-search level.
* Use `time_range` when you need to restrict results to a specific recency window.
* Use the Search Jobs endpoint when you need to find previous search jobs.
* Use the Search Job Details endpoint to retrieve the full details and results of a completed search job.
***
## Next Steps
You now understand the main Search API workflow.
Continue with:
* [Getting Started](/docs/scraper-api/guides/search/02_getting_started) — Make your first Search API request.
# Getting Started with Search (/docs/scraper-api/guides/search/02_getting_started)
In this guide, you'll make your first Search API request and see how the API returns search results.
## Before You Begin
Before making a Search request, make sure you have:
* A Geonode API key
* The Scraper API base URL
If you haven't completed the initial setup, see [Quick Start Guide](/docs/scraper-api/quick-start).
## Make Your First Search
Create a search by sending a `POST` request to the Search endpoint.
```http
POST /v1/search
```
For your first search, you only need to provide a search query.
### Example Request
```bash
curl -X POST "https://scraper.geonode.io/v1/search" \
-H "X-Api-Key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"query": "web scraping best practice"
}'
```
You can also send the request body as JSON:
```json
{
"query": "web scraping best practice"
}
```
## Example Response
A successful request returns the search results together with a unique job ID.
```json
{
"job_id": "a478199c-54d2-4884-ba8d-d6d7568614e9",
"query": "web scraping best practice",
"page": 1,
"attempts": 2,
"results": [
{
"position": 1,
"title": "Web Scraping Best Practices in 2026",
"url": "https://www.scrapingbee.com/blog/web-scraping-best-practices/",
"snippet": "Web scraping is the automated process of retrieving data from websites and transforming raw HTML or other web data into structured formats for analysis or use. Whether you are working on a small web scraping project or managing large-scale data collection activities, choosing the right web scraping tool and following best practices is essential. In this article, I'll walk you through the best ...",
"thumbnail": null,
"displayed_url": "www.scrapingbee.com",
"source_host": "www.scrapingbee.com"
}
],
"suggestions": [],
"spelling_correction": null
}
```
The response contains:
| Field | Description |
| --------------------- | ------------------------------------------------ |
| `job_id` | Unique identifier for the search job. |
| `query` | The search query submitted to the API. |
| `page` | The result page returned by the search. |
| `attempts` | Number of attempts made for the search. |
| `results` | Search results returned for the query. |
| `suggestions` | Search suggestions returned by the API. |
| `spelling_correction` | Spelling correction information, when available. |
The `results` array contains the webpages returned for the search query.
You can learn more about the individual result fields in [Understanding Search Results](/docs/scraper-api/guides/search/04_understanding_search_results).
## What Happens Next?
Your first Search request is now complete.
The next guides explain how to customize and work with Search:
* [Search Parameters](/docs/scraper-api/guides/search/03_search_parameters) — Learn how to configure your search request.
# Search Parameters (/docs/scraper-api/guides/search/03_search_parameters)
The Search API provides several parameters to control what results are returned. The `query` parameter is required, while `locale`, `page`, `safe`, and `time_range` are optional.
## query
The `query` parameter contains the search query you want to submit.
It is the only required parameter.
### Requirements
* Type: `string`
* Minimum length: `1`
* Maximum length: `1000`
### Example
```json
{
"query": "web scraping best practice"
}
```
Every Search request must include `query`. The value must contain between 1 and 1,000 characters.
***
## locale
The `locale` parameter specifies the locale for the search.
It accepts:
* A language code, such as `en`
* A language-region tag, such as `en-US`
* `all`
### Example
```json
{
"query": "web scraping",
"locale": "en-US"
}
```
The `locale` parameter is optional.
***
## page
The `page` parameter specifies which result page to fetch.
### Requirements
* Type: `integer`
* Minimum: `1`
* Maximum: `20`
* Default: `1`
### Example
```json
{
"query": "web scraping",
"page": 2
}
```
The OpenAPI defines example values of `1`, `2`, and `20`.
The `page` value must be between `1` and `20`.
***
## safe
The `safe` parameter controls the safe-search filtering level.
The accepted values are:
| Value | Description |
| ---------- | ---------------------------------- |
| `off` | Safe-search filtering is disabled. |
| `moderate` | Moderate safe-search filtering. |
| `strict` | Strict safe-search filtering. |
The default value is `off`.
### Example
```json
{
"query": "web scraping",
"safe": "moderate"
}
```
The OpenAPI only defines these three values for `safe`.
Use only `off`, `moderate`, or `strict` for the `safe` parameter.
***
## time\_range
The `time_range` parameter restricts results to a recency window.
The accepted values are:
| Value | Recency window |
| ------- | -------------- |
| `day` | Day |
| `week` | Week |
| `month` | Month |
| `year` | Year |
### Example
```json
{
"query": "web scraping",
"time_range": "week"
}
```
The parameter is optional and can be set to `day`, `week`, `month`, or `year`.
***
## Using Multiple Parameters
You can combine the optional parameters with the required `query` parameter in a single request.
For example:
```json
{
"query": "web scraping best practice",
"locale": "en-US",
"page": 2,
"safe": "moderate",
"time_range": "month"
}
```
The same parameters can be sent to:
```http
POST /v1/search
```
***
## Complete Request Example
The following request uses all available Search request parameters:
```bash
curl -X POST "YOUR_SCRAPER_API_URL/v1/search" \
-H "X-Api-Key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"query": "web scraping best practice",
"locale": "en-US",
"page": 2,
"safe": "moderate",
"time_range": "month"
}'
```
***
## Parameter Summary
| Parameter | Required | Type | Default | Accepted values / limits |
| ------------ | -------- | --------- | ------- | -------------------------------------------- |
| `query` | Yes | `string` | — | 1–1000 characters |
| `locale` | No | `string` | — | Language code, language-region tag, or `all` |
| `page` | No | `integer` | `1` | `1–20` |
| `safe` | No | `string` | `off` | `off`, `moderate`, `strict` |
| `time_range` | No | `string` | — | `day`, `week`, `month`, `year` |
These are the complete request parameters defined by the current `SearchRequest` schema.
***
## What's Next?
You now know how to configure a Search request.
Continue with [Understanding Search Results](/docs/scraper-api/guides/search/04_understanding_search_results) to learn how the API structures the returned search results.
# Understanding Search Results (/docs/scraper-api/guides/search/04_understanding_search_results)
The Search API returns a `SearchResponse` containing the submitted query, the returned result page, search results, and additional search information.
## Search Response
A successful Search request returns the following top-level fields:
| Field | Description |
| --------------------- | ----------------------------------------- |
| `job_id` | Unique job identifier for this search. |
| `query` | Search query that was submitted. |
| `page` | Result page that was returned. |
| `attempts` | Number of upstream attempts made. |
| `results` | Search results returned by the API. |
| `suggestions` | Query suggestions returned by the engine. |
| `spelling_correction` | Spelling correction applied to the query. |
The OpenAPI defines `job_id`, `query`, `page`, `attempts`, `results`, and `suggestions` as required response fields. `spelling_correction` can be a string or `null`.
## Example Response
The following is a real response from the Search API:
```json
{
"job_id": "a478199c-54d2-4884-ba8d-d6d7568614e9",
"query": "web scraping best practice",
"page": 1,
"attempts": 2,
"results": [
{
"position": 1,
"title": "Web Scraping Best Practices in 2026",
"url": "https://www.scrapingbee.com/blog/web-scraping-best-practices/",
"snippet": "Web scraping is the automated process of retrieving data from websites and transforming raw HTML or other web data into structured formats for analysis or use. Whether you are working on a small web scraping project or managing large-scale data collection activities, choosing the right web scraping tool and following best practices is essential. In this article, I'll walk you through the best ...",
"thumbnail": null,
"displayed_url": "www.scrapingbee.com",
"source_host": "www.scrapingbee.com"
}
],
"suggestions": [],
"spelling_correction": null
}
```
The response above uses the fields defined by the `SearchResponse` and `SearchHitModel` schemas.
## Search Results
The `results` field is an array of `SearchHitModel` objects.
Each search result contains three required fields:
* `position`
* `title`
* `url`
It can also contain:
* `snippet`
* `thumbnail`
* `displayed_url`
* `source_host`
### Result Fields
| Field | Required | Description |
| --------------- | -------- | ------------------------------------------ |
| `position` | Yes | 1-based rank of this result on the page. |
| `title` | Yes | Result title. |
| `url` | Yes | Result target URL. |
| `snippet` | No | Result preview text. |
| `thumbnail` | No | Thumbnail image URL, if available. |
| `displayed_url` | No | Human-readable URL as shown by the engine. |
| `source_host` | No | Hostname the result was served from. |
The OpenAPI defines `position`, `title`, and `url` as the required fields for each result. `thumbnail`, `displayed_url`, and `source_host` can be `null`.
## Working With a Result
A result can be accessed from the `results` array.
For example, the first result in the response is:
```json
{
"position": 1,
"title": "Web Scraping Best Practices in 2026",
"url": "https://www.scrapingbee.com/blog/web-scraping-best-practices/",
"snippet": "Web scraping is the automated process of retrieving data from websites and transforming raw HTML or other web data into structured formats for analysis or use.",
"thumbnail": null,
"displayed_url": "www.scrapingbee.com",
"source_host": "www.scrapingbee.com"
}
```
The `position` identifies the result's 1-based rank on the returned page.
The `title` and `url` identify the result, while the remaining fields provide additional result information when available.
## Suggestions
The `suggestions` field contains query suggestions returned by the engine.
In the example response, no suggestions were returned:
```json
{
"suggestions": []
}
```
The OpenAPI defines `suggestions` as an array of strings.
## Spelling Correction
The `spelling_correction` field contains the spelling correction applied to the query.
It can contain a string or `null`.
In the example response:
```json
{
"spelling_correction": null
}
```
The OpenAPI defines this field as either a string or `null`.
## Complete Response Structure
The Search response can be represented by the following structure:
```text
SearchResponse
├── job_id
├── query
├── page
├── attempts
├── results[]
│ ├── position
│ ├── title
│ ├── url
│ ├── snippet
│ ├── thumbnail
│ ├── displayed_url
│ └── source_host
├── suggestions[]
└── spelling_correction
```
This structure follows the `SearchResponse` and `SearchHitModel` schemas defined in the OpenAPI specification.
## What's Next?
You now understand the structure of a Search API response.
Continue with [Pagination and Filters](/docs/scraper-api/guides/search/05_pagination_and_filters) to learn how to work with result pages and the available Search request filters.
# Pagination and Filters (/docs/scraper-api/guides/search/05_pagination_and_filters)
The Search API provides pagination and optional filters that you can include in the search request.
You can use `page` to request a specific result page and use `locale`, `safe`, and `time_range` to configure the search request.
## Pagination
Use the `page` parameter to specify which result page to fetch.
The OpenAPI defines the following limits:
* Minimum: `1`
* Maximum: `20`
* Default: `1`
For example, to request the second result page:
```json
{
"query": "web scraping best practice",
"page": 2
}
```
The `page` value must be an integer between `1` and `20`.
### Request a Specific Page
You can request any page from `1` through `20`.
```bash
curl -X POST "YOUR_SCRAPER_API_URL/v1/search" \
-H "X-Api-Key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"query": "web scraping best practice",
"page": 2
}'
```
The returned `page` field identifies the result page that was returned.
***
## Locale
Use the `locale` parameter to specify the locale for the search.
The OpenAPI accepts:
* A language code, such as `en`
* A language-region tag, such as `en-US`
* `all`
For example:
```json
{
"query": "web scraping",
"locale": "en-US"
}
```
The `locale` parameter is optional.
***
## Safe Search
Use the `safe` parameter to specify the safe-search level.
The supported values are:
| Value | Description |
| ---------- | ------------------------------------ |
| `off` | Safe-search level set to `off`. |
| `moderate` | Safe-search level set to `moderate`. |
| `strict` | Safe-search level set to `strict`. |
The default value is `off`.
### Example
```json
{
"query": "web scraping",
"safe": "moderate"
}
```
Only `off`, `moderate`, and `strict` are accepted values for this parameter.
***
## Time Range
Use the `time_range` parameter to restrict results to a recency window.
The supported values are:
| Value |
| ------- |
| `day` |
| `week` |
| `month` |
| `year` |
For example:
```json
{
"query": "web scraping",
"time_range": "week"
}
```
The `time_range` parameter is optional.
***
## Combining Parameters
You can include multiple optional parameters in the same Search request.
For example:
```json
{
"query": "web scraping best practice",
"locale": "en-US",
"page": 2,
"safe": "moderate",
"time_range": "month"
}
```
This request uses:
* `locale` to specify `en-US`
* `page` to request page `2`
* `safe` with the `moderate` value
* `time_range` with the `month` value
All four parameters are defined as optional in the `SearchRequest` schema.
***
## Complete Request
The following example combines all available Search request options:
```bash
curl -X POST "YOUR_SCRAPER_API_URL/v1/search" \
-H "X-Api-Key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"query": "web scraping best practice",
"locale": "en-US",
"page": 2,
"safe": "moderate",
"time_range": "month"
}'
```
### Parameter Summary
| Parameter | Required | Type | Default | Values / Limits |
| ------------ | -------- | ------------------ | ------- | -------------------------------------------- |
| `query` | Yes | `string` | — | 1–1000 characters |
| `locale` | No | `string` or `null` | — | Language code, language-region tag, or `all` |
| `page` | No | `integer` | `1` | `1–20` |
| `safe` | No | `string` | `off` | `off`, `moderate`, `strict` |
| `time_range` | No | `string` or `null` | — | `day`, `week`, `month`, `year` |
These are the request parameters defined by the current `SearchRequest` schema.
***
## What's Next?
You now know how to paginate Search results and configure the available Search filters.
Continue with [Search Jobs](/docs/scraper-api/guides/search/06_search_jobs) to learn how to list and retrieve search jobs.
# Search Jobs (/docs/scraper-api/guides/search/06_search_jobs)
Every Search request returns a unique `job_id`. You can use this ID to retrieve the details of a completed search job, or use the Search Jobs endpoint to find previous searches.
## Search Job Workflow
The following diagram shows how the Search Job endpoints work together.
## List Search Jobs
Use the Search Jobs endpoint to list search jobs for the authenticated user.
```http
GET /v1/search/jobs
```
The endpoint supports optional filtering and pagination.
### Example Request
```bash
curl "YOUR_SCRAPER_API_URL/v1/search/jobs" \
-H "X-Api-Key: YOUR_API_KEY"
```
### Example Response
The following is a real response from the Search API:
```json
{
"jobs": [
{
"job_id": "a478199c-54d2-4884-ba8d-d6d7568614e9",
"query": "web scraping best practice",
"status": "completed",
"results_count": 10,
"duration_ms": 1800,
"block_reason": null,
"error_code": null,
"created_at": "2026-08-16T17:44:07.447657Z",
"completed_at": "2026-08-16T17:44:09.673985Z"
},
{
"job_id": "ed9f8595-be2a-4d3e-873b-1a1c52620a0e",
"query": "https://scraper.geonode.io",
"status": "completed",
"results_count": 10,
"duration_ms": 5077,
"block_reason": null,
"error_code": null,
"created_at": "2026-08-16T17:43:29.525681Z",
"completed_at": "2026-08-16T17:43:35.053054Z"
}
],
"page": 1,
"page_size": 10,
"page_count": 1
}
```
The response contains:
| Field | Description |
| ------------ | --------------------------------- |
| `jobs` | Search jobs on the current page. |
| `page` | Current page number. |
| `page_size` | Number of jobs returned per page. |
| `page_count` | Total number of pages available. |
These fields are defined by the `SearchJobsResponse` schema.
## Search Job Fields
Each item in the `jobs` array is a `SearchListItemResponse`.
The available fields are:
| Field | Description |
| --------------- | ---------------------------------------------- |
| `job_id` | Unique job identifier. |
| `query` | Search query that was submitted. |
| `status` | Job status. |
| `results_count` | Number of results returned. |
| `duration_ms` | Execution time in milliseconds. |
| `block_reason` | Upstream block classification for failed jobs. |
| `error_code` | Machine-readable error code for failed jobs. |
| `created_at` | Time when the search job was created. |
| `completed_at` | Time when the search job finished. |
The OpenAPI defines `job_id`, `query`, `status`, and `created_at` as required fields. The remaining fields can be absent or nullable according to the schema.
## Filter Search Jobs
You can filter the jobs returned by `GET /v1/search/jobs`.
The available filters are:
| Parameter | Description |
| ------------ | --------------------------------------------- |
| `query` | Filter by search query using a partial match. |
| `status` | Filter by job status. |
| `start_date` | Filter jobs created on or after this date. |
| `end_date` | Filter jobs created on or before this date. |
For example, to filter jobs by query:
```bash
curl "YOUR_SCRAPER_API_URL/v1/search/jobs?query=web%20scraping" \
-H "X-Api-Key: YOUR_API_KEY"
```
To filter by status:
```bash
curl "YOUR_SCRAPER_API_URL/v1/search/jobs?status=completed" \
-H "X-Api-Key: YOUR_API_KEY"
```
The OpenAPI defines these four filters for the Search Jobs endpoint.
## Paginate Search Jobs
The Search Jobs endpoint supports pagination using:
* `page`
* `page_size`
`page` specifies the page number and defaults to `1`.
`page_size` specifies the number of results per page. It must be between `1` and `100` and defaults to `10`.
For example:
```bash
curl "YOUR_SCRAPER_API_URL/v1/search/jobs?page=1&page_size=20" \
-H "X-Api-Key: YOUR_API_KEY"
```
The response includes `page`, `page_size`, and `page_count` so you can determine the current page and the total number of available pages.
The `page_size` parameter accepts values from `1` through `100`.
## Retrieve Search Job Details
Once you have a `job_id`, use the Search Job endpoint to retrieve the full details and results for a completed search job.
```http
GET /v1/search/{job_id}
```
For example:
```bash
curl "YOUR_SCRAPER_API_URL/v1/search/a478199c-54d2-4884-ba8d-d6d7568614e9" \
-H "X-Api-Key: YOUR_API_KEY"
```
### Example Response
The following is a real response for the search job used in the examples above:
```json
{
"job_id": "a478199c-54d2-4884-ba8d-d6d7568614e9",
"query": "web scraping best practice",
"locale": null,
"page": 1,
"safe": "off",
"time_range": null,
"status": "completed",
"results": [
{
"position": 1,
"title": "Web Scraping Best Practices in 2026",
"url": "https://www.scrapingbee.com/blog/web-scraping-best-practices/",
"snippet": "Web scraping is the automated process of retrieving data from websites and transforming raw HTML or other web data into structured formats for analysis or use. Whether you are working on a small web scraping project or managing large-scale data collection activities, choosing the right web scraping tool and following best practices is essential. In this article, I'll walk you through the best ...",
"thumbnail": null,
"displayed_url": "www.scrapingbee.com",
"source_host": "www.scrapingbee.com"
},
{
"position": 2,
"title": "10 Best Sample Websites for Web Scraping Practice in 2026",
"url": "https://thunderbit.com/blog/best-web-scraping-test-sites",
"snippet": "Practice web scraping on the best sample sites for all skill levels. Thunderbit helps automate extraction, handle complex sites, and streamline data workflows.",
"thumbnail": null,
"displayed_url": "thunderbit.com",
"source_host": "thunderbit.com"
},
{
"position": 3,
"title": "11 Web Scraping Best Practices for Reliable Data Collection",
"url": "https://scrapfly.io/blog/posts/web-scraping-best-practices",
"snippet": "11 web scraping best practices for 2026: robots.txt, rate limiting, hidden APIs, proxy rotation, retries, validation, and monitoring, with working code.",
"thumbnail": null,
"displayed_url": "scrapfly.io",
"source_host": "scrapfly.io"
}
],
"results_count": 10,
"suggestions": [],
"spelling_correction": null,
"attempts": 2,
"final_url": "http://searxng:8080/search?q=web+scraping+best+practice&format=json&engines=duckduckgo&pageno=1&safesearch=0",
"duration_ms": 1800,
"tokens_charged": 1,
"block_reason": null,
"error_code": null,
"error_message": null,
"created_at": "2026-08-16T17:44:07.447657Z",
"completed_at": "2026-08-16T17:44:09.673985Z"
}
```
The `GET /v1/search/{job_id}` endpoint is documented as retrieving the full details and results for a completed search job.
## Search Job Details
The detailed response contains the search configuration, status, results, and execution information.
| Field | Description |
| --------------------- | ---------------------------------------------- |
| `job_id` | Unique job identifier. |
| `query` | Search query that was submitted. |
| `locale` | Locale requested for the search. |
| `page` | Result page that was requested. |
| `safe` | Safe-search level requested. |
| `time_range` | Time range filter, if requested. |
| `status` | Job status. |
| `results` | Search results. |
| `results_count` | Number of results returned. |
| `suggestions` | Query suggestions returned by the engine. |
| `spelling_correction` | Spelling correction applied to the query. |
| `attempts` | Number of upstream attempts made. |
| `final_url` | Final upstream URL that produced the results. |
| `duration_ms` | Execution time in milliseconds. |
| `tokens_charged` | Tokens charged for the job. |
| `block_reason` | Upstream block classification for failed jobs. |
| `error_code` | Machine-readable error code. |
| `error_message` | Human-readable error description. |
| `created_at` | Time when the search job was created. |
| `completed_at` | Time when the search job finished. |
These fields are defined by `SearchJobDetailResponse` in the OpenAPI.
## Job Status
Search jobs use the `JobStatus` schema.
The available statuses are:
```text
queued
processing
completed
failed
cancelled
```
The Search Jobs response uses this status field for each listed search job.
## Complete Search Job Workflow
## What's Next?
You now know how to list Search jobs, filter and paginate the job list, and retrieve the details of a completed search job.
Continue with [Search Errors](/docs/scraper-api/guides/search/07_search_errors) to learn about the Search API error responses.
# Search Errors (/docs/scraper-api/guides/search/07_search_errors)
The Search API can return different error responses when a request cannot be completed.
The response format depends on the type of error. The OpenAPI defines `SearchErrorResponse` for Search-specific errors and `ErrorResponse` for other API errors.
## Search Error Response
Search-specific errors use the following response structure:
```json
{
"error": "error_code",
"message": "Human-readable error description",
"details": {}
}
```
The `error` and `message` fields are required. The `details` field is optional and can contain additional error context.
### Error Fields
| Field | Description |
| --------- | ----------------------------------------- |
| `error` | Machine-readable error code. |
| `message` | Human-readable error description. |
| `details` | Additional error context, when available. |
The OpenAPI gives `no_results` and `upstream_blocked` as examples of the `error` field.
***
## Search-Specific Errors
The Search endpoint documents two responses that use `SearchErrorResponse`.
### `422` — Unprocessable Entity
The Search endpoint can return `422` with a `SearchErrorResponse`.
```http
422 Unprocessable Entity
```
This response uses the following structure:
```json
{
"error": "no_results",
"message": "Human-readable error description",
"details": {}
}
```
The OpenAPI does not define a fixed message or `details` structure for this response, so use the values returned by the API.
### `502` — Bad Gateway
The Search endpoint can return `502` with a `SearchErrorResponse`.
```http
502 Bad Gateway
```
The response uses the same Search-specific error structure:
```json
{
"error": "upstream_blocked",
"message": "Human-readable error description",
"details": {}
}
```
The OpenAPI documents `upstream_blocked` as an example machine-readable error code.
The exact `error`, `message`, and `details` values depend on the response returned by the API. Do not assume a specific message or details object.
***
## Other Search API Errors
The Search endpoint also documents the following HTTP responses:
| Status | Description | Response |
| ------ | ------------------------------- | ---------------------------- |
| `401` | Unauthorized | `ErrorResponse` |
| `402` | Insufficient token balance | `ErrorResponse` |
| `408` | Request timed out | `ErrorResponse` |
| `422` | Unprocessable Entity | `SearchErrorResponse` |
| `429` | Request throttled | No response schema specified |
| `500` | Internal server error | `ErrorResponse` |
| `502` | Bad Gateway | `SearchErrorResponse` |
| `503` | Service temporarily unavailable | `ErrorResponse` |
These responses and their associated schemas are defined on `POST /v1/search` in the OpenAPI specification.
***
## Handle Errors in Your Application
Check the HTTP status code before processing a successful Search response.
For Search-specific errors, inspect the returned `error`, `message`, and `details` fields.
For example:
```python
response = requests.post(
"YOUR_SCRAPER_API_URL/v1/search",
headers={
"X-Api-Key": "YOUR_API_KEY",
"Content-Type": "application/json"
},
json={
"query": "web scraping best practice"
}
)
if response.status_code == 200:
data = response.json()
else:
error = response.json()
print(error)
```
The exact error fields available depend on the response documented for that HTTP status.
***
## Search Error Summary
The Search endpoint documents these HTTP responses:
```text
200 Successful Response
401 Unauthorized
402 Insufficient token balance
408 Request timed out
422 Unprocessable Entity
429 Request throttled
500 Internal server error
502 Bad Gateway
503 Service temporarily unavailable
```
The `422` and `502` responses use `SearchErrorResponse`. The other documented error responses use `ErrorResponse`, except `429`, for which the OpenAPI does not specify a response schema.
***
## What's Next?
You now know the error responses documented for the Search API.
For the complete request and response definitions, see the [Search API Reference](/docs/scraper-api/v1/reference).
# Extract a JavaScript-Rendered Product Page (/docs/scraper-api/real-world/ecommerce/01_javascript_rendered_product_page)
This real-world example shows how to extract product content from a JavaScript-rendered e-commerce page using the Scraper API.
We will use the ScrapingCourse JavaScript Rendering page and compare the same extraction with JavaScript rendering disabled and enabled.
## Use Case
Modern websites often use JavaScript to render or load content after the initial page response.
For this example, we use:
```text
https://www.scrapingcourse.com/javascript-rendering
```
The page contains multiple products with their names, prices, and product links.
The extracted content includes products such as:
* Chaz Kangeroo Hoodie — $52
* Teton Pullover Hoodie — $70
* Bruno Compete Hoodie — $63
* Frankie Sweatshirt — $60
* Hollister Backyard Sweatshirt — $52
* Stark Fundamental Hoodie — $42
* Hero Hoodie — $54
* Oslo Trek Hoodie — $42
* Abominable Hoodie — $69
* Mach Street Sweatshirt — $62
* Grayson Crewneck Sweatshirt — $64
* Ajax Full-Zip Sweatshirt — $69
## Before You Start
You need:
* A Geonode API key
* Python 3.9 or later
* The `requests` package
Store your API key in an environment variable:
```bash
export GEONODE_SCRAPER_API_KEY="YOUR_API_KEY"
```
On Windows PowerShell:
```powershell
$env:GEONODE_SCRAPER_API_KEY="YOUR_API_KEY"
```
Never put your Geonode API key directly in your source code or commit it to your repository. Store it in an environment variable or `.env` file instead.
## Install the Dependency
Install the `requests` package:
```bash
pip install requests
```
## Case 1: Extract Without JavaScript Rendering
First, send the request with JavaScript rendering disabled.
```python
import os
import requests
api_key = os.environ["GEONODE_SCRAPER_API_KEY"]
response = requests.post(
"https://scraper.geonode.io/v1/extract",
headers={
"X-Api-Key": api_key,
"Content-Type": "application/json",
},
json={
"url": "https://www.scrapingcourse.com/javascript-rendering",
"formats": ["markdown"],
"render_js": False,
"processing_mode": "sync",
},
)
response.raise_for_status()
result = response.json()
print(result["data"]["markdown"])
```
### Case 1 Result
The actual extraction returned the following Markdown:
```markdown
---
meta-viewport: width=device-width, initial-scale=1.0
title: JS Rendering Challenge to Learn Web Scraping - ScrapingCourse.com
---
# JS Rendering

## Challenge
Enable JavaScript to see products
[Chaz Kangeroo Hoodie Chaz Kangeroo Hoodie $52](https://scrapingcourse.com/ecommerce/product/chaz-kangeroo-hoodie)
[Teton Pullover Hoodie Teton Pullover Hoodie $70](https://scrapingcourse.com/ecommerce/product/teton-pullover-hoodie)
[Bruno Compete Hoodie Bruno Compete Hoodie $63](https://scrapingcourse.com/ecommerce/product/bruno-compete-hoodie)
[Frankie Sweatshirt Frankie Sweatshirt $60](https://scrapingcourse.com/ecommerce/product/frankie-sweatshirt)
[Hollister Backyard Sweatshirt Hollister Backyard Sweatshirt $52](https://scrapingcourse.com/ecommerce/product/hollister-backyard-sweatshirt)
[Stark Fundamental Hoodie Stark Fundamental Hoodie $42](https://scrapingcourse.com/ecommerce/product/stark-fundamental-hoodie)
[Hero Hoodie Hero Hoodie $54](https://scrapingcourse.com/ecommerce/product/hero-hoodie)
[Oslo Trek Hoodie Oslo Trek Hoodie $42](https://scrapingcourse.com/ecommerce/product/oslo-trek-hoodie)
[Abominable Hoodie Abominable Hoodie $69](https://scrapingcourse.com/ecommerce/product/abominable-hoodie)
[Mach Street Sweatshirt Mach Street Sweatshirt $62](https://scrapingcourse.com/ecommerce/product/mach-street-sweatshirt)
[Grayson Crewneck Sweatshirt Grayson Crewneck Sweatshirt $64](https://scrapingcourse.com/ecommerce/product/grayson-crewneck-sweatshirt)
[Ajax Full-Zip Sweatshirt Ajax Full-Zip Sweatshirt $69](https://scrapingcourse.com/ecommerce/product/ajax-full-zip-sweatshirt)
```
## Case 2: Extract With JavaScript Rendering
Now run the same extraction with JavaScript rendering enabled.
The only change is:
```json
"render_js": true
```
```python
import os
import requests
api_key = os.environ["GEONODE_SCRAPER_API_KEY"]
response = requests.post(
"https://scraper.geonode.io/v1/extract",
headers={
"X-Api-Key": api_key,
"Content-Type": "application/json",
},
json={
"url": "https://www.scrapingcourse.com/javascript-rendering",
"formats": ["markdown"],
"render_js": True,
"processing_mode": "sync",
},
)
response.raise_for_status()
result = response.json()
print(result["data"]["markdown"])
```
### Case 2 Result
The actual extraction returned the following Markdown:
```markdown
---
meta-viewport: width=device-width, initial-scale=1.0
title: JS Rendering Challenge to Learn Web Scraping - ScrapingCourse.com
---
# JS Rendering

## Challenge
Enable JavaScript to see products
[Chaz Kangeroo Hoodie Chaz Kangeroo Hoodie $52](https://scrapingcourse.com/ecommerce/product/chaz-kangeroo-hoodie)
[Teton Pullover Hoodie Teton Pullover Hoodie $70](https://scrapingcourse.com/ecommerce/product/teton-pullover-hoodie)
[Bruno Compete Hoodie Bruno Compete Hoodie $63](https://scrapingcourse.com/ecommerce/product/bruno-compete-hoodie)
[Frankie Sweatshirt Frankie Sweatshirt $60](https://scrapingcourse.com/ecommerce/product/frankie-sweatshirt)
[Hollister Backyard Sweatshirt Hollister Backyard Sweatshirt $52](https://scrapingcourse.com/ecommerce/product/hollister-backyard-sweatshirt)
[Stark Fundamental Hoodie Stark Fundamental Hoodie $42](https://scrapingcourse.com/ecommerce/product/stark-fundamental-hoodie)
[Hero Hoodie Hero Hoodie $54](https://scrapingcourse.com/ecommerce/product/hero-hoodie)
[Oslo Trek Hoodie Oslo Trek Hoodie $42](https://scrapingcourse.com/ecommerce/product/oslo-trek-hoodie)
[Abominable Hoodie Abominable Hoodie $69](https://scrapingcourse.com/ecommerce/product/abominable-hoodie)
[Mach Street Sweatshirt Mach Street Sweatshirt $62](https://scrapingcourse.com/ecommerce/product/mach-street-sweatshirt)
[Grayson Crewneck Sweatshirt Grayson Crewneck Sweatshirt $64](https://scrapingcourse.com/ecommerce/product/grayson-crewneck-sweatshirt)
[Ajax Full-Zip Sweatshirt Ajax Full-Zip Sweatshirt $69](https://scrapingcourse.com/ecommerce/product/ajax-full-zip-sweatshirt)
```
## Compare the Results
You can compare both requests directly:
### Request
```json
{
"url": "https://www.scrapingcourse.com/javascript-rendering",
"formats": ["markdown"],
"render_js": false,
"processing_mode": "sync"
}
```
### Result
The captured result contains the page heading, the JavaScript challenge message, and the 12 product links with their prices.
### Request
```json
{
"url": "https://www.scrapingcourse.com/javascript-rendering",
"formats": ["markdown"],
"render_js": true,
"processing_mode": "sync"
}
```
### Result
The captured result also contains the page heading, the JavaScript challenge message, and the 12 product links with their prices.
For this particular target and the captured runs, both cases returned the same Markdown content. The example therefore demonstrates how to enable JavaScript rendering, but the captured result does not establish that `render_js` is required for this page.
## Extracted Products
The returned Markdown contains 12 product entries.
| Product | Price |
| ----------------------------- | ----: |
| Chaz Kangeroo Hoodie | $52 |
| Teton Pullover Hoodie | $70 |
| Bruno Compete Hoodie | $63 |
| Frankie Sweatshirt | $60 |
| Hollister Backyard Sweatshirt | $52 |
| Stark Fundamental Hoodie | $42 |
| Hero Hoodie | $54 |
| Oslo Trek Hoodie | $42 |
| Abominable Hoodie | $69 |
| Mach Street Sweatshirt | $62 |
| Grayson Crewneck Sweatshirt | $64 |
| Ajax Full-Zip Sweatshirt | $69 |
## Understanding the Response
The extracted page content is available under:
```text
data.markdown
```
You can use the returned Markdown directly or process it further in your application.
For example:
```python
markdown = result["data"]["markdown"]
print(markdown)
```
The Scraper API returns page content rather than automatically converting the page into structured product fields.
If your application needs fields such as:
```text
name
price
url
sku
availability
```
you can parse the returned Markdown or HTML in your own application.
The Scraper API returns extracted page content as Markdown or HTML. It does not automatically return e-commerce fields such as product name, price, or SKU as structured fields.
## When to Use JavaScript Rendering
Use:
```json
"render_js": true
```
when the content you need depends on JavaScript execution.
Common examples include:
* Product information loaded after the page opens
* Client-side rendered product listings
* Content populated dynamically by JavaScript
* Pages where the initial HTML does not contain the content you need
For pages where the required content is already available without browser rendering, you can use:
```json
"render_js": false
```
This avoids browser rendering when it is not needed.
## What About `wait_config`?
This example does not include `wait_config`.
The focused Case 2 probe successfully completed with:
```json
{
"render_js": true,
"processing_mode": "sync",
"wait_config": null
}
```
It reported 12 product links and 12 prices in the returned Markdown.
If a JavaScript-rendered page requires an explicit browser wait condition, you can use `wait_config` to control when the browser should consider the page ready for extraction.
## Complete Example
Once you understand the difference between the two requests, the JavaScript-rendered version can be reduced to this:
```python
import os
import requests
API_KEY = os.environ["GEONODE_SCRAPER_API_KEY"]
URL = "https://www.scrapingcourse.com/javascript-rendering"
response = requests.post(
"https://scraper.geonode.io/v1/extract",
headers={
"X-Api-Key": API_KEY,
"Content-Type": "application/json",
},
json={
"url": URL,
"formats": ["markdown"],
"render_js": True,
"processing_mode": "sync",
},
)
response.raise_for_status()
result = response.json()
print(result["data"]["markdown"])
```
## Request Parameters
| Parameter | Value | Purpose |
| ----------------- | ----------------------------------------------------- | ----------------------------------------- |
| `url` | `https://www.scrapingcourse.com/javascript-rendering` | Target page to extract |
| `formats` | `["markdown"]` | Returns the extracted content as Markdown |
| `render_js` | `true` | Enables JavaScript rendering |
| `processing_mode` | `sync` | Waits for the extraction to complete |
## Next Steps
Now that you can extract a page with JavaScript rendering, continue to **Compare a Product Page Across Locations** to extract the same product page from different countries with proxy geo-targeting.
You can also explore:
* Extract multiple product URLs with Batch
* Discover product URLs with Map and extract them with Batch
* Crawl a product or content section
* Run longer extractions asynchronously
# Compare a Product Page Across Locations (/docs/scraper-api/real-world/ecommerce/02_geo_targeted_product_page)
Retail sites often change what they show based on where the request comes from. That is useful when you are checking localized messaging, market selectors, shipping copy, or regional pricing — but it is hard to do from a single office IP.
This example uses the Scraper API to fetch the **same product URL** twice: once through a United States residential proxy, and once through a United Kingdom residential proxy. Then we compare the Markdown.
## What you will get
By the end, you will have:
* Two Markdown extracts of the same Nike product page
* Confirmed `metadata.proxy.country` values for each run
* A simple side-by-side check for location-specific content
## Target page
```text
https://www.nike.com/t/air-force-1-07-mens-shoes-jBrhbr/CW2288-111
```
Why this page works well for a demo:
* It is a real public product detail page
* It responds differently by exit country
* In our captured UK run, Nike surfaced an explicit location banner
## Setup
You need a Geonode API key and Python with `requests` installed.
```bash
pip install requests
export GEONODE_SCRAPER_API_KEY="YOUR_API_KEY"
```
PowerShell:
```powershell
pip install requests
$env:GEONODE_SCRAPER_API_KEY="YOUR_API_KEY"
```
Do not hardcode the key in source or commit it to git. Use an environment variable or a local `.env` file.
## The idea
Keep everything identical except the country:
```json
{
"url": "https://www.nike.com/t/air-force-1-07-mens-shoes-jBrhbr/CW2288-111",
"formats": ["markdown"],
"render_js": true,
"processing_mode": "sync",
"proxy": {
"country": "US",
"type": "residential"
}
}
```
Then repeat the same request with `"country": "GB"`.
That isolates geo-targeting as the only intentional variable.
| Field | Why it is set |
| ------------------------- | ---------------------------------------- |
| `formats: ["markdown"]` | Easy to read and compare |
| `render_js: true` | Nike's PDP is JS-heavy |
| `processing_mode: sync` | Wait for the full result in one response |
| `proxy.country` | Force US or GB exit |
| `proxy.type: residential` | Better fit for consumer retail sites |
## Implementation
```python
import json
import os
from pathlib import Path
import requests
API_KEY = os.environ["GEONODE_SCRAPER_API_KEY"]
URL = "https://www.nike.com/t/air-force-1-07-mens-shoes-jBrhbr/CW2288-111"
COUNTRIES = ["US", "GB"]
OUT = Path("output")
OUT.mkdir(exist_ok=True)
def extract(country: str) -> dict:
response = requests.post(
"https://scraper.geonode.io/v1/extract",
headers={
"X-Api-Key": API_KEY,
"Content-Type": "application/json",
},
json={
"url": URL,
"formats": ["markdown"],
"render_js": True,
"processing_mode": "sync",
"proxy": {
"country": country,
"type": "residential",
},
},
timeout=300,
)
response.raise_for_status()
return response.json()
results = {}
for country in COUNTRIES:
result = extract(country)
markdown = result["data"]["markdown"]
path = OUT / f"nike_af1_{country}.md"
path.write_text(markdown, encoding="utf-8")
results[country] = {
"proxy": result["metadata"]["proxy"],
"tokens_charged": result["tokens_charged"],
"markdown_length": len(markdown),
"has_uk_banner": "We think you are in United Kingdom" in markdown,
"path": str(path),
}
print(country, results[country])
us = (OUT / "nike_af1_US.md").read_text(encoding="utf-8")
gb = (OUT / "nike_af1_GB.md").read_text(encoding="utf-8")
summary = {
"target_url": URL,
"markdown_differs": us != gb,
"results": results,
}
(OUT / "geo_comparison.json").write_text(
json.dumps(summary, indent=2),
encoding="utf-8",
)
print("markdown_differs:", summary["markdown_differs"])
```
Run it:
```bash
python geo_targeted_extract.py
```
You should get:
* `output/nike_af1_US.md`
* `output/nike_af1_GB.md`
* `output/geo_comparison.json`
## What you get in the response
A successful sync extract looks like this shape:
```json
{
"data": {
"markdown": "# Nike Air Force 1 '07\n\n$115\n..."
},
"metadata": {
"url": "https://www.nike.com/t/air-force-1-07-mens-shoes-jBrhbr/CW2288-111",
"render_js": true,
"http_status": 200,
"duration_ms": 9069,
"formats": ["markdown"],
"proxy": {
"country": "US",
"type": "residential"
},
"processing_mode": "sync",
"headers": {},
"wait_config": null
},
"tokens_charged": 1
}
```
You mainly care about:
| Field | Meaning |
| ---------------------- | ------------------------------------ |
| `data.markdown` | Clean page content for that country |
| `metadata.proxy` | Country and proxy type actually used |
| `metadata.http_status` | Status from the target site |
| `metadata.duration_ms` | How long the extract took |
| `tokens_charged` | Tokens billed for the request |
You do **not** get structured product fields such as `name`, `price`, or `sku`. Those must be parsed from `data.markdown` if you need them.
## Captured US vs GB results
From the live runs used for this example:
| | US | GB |
| ------------------------ | ------------- | ------------- |
| HTTP status | `200` | `200` |
| `tokens_charged` | `1` | `1` |
| `metadata.proxy.country` | `US` | `GB` |
| `metadata.proxy.type` | `residential` | `residential` |
| `metadata.render_js` | `true` | `true` |
| `metadata.duration_ms` | `9069` | `13864` |
| Markdown length | `67817` | `42453` |
| Markdown differs | yes | yes |
Both returned the product:
```markdown
# Nike Air Force 1 '07
$115
```
The UK extract also included this location banner:
```markdown
# We think you are in United Kingdom. Update your location?
```
That banner was not present in the captured US extract.
Geo-targeting does not guarantee a currency change. In this captured run, both markets still showed `$115`. The useful proof is that the page reacted to exit country — here via Nike's location banner and different Markdown overall.
## How to judge success
1. **Proxy applied** — `metadata.proxy.country` matches what you requested
2. **Content extracted** — `data.markdown` contains the product page
3. **Meaningful difference** — the two Markdown files differ, or one country shows a clear market signal
Do not judge only on price. Banners, market pickers, shipping text, and locale strings all count.
## When to use this pattern
Use geo-targeted extract when location changes the page you care about:
* Market localization QA
* Competitor monitoring by region
* Shipping / availability messaging checks
* Detecting country-specific offers or banners
If location does not matter, omit `proxy.country` and let the API use default routing.
## Next
Once this pattern is working, natural follow-ons are:
* Batch-extract a list of SKUs from one country
* Map a storefront, filter product URLs, then batch-extract them
* Combine `proxy.country` with custom headers such as `Accept-Language`
# Cancel a Batch Job (/docs/scraper-api/v1/batch/cancel-batch-job)
Stop scheduling new batch items and let in-flight child extractions drain
# Get Batch Job Status (/docs/scraper-api/v1/batch/get-batch-job-status)
Poll for the current status and partial results of a batch job
# List batch jobs (/docs/scraper-api/v1/batch/list-batch-job)
List batch jobs for the authenticated user with optional filtering and pagination
# Start a Batch Job (/docs/scraper-api/v1/batch/start-batch-job)
Queue a batch of URLs for asynchronous extraction
# Cancel a Crawl Job (/docs/scraper-api/v1/crawl/cancel-crawl-job)
Stop scheduling new crawl pages and let in-flight page work drain
# Get Crawl Job Status (/docs/scraper-api/v1/crawl/get-crawl-job-status)
Poll for the current status and results of a crawl job
# List crawl jobs (/docs/scraper-api/v1/crawl/list-crawl-jobs)
List crawl jobs for the authenticated user with optional filtering and pagination
# Start a Crawl Job (/docs/scraper-api/v1/crawl/start-crawl-job)
Crawl a website starting from a seed URL up to a given depth and page limit
# Extract Content (/docs/scraper-api/v1/extraction/extract-content)
Extract clean markdown or HTML from any URL with sync or async mode support
# Get Extraction Job (/docs/scraper-api/v1/extraction/get-extraction-job)
Poll for job progress and retrieve extraction results (async mode only)
# List Extraction Jobs (/docs/scraper-api/v1/extraction/list-extraction-jobs)
List and filter extraction jobs by job ID, URL, status, output format, and date range
# Get map job details (/docs/scraper-api/v1/map/get-map-job)
Retrieve the full details and discovered links for a completed map job
# List map jobs (/docs/scraper-api/v1/map/list-map-jobs)
List map jobs for the authenticated user with optional filtering and pagination.
# Map URLs (/docs/scraper-api/v1/map/map-urls)
Returns the list of URLs found under the given base URL by combining sitemap parsing with HTML link extraction from the seed page. The optional `search` parameter filters the discovered URLs by case-insensitive substring match — it does NOT query a search engine.
# Get search job details (/docs/scraper-api/v1/search/get-search-job)
Retrieve the full details and results for a completed search job
# List search jobs (/docs/scraper-api/v1/search/list-search-jobs)
List search jobs for the authenticated user with optional filtering and pagination
{/* This file was generated by Fumadocs. Do not edit this file directly. Any changes should be made by running the generation command again. */}
# Search (/docs/scraper-api/v1/search/start-search-job)
Start a new search job
{/* This file was generated by Fumadocs. Do not edit this file directly. Any changes should be made by running the generation command again. */ }
{/* This file was generated by Fumadocs. Do not edit this file directly. Any changes should be made by running the generation command again. */}
# Get Statistics (/docs/scraper-api/v1/statistics/get-statistics)
This endpoint is documented but not yet available in production. The contract
below reflects the planned behavior. Reach out via
[support](https://geonode.com/contact) for early access or launch notification.
Retrieve aggregated extraction statistics for a given date range
# Health Check (/docs/scraper-api/v1/system/health-check)
{/* This file was generated by Fumadocs. Do not edit this file directly. Any changes should be made by running the generation command again. */}
# Get Concurrency Usage (/docs/scraper-api/v1/usage/get-concurrency-usage)
Return the caller's live work-concurrency slot usage against their plan limit.
# Create a Webhook (/docs/scraper-api/v1/webhooks/create-webhook)
This endpoint is documented but not yet available in production. The contract
below reflects the planned behavior. Reach out via
[support](https://geonode.com/contact) for early access or launch notification.
Create a webhook subscription for a given event type and return its generated signing secret
# Delete a Webhook (/docs/scraper-api/v1/webhooks/delete-webhook)
This endpoint is documented but not yet available in production. The contract
below reflects the planned behavior. Reach out via
[support](https://geonode.com/contact) for early access or launch notification.
Permanently remove a webhook subscription
# Get a Webhook (/docs/scraper-api/v1/webhooks/get-webhook)
This endpoint is documented but not yet available in production. The contract
below reflects the planned behavior. Reach out via
[support](https://geonode.com/contact) for early access or launch notification.
Retrieve a single webhook by its identifier
# List Webhook Deliveries (/docs/scraper-api/v1/webhooks/list-webhook-deliveries)
This endpoint is documented but not yet available in production. The contract
below reflects the planned behavior. Reach out via
[support](https://geonode.com/contact) for early access or launch notification.
List webhook delivery attempts with optional status filter and pagination
# List Webhooks (/docs/scraper-api/v1/webhooks/list-webhooks)
This endpoint is documented but not yet available in production. The contract
below reflects the planned behavior. Reach out via
[support](https://geonode.com/contact) for early access or launch notification.
List webhooks registered for the current user with pagination
# Rotate Webhook Secret (/docs/scraper-api/v1/webhooks/rotate-webhook-secret)
This endpoint is documented but not yet available in production. The contract
below reflects the planned behavior. Reach out via
[support](https://geonode.com/contact) for early access or launch notification.
Generate a new signing secret for the webhook and invalidate the previous one
# Update a Webhook (/docs/scraper-api/v1/webhooks/update-webhook)
This endpoint is documented but not yet available in production. The contract
below reflects the planned behavior. Reach out via
[support](https://geonode.com/contact) for early access or launch notification.
Partially update a webhook's url, description, event type, or active flag
# Perform Bandwidth-Limited Proxy Session (/docs/proxies/api-reference/sticky-session/get-sticky-session/get-bandwidth-limited)
Create a proxy session with bandwidth limitations to control the data transfer rate for your connection.
This API requires Basic Authentication. Include the following in your request:
* **Username**: `-session--limit-`
* **Password**: ``
* **Proxy**: `http://proxy.geonode.io:`
* **Accept Header**: `application/json`
## Request
```bash
curl --request GET \
-x "http://proxy.geonode.io:" \
--user "-session--limit-:" \
--url "http://ip-api.com/json"
```
## Response
### 200 Success
Bandwidth-limited proxy session executed successfully.
#### Response Fields
| Field | Type | Description |
| ------------- | ------ | ---------------------------------------------------- |
| `status` | string | The status of the request (e.g., "success") |
| `country` | string | The full name of the country where the IP is located |
| `countryCode` | string | The two-letter country code |
| `region` | string | The region code |
| `regionName` | string | The full name of the region |
| `city` | string | The name of the city |
| `zip` | string | The postal code associated with the IP |
| `lat` | number | The latitude coordinate |
| `lon` | number | The longitude coordinate |
| `timezone` | string | The timezone of the IP location |
| `isp` | string | The name of the Internet Service Provider |
| `org` | string | The organization that owns the IP address |
| `as` | string | The Autonomous System (AS) number |
| `query` | string | The queried IP address |
#### Example Response
```json
{
"status": "success",
"country": "Algeria",
"countryCode": "DZ",
"region": "22",
"regionName": "Sidi Bel Abbès",
"city": "Sidi Bel Abbes",
"zip": "22000",
"lat": 34.8934,
"lon": -0.6526,
"timezone": "Africa/Algiers",
"isp": "4 djaweb de AS fawri",
"org": "",
"as": "AS36947 Telecom Algeria",
"query": "197.203.245.147"
}
```
# Create a New Sticky Session (/docs/proxies/api-reference/sticky-session/get-sticky-session/get-create)
Create a new sticky session using proxy credentials. Sticky sessions let you keep the same IP address longer, which is useful for tasks that need a stable connection, like managing social media accounts or web scraping. Knowing how to use sticky sessions can improve your proxy experience by making it more efficient and reliable.
Geonode provides specific ports for sticky sessions, ensuring a persistent IP session:
| Protocol | Port Range |
| -------------- | ------------- |
| **HTTP/HTTPS** | 10000 - 10900 |
| **SOCKS5** | 12000 - 12010 |
Use sticky ports when you need a stable IP for session-based activities.
## Request
Make a request using a sticky session by including the session ID in your username string:
```bash
curl -X "http://proxy.geonode.io:" \
--user "-country--session-:" \
--url "http://ip-api.com/json" \
--header "Accept: application/json"
```
### Parameters
The session ID can be any custom string (1-25 alphanumeric characters or underscores). As long as you keep sending requests with the same session ID, you will maintain the same proxy IP.
## Response
### 200 Success
The response contains detailed geolocation information about the IP address assigned to your sticky session.
#### Response Fields
| Field | Type | Description |
| --------------- | ------- | --------------------------------------------------------- |
| `status` | string | The status of the request (e.g., "success") |
| `continent` | string | The continent where the IP is located |
| `continentCode` | string | The two-letter continent code |
| `country` | string | The full name of the country |
| `countryCode` | string | The two-letter country code (ISO 3166-1 alpha-2) |
| `region` | string | The region code |
| `regionName` | string | The full name of the region |
| `city` | string | The name of the city |
| `district` | string | The district name, if available |
| `zip` | string | The postal code associated with the IP |
| `lat` | number | The latitude coordinate |
| `lon` | number | The longitude coordinate |
| `timezone` | string | The timezone of the IP location |
| `offset` | integer | The time offset in seconds from UTC |
| `currency` | string | The currency code of the country |
| `isp` | string | The name of the Internet Service Provider |
| `org` | string | The organization that owns the IP address |
| `as` | string | The Autonomous System (AS) number |
| `asname` | string | The name associated with the AS number |
| `mobile` | boolean | Indicates whether the connection is from a mobile network |
| `proxy` | boolean | Indicates whether the IP is a known proxy |
| `hosting` | boolean | Indicates whether the IP is from a hosting provider |
| `query` | string | The queried IP address |
#### Example Response
```json
{
"status": "success",
"continent": "Africa",
"continentCode": "AF",
"country": "Algeria",
"countryCode": "DZ",
"region": "05",
"regionName": "Batna",
"city": "Batna City",
"district": "",
"zip": "05000",
"lat": 35.5064,
"lon": 6.0707,
"timezone": "Africa/Algiers",
"offset": 3600,
"currency": "DZD",
"isp": "4 djaweb de AS fawri",
"org": "",
"as": "AS36947 Telecom Algeria",
"asname": "ALGTEL-AS",
"mobile": false,
"proxy": false,
"hosting": false,
"query": "XXX.XXX.XXX.XXX"
}
```
# List All Active Sessions (/docs/proxies/api-reference/sticky-session/get-sticky-session/get-list-active)
Retrieve a paginated list of all currently active proxy sessions for your account. This endpoint helps you monitor and manage your active connections.
## Request
```bash
curl -X GET "https://app-api.geonode.com/api/sessions/proxies" \
-H "Authorization: Basic base64(username:password)"
```
### Query Parameters
| Parameter | Type | Required | Default | Description |
| ---------- | ------- | -------- | ------- | ------------------------------------- |
| `page` | integer | No | 1 | Page number for pagination |
| `pageSize` | integer | No | 250 | Number of sessions per page (max 250) |
## Response
### 200 Success
A list of active sessions retrieved successfully.
#### Response Fields
| Field | Type | Description |
| -------------------------------------- | ------- | ----------------------------------------------------------- |
| `sessions` | array | A list of currently active proxy sessions |
| `sessions[].id` | string | A unique identifier for the session |
| `sessions[].userSessionId` | string | The user-defined session identifier |
| `sessions[].userId` | string | The Geonode user ID associated with the session |
| `sessions[].port` | string | The port number assigned to the session |
| `sessions[].domain` | string | The target domain being accessed in the session |
| `sessions[].country` | string | The country code representing the location |
| `sessions[].rotatingIntervalInSeconds` | number | The interval (in seconds) at which the session rotates |
| `sessions[].durationInSeconds` | number | The total duration (in seconds) the session has been active |
| `count` | integer | The number of active sessions returned in the current page |
| `total` | integer | The total number of active sessions across all pages |
| `page` | integer | The current page number |
| `pageSize` | integer | The number of sessions per page |
#### Example Response
```json
{
"sessions": [
{
"id": "a7f7c7a3-aaaa-4908-b7bd-71091daadaa6",
"userSessionId": "vzzzzd",
"userId": "geonode_userid",
"port": "10000",
"domain": "ip-api.com",
"country": "dz",
"rotatingIntervalInSeconds": 179.32,
"durationInSeconds": 2.785
}
],
"count": 1,
"total": 1,
"page": 1,
"pageSize": 250
}
```
# Configuring Proxy Session ID & Lifetime (/docs/proxies/api-reference/sticky-session/get-sticky-session/get-proxy-session)
Control the duration of your proxy sessions by setting a custom session ID and lifetime. This allows you to maintain stable connections for extended periods, which is essential for tasks requiring persistent IP addresses.
The `lifetime` parameter is measured in **minutes** and defines how long the session remains active:
| Constraint | Value |
| ----------- | ----------------------- |
| **Minimum** | 3 minutes |
| **Maximum** | 1440 minutes (24 hours) |
| **Default** | 10 minutes |
**Examples:**
* `lifetime: 5` → 5 minutes
* `lifetime: 30` → 30 minutes
* `lifetime: 60` → 1 hour
* `lifetime: 180` → 3 hours
Your custom session ID must follow these rules:
* **Format**: 1-25 alphanumeric characters or underscores (`_`)
* **Regex pattern**: `^[A-Za-z0-9_]{1,25}$`
* **Valid characters**: Letters (A-Z, a-z), digits (0-9), and underscores
* **No spaces or special symbols** allowed
**Examples of valid session IDs:**
* `session123`
* `my_session_01`
* `ABC123`
## Request
```bash
curl --request GET \
-x "http://proxy.geonode.io:" \
--user "-session--lifetime-:" \
--url "http://ip-api.com/json"
```
## Response
### 200 Success
Proxy session created successfully with the specified lifetime.
#### Response Fields
| Field | Type | Description |
| --------------- | ------- | --------------------------------------------------------- |
| `status` | string | The status of the request (e.g., "success") |
| `continent` | string | The continent where the IP is located |
| `continentCode` | string | The two-letter continent code |
| `country` | string | The full name of the country |
| `countryCode` | string | The two-letter country code |
| `region` | string | The region code |
| `regionName` | string | The full name of the region |
| `city` | string | The name of the city |
| `district` | string | The district name, if available |
| `zip` | string | The postal code associated with the IP |
| `lat` | number | The latitude coordinate |
| `lon` | number | The longitude coordinate |
| `timezone` | string | The timezone of the IP location |
| `offset` | integer | The time offset in seconds from UTC |
| `currency` | string | The currency code of the country |
| `isp` | string | The name of the Internet Service Provider |
| `org` | string | The organization that owns the IP address |
| `as` | string | The Autonomous System (AS) number |
| `asname` | string | The name associated with the AS number |
| `mobile` | boolean | Indicates whether the connection is from a mobile network |
| `proxy` | boolean | Indicates whether the IP is a known proxy |
| `hosting` | boolean | Indicates whether the IP is from a hosting provider |
| `query` | string | The queried IP address |
#### Example Response
```json
{
"status": "success",
"continent": "Africa",
"continentCode": "AF",
"country": "Algeria",
"countryCode": "DZ",
"region": "05",
"regionName": "Batna",
"city": "Batna City",
"district": "",
"zip": "05000",
"lat": 35.5064,
"lon": 6.0707,
"timezone": "Africa/Algiers",
"offset": 3600,
"currency": "DZD",
"isp": "4 djaweb de AS fawri",
"org": "",
"as": "AS36947 Telecom Algeria",
"asname": "ALGTEL-AS",
"mobile": false,
"proxy": false,
"hosting": false,
"query": "192.168.1.1"
}
```
# Release Sticky Session by Session ID and Port (/docs/proxies/api-reference/sticky-session/release-sticky-session/put-release-by-session-id-and-port)
Release one or more sticky sessions by specifying their session IDs and ports. This allows you to free up resources and terminate specific proxy connections.
## Request
```bash
curl -X PUT "https://app-api.geonode.com/api/sessions/release/proxies" \
-H "Authorization: Basic base64(username:password)" \
-H "Content-Type: application/json" \
-d '{"data":[{"sessionId":"random0001","port":10000}]}'
```
### Request Body Parameters
| Field | Type | Required | Description |
| ------------------ | ------- | -------- | ------------------------------------------- |
| `data` | array | Yes | Array containing session details |
| `data[].sessionId` | string | Yes | The session ID to release |
| `data[].port` | integer | Yes | The port number associated with the session |
### Example Request Body
```json
{
"data": [
{
"sessionId": "random0001",
"port": 10000
}
]
}
```
## Response
### 200 Success
Session released successfully.
#### Response Fields
| Field | Type | Description |
| --------- | ------- | ---------------------------------------------------- |
| `success` | boolean | Indicates whether the session release was successful |
#### Example Response
```json
{
"success": true
}
```
# Release Proxy Session by Multiple Port (/docs/proxies/api-reference/sticky-session/release-sticky-session/put-release-proxy-session-by-multiple-port)
Release sticky sessions running on multiple ports in a single request. This is useful when you need to free several proxy connections at once.
## Request
```bash
curl -X PUT "https://monitor.geonode.com/sessions/release/proxies" \
-H "Authorization: Basic base64(username:password)" \
-H "Content-Type: application/json" \
-d '{"data":[{"port":10000},{"port":10001},{"port":10002}]}'
```
### Request Body Parameters
| Field | Type | Required | Description |
| ------------- | ------- | -------- | -------------------------------------------- |
| `data` | array | Yes | Array containing session details |
| `data[].port` | integer | Yes | Port number of each proxy session to release |
### Example Request Body
```json
{
"data": [{ "port": 10000 }, { "port": 10001 }, { "port": 10002 }]
}
```
## Response
### 200 Success
Successfully released the specified ports.
#### Response Fields
| Field | Type | Description |
| --------- | ------- | ---------------------------------------------------- |
| `success` | boolean | Indicates whether the session release was successful |
#### Example Response
```json
{
"success": true
}
```
# Release Proxy Session by Port (/docs/proxies/api-reference/sticky-session/release-sticky-session/put-release-proxy-session-by-port)
Release a sticky session by specifying its port number. Use this when you need to free up a particular proxy connection without affecting other sessions.
## Request
```bash
curl -X PUT "https://monitor.geonode.com/sessions/release/proxies" \
-H "Authorization: Basic base64(username:password)" \
-H "Content-Type: application/json" \
-d '{"data":[{"port":10001}]}'
```
### Request Body Parameters
| Field | Type | Required | Description |
| ------------- | ------- | -------- | ----------------------------------------------- |
| `data` | array | Yes | Array containing session details |
| `data[].port` | integer | Yes | The port number of the proxy session to release |
### Example Request Body
```json
{
"data": [
{
"port": 10001
}
]
}
```
## Response
### 200 Success
Session released successfully.
#### Response Fields
| Field | Type | Description |
| --------- | ------- | ---------------------------------------------------- |
| `success` | boolean | Indicates whether the session release was successful |
#### Example Response
```json
{
"success": true
}
```
# Overview (/docs/proxies/api-reference/sticky-session/release-sticky-session/release-sticky-session)
Release sticky sessions for specified Geonode services by session ID and port. This allows you to free up resources and terminate specific proxy connections.
## What is Session Release?
Session release is the process of terminating active sticky sessions before they naturally expire. When you release a session, you're telling the Geonode proxy system to immediately free up the IP address and resources associated with that session.
## Why Release Sessions?
Release sessions when you need to free up resources or terminate specific proxy connections before they naturally expire.
## How It Works
To release a sticky session, you need to provide:
* **Session ID**: The unique identifier for the session you want to release
* **Port**: The port number associated with the session
The system will immediately terminate the session and free up the associated resources. Once released, the IP address becomes available for other users, and you'll need to create a new session if you want to continue using a sticky connection.
## Available Operations
This section provides endpoints for managing sticky session releases:
* **[Release Sticky Session by Session ID and Port](/docs/proxies/api-reference/sticky-session/release-sticky-session/put-release-by-session-id-and-port)**: Release one or more specific sessions by providing their session IDs and ports
Before releasing sessions, you can view all your active sessions using the
[List All Active
Sessions](/docs/proxies/api-reference/sticky-session/get-sticky-session/get-list-active) endpoint to
identify which sessions you want to release.
## Best Practices
Before releasing sessions, you can view all your active sessions to identify which ones you want to release. Once a session is released, it cannot be restored, so make sure you're ready to release it before executing the operation.
Once a session is released, it cannot be restored. You'll need to create a new
session if you want to continue using a sticky connection. Make sure you're
ready to release a session before executing the release operation.
# Port Configuration (/docs/proxies/getting-started/setup_and_configuration/advance-configuration/port-configuration)
import SupportParagraph from "../../../../../snippets/support-paragraph.mdx";
The **Port Configuration** section in your Geonode dashboard lets you set up and manage proxy ports with precise geo-targeting and connection options.
***
## Step 1 — Open the Port Configuration page
1. Log in to your Geonode dashboard.
2. Select **Port Configuration** from the top navigation bar.
You’ll see configuration options on the right and a list of added locations in the center.
***
## Step 2 — Configure your port settings
### 1. Country targeting (required)
* Open the **Countries** dropdown.
* Select the country you want to target — this field is required for geo-targeting.
### 2. State targeting (optional)
* Once a country is selected, the **States** dropdown becomes active.
* Choose a specific state if needed.
### 3. City targeting (optional)
* After choosing a state, the **Cities** dropdown will be enabled.
* Select a city for more precise targeting.
### 4. Protocol selection
* Click the **HTTP/HTTPS** dropdown.
* Choose a protocol:
* **HTTP/HTTPS** — standard web traffic.
* **SOCKS5** — advanced, more flexible option (if available).
→ [Understanding Protocol Type for Proxy Configuration](/docs/proxies/getting-started/knowledge-base/protocol-type)
### 5. Session type
* Select a session mode:
* **Rotating** — IP changes periodically (best for scraping/automation).
* **Sticky** — one IP per session (best for login or persistent tasks).
→ [Understanding Session Type for Proxy Configuration](/docs/proxies/getting-started/knowledge-base/session-type)
### 6. Port selection (required)
* Click **Select Port** to choose from available ports.
* This field is required to complete setup.
→ [Understanding Port Usage for Proxy Configuration](/docs/proxies/getting-started/knowledge-base/proxy-usage)
***
## Step 3 — Add the configuration
When all required fields are filled, the **Add** button becomes active.\
Click **Add** to save your configuration.
***
## Step 4 — Review added locations
After saving, all configured ports and locations appear in the list.\
You can view, edit, or delete configurations at any time.
***
## Port in use
When a port is assigned to a specific country, it can’t be reused elsewhere.
***
## Final result
You’re now ready to manage ports efficiently in Geonode.
* **Rotating Example**
* **Sticky Example**
***
## Troubleshooting
* **Add button disabled:** make sure Country and Port are selected.
* **State/City inactive:** these fields unlock only after a country is chosen.
* **Connection issues:** verify your protocol and session settings.
***
***
## FAQs
The button stays inactive until all required fields are selected. Choose both a country and a port before proceeding.
{" "}
No. Each port can only be assigned to one country.
{" "}
Rotating sessions periodically change your IP, ideal for scraping and
automation.
{" "}
Yes, you can select either protocol during configuration. SOCKS5 support
depends on your plan.
State and city fields are only active after selecting a country.
# Proxy Authorization in API Using Headers (/docs/proxies/getting-started/setup_and_configuration/advance-configuration/proxy-authentication-in-api)
import SupportParagraph from "../../../../../snippets/support-paragraph.mdx";
***
API authorization ensures secure access to Geonode’s proxy services.\
This guide explains how to authenticate requests using the `Authorization` header.
***
## Step 1 — Get your API credentials
Before setting up authentication, make sure you have your Geonode API credentials.
→ [How to access your Geonode API credentials](/docs/proxies/getting-started/prerequisites/access-credentials)
***
## Step 2 — Generate the Authorization header
Geonode’s API uses **Basic Authentication**, which requires your username and password encoded in **Base64**.
1. Open the API documentation for the endpoint you want to test (for example, *Retrieve Usage Statistics*).
2. Click **Try it** on the API page.
3. A popup will appear asking for your **username** and **password**. Enter them.
4. The system automatically generates the `Authorization` header for you — containing the Base64-encoded string.
5. Copy the generated header or the encoded string.\
You can use it directly in your API requests or store the Base64 value in your code securely.
***
## Step 3 — Follow best practices
* **Generate once, reuse:** Create your token once and reuse it for multiple requests.
* **Store securely:** Keep it in a `.env` file or secret manager.
* **Avoid hardcoding:** Never paste credentials directly into your source code.
* **Always use HTTPS:** This encrypts your traffic and protects sensitive data.
* **Rotate regularly:** Update your credentials periodically.\
When you do, generate a new token.
***
## Troubleshooting
* Check that your header is formatted correctly:
* Encode exactly `username:password` — no extra spaces or characters.
* Use HTTPS in all requests; avoid using insecure HTTP.
* Test your setup with tools like **Postman** or **cURL** to confirm it works.
* If your credentials are compromised, **change your password**, regenerate your API key, and update the token.
***
***
## FAQs
No. The Geonode API documentation automatically generates the Base64-encoded string for you.
Not at the moment — all proxy API requests require Basic Authorization.
Currently, Geonode only supports Basic Authorization. Check future API updates
for new methods.
No. Base64 is not encryption — it’s only encoding. Always use HTTPS to keep
credentials safe.
Rotate them regularly or immediately if you suspect compromise.
# Proxy Configuration (/docs/proxies/getting-started/setup_and_configuration/advance-configuration/proxy-configuration)
***
## What is the Endpoint Generator
The Endpoint Generator helps you create customized proxy lists based on your preferences.\
You can select **IP type**, **host**, **location**, **protocol**, and **session type**, then export endpoints in `.txt` format.
***
## Step 1 — Access the Dashboard
1. Log in to your Geonode dashboard.
2. Scroll down to **Proxy Configuration**.
***
## Step 2 — Choose the Endpoint Format
There are **six formats** available for generating endpoints, depending on your authentication and connection preferences.
→ [Understanding Different Proxy Endpoint Formats and Their Uses](/docs/proxies/getting-started/knowledge-base/endpoint-formats)
***
## Step 3 — Set the Endpoint Count
Specify how many endpoints you want to generate.
***
## Step 4 — Configure Proxy Parameters
There are **seven options** available for customizing proxy endpoints.
***
### 1. IP Type
Select the type of IPs you need.\
Three types are available:
* **Residential**
* **Datacenter**
* **Mixed**
→ [Learn about Residential, Datacenter, and Mixed IPs and their best use cases](/docs/proxies/getting-started/knowledge-base/ip-type)
***
### 2. Gateway (Host)
Choose the appropriate host for your location.\
Three options are available:
| Location | Host Address |
| ------------- | ------------------------------------ |
| France | `proxy.geonode.io` |
| United States | `us.premium-residential.geonode.com` |
| Singapore | `sg.premium-residential.geonode.com` |
Selecting a host automatically updates the corresponding address.
→ [Learn about Geonode’s proxy gateways and how they work](/docs/proxies/getting-started/knowledge-base/gateway)
***
### 3. Geo-Targeting (Country, State, City)
Select your desired **country**, **state**, or **city**.\
By default, the system uses **Any**, meaning proxies can come from any location.
**Example:**
* **Host:** Singapore
* **Target Country:** China
This setup provides faster connectivity to nearby Chinese servers.
→ [What is Geo-Targeting](/docs/proxies/getting-started/knowledge-base/geo-targeting)
***
### 4. Protocol
Choose how your proxy handles connections:
1. **HTTP/HTTPS** — standard web traffic.
2. **SOCKS5** — flexible and secure (if supported).
→ [Understanding Protocol Type for Proxy Configuration](/docs/proxies/getting-started/knowledge-base/protocol-type)
When you change the protocol, port numbers in generated endpoints update automatically.
→ [Learn how ports work in Geonode’s proxy configuration](/docs/proxies/getting-started/knowledge-base/proxy-usage)
***
### 5. Session Type
Define how IPs are handled within a session:
1. **Rotating Session** — IP changes periodically (best for scraping and automation).
2. **Sticky Session** — IP stays the same for the entire session (best for logins or persistent tasks).
→ [A Guide to Rotating and Sticky Sessions](/docs/proxies/getting-started/knowledge-base/session-type)
***
### 6. Rotating Interval (Sticky Sessions Only)
If you use Sticky Sessions, you can set a rotation interval to control how often IPs refresh — in minutes or hours.
This keeps your connection stable while maintaining periodic IP rotation for security and reliability.
***
## Step 5 — Select the Output Format
After completing the configuration, you can:
* **Copy** endpoints directly for immediate use, or
* **Download** them as a `.txt` file for later.
***
✅ **You’re all set!**\
You’ve successfully created a ready-to-use list of proxy endpoints.\
Your setup is now fully configured and ready to integrate into your applications.
# Whitelist IP (/docs/proxies/getting-started/setup_and_configuration/advance-configuration/whitelist-ip)
import VerifyProxyConnectionComponent from "../../../../../snippets/verify-proxy-connection-component.mdx";
import SupportParagraph from "../../../../../snippets/support-paragraph.mdx";
Whitelisting your IP lets you access proxy services without re-entering credentials each time.
## Key limits and requirements
* Authentication: Basic Auth (username:password encoded in Base64).
* IP limit: up to 150 whitelisted IPs per user.
* Batch size: up to 10 IPs per request (add/remove).
* Validation: the API rejects invalid or duplicate IPs.
* Rate limit: up to 100 requests per minute.
## Add an IP to the whitelist
You can use either the dashboard (recommended for non-developers) or the API.
### Using the dashboard
#### Add IP addresses
1. Open **Whitelisted IPs** in settings.
2. Enter the IP to whitelist. For a quick setup, click **Detect My IP**.\
Optionally add a description (e.g., Home, Office, Server). Click **Add**.
3. Verify that the IP appears in the table with correct details.
#### Remove an IP
1. Find the IP in the whitelist table.
2. Click **Delete** next to the IP.
3. Confirm the deletion and check the notification.
### Using the API
Available endpoints:
* [Retrieve whitelisted IPs](/docs/proxies/api-reference/whitelisting-ip/get) `GET`
* [Add whitelisted IPs](/docs/proxies/api-reference/whitelisting-ip/post) `POST`
* [Update IP description](/docs/proxies/api-reference/whitelisting-ip/put) `PUT`
* [Remove whitelisted IPs](/docs/proxies/api-reference/whitelisting-ip/delete) `DELETE`
For request/response formats and examples, see the Geonode API docs:\
[Geonode API Documentation](/docs/proxies/api-reference/whitelisting-ip/whitelisting-ips)
***
***
## FAQs
Basic Authentication ensures only authorized users can modify the whitelist.
Encode your username and password in Base64 and include it in the request.
Yes. Use the Geonode API to add, update, and remove IPs programmatically.
Requests that push the total over 150 IPs are rejected. Remove unused entries
first.
Delete the incorrect IP from the dashboard or via the API, then add the
correct one.
Whitelisting lets trusted devices connect without re-entering credentials,
simplifying access for known locations.
***
# Selenium (/docs/proxies/getting-started/setup_and_configuration/automation-frameworks/selenium)
import ProxyInfoFromGeonode from "../../../../../snippets/get-proxy-info-from-geonode.mdx";
import VerifyProxyConnectionComponent from "../../../../../snippets/verify-proxy-connection-component.mdx";
import SupportParagraph from "../../../../../snippets/support-paragraph.mdx";
***
This guide will help you integrate the Geonode API with Selenium to manage proxies effectively while running headless browsers.
**We will be using Python for this guide.**
***
***
## Prerequisites
* Python installed on your system
* Geonode API credentials (username and password)
* ChromeDriver installed (compatible with your Chrome version)
## Steps: Setting Up a Proxy in Brave
Follow these steps to configure a proxy in Selenium:
### Step 1: Set Up a Virtual Environment
Creating a virtual environment helps isolate dependencies and avoid conflicts.
```
# Install virtualenv if not already installed
pip install virtualenv
# Create a virtual environment
python -m venv .venv
# Activate the virtual environment
## On Windows
.venv\Scripts\activate
## On macOS/Linux
source .venv/bin/activate
```
### Step 2: Install Required Libraries
```
pip install selenium python-dotenv selenium-wire
```
* **Selenium:** For browser automation
* **python-dotenv:** For managing environment variables
* **selenium-wire:** To handle proxy authentication, as Selenium doesn't provide it natively
### Step 3: Configure Geonode Proxy Endpoint
With the help of the **Endpoint Generator**, you can easily generate the proxy with specific configurations such as:
* Target country
* Port
* Session persistence
* And many more
Refer to the guide **[How to Use the Endpoint Generator](/docs/proxies/getting-started/knowledge-base/geo-targeting)** to generate your endpoints.
### Step 4: Setting up environment variables
1. Create a `.env` file in your project directory.
2. Add your credentials:
```
GEONODE_USERNAME=your_geonode_username
GEONODE_PASSWORD=your_geonode_password
GEONODE_HOST=proxy.geonode.io
GEONODE_PORT=9000
GEONODE_DNS=your_geonode_dns
```
Never upload your
`.env`
file to the internet.
## Code Implementation
### I. Import
```
import os
from dotenv import load_dotenv
from seleniumwire import webdriver
from selenium.webdriver.chrome.options import Options
```
### ii. Load environment variables
```
load_dotenv()
proxy_host = os.getenv('GEONODE_HOST')
proxy_port = os.getenv('GEONODE_PORT')
username = os.getenv('GEONODE_USERNAME')
password = os.getenv('GEONODE_PASSWORD')
GEONODE_DNS = os.getenv('GEONODE_PROXY')
```
### iii. Configure Proxy options
```
proxy_options = {
'proxy': {
'http': f'http://{username}:{password}@{proxy_host}:{proxy_port}',
'https': f'https://{username}:{password}@{proxy_host}:{proxy_port}',
}
}
```
### iv. Configure Browser Options
```
chrome_options = Options()
chrome_options.add_argument('--disable-gpu')
chrome_options.add_argument('--start-maximized')
chrome_options.add_argument('--ignore-certificate-errors')
```
### v. Initialize Browser
```
browser = webdriver.Chrome(seleniumwire_options=proxy_options, options=chrome_options)
```
### vi. Open IP address website to check
```
urlToGet = "https://ip-api.com/"
browser.get(urlToGet)
```
**Output:**
### vii. Keep the browser open
```
input("Press Enter to close the browser...")
```
### viii. Quit the browser
```
browser.quit()
```
### xi. Full Code
```
import os
from dotenv import load_dotenv
from seleniumwire import webdriver
from selenium.webdriver.chrome.options import Options
load_dotenv()
proxy_host = os.getenv('GEONODE_HOST')
proxy_port = os.getenv('GEONODE_PORT')
username = os.getenv('GEONODE_USERNAME')
password = os.getenv('GEONODE_PASSWORD')
GEONODE_DNS = os.getenv('GEONODE_PROXY')
proxy_options = {
'proxy': {
'http': f'http://{username}:{password}@{proxy_host}:{proxy_port}',
'https': f'https://{username}:{password}@{proxy_host}:{proxy_port}',
}
}
chrome_options = Options()
chrome_options.add_argument('--disable-gpu')
chrome_options.add_argument('--start-maximized')
chrome_options.add_argument('--ignore-certificate-errors')
browser = webdriver.Chrome(seleniumwire_options=proxy_options, options=chrome_options)
urlToGet = "https://ip-api.com/"
browser.get(urlToGet)
input("Press Enter to close the browser...")
browser.quit()
```
### x. Folder Stucture
```
Project Root
├── .venv/
├── .env
└── app.py
```
## Source Code
You can find the full source code for this script on GitHub at the following link:
[https://github.com/geonodecom/proxy-testing-toolkit/tree/automation-framework/selenium](https://github.com/geonodecom/proxy-testing-toolkit/tree/automation-framework/selenium)
***
## Use Cases of Integrating Geonode with Selenium
Integrating Geonode with Selenium can be beneficial for a wide range of applications, including:
1. **Web Scraping:** Collect data from websites while maintaining anonymity to avoid IP bans.
2. **Ad Verification:** Test and verify advertisements across different geographies to ensure proper delivery.
3. **Price Monitoring:** Track pricing changes on e-commerce platforms without getting blocked.
4. **SEO Monitoring:** Monitor search engine results and competitor websites without affecting personalized search results.
5. **Market Research:** Gather data from various sources to analyze trends and competitor performance.
6. **Social Media Automation:** Manage multiple social media accounts while avoiding detection.
7. **Fraud Detection:** Simulate real-world traffic for security testing and fraud detection systems.
Explore more use cases here *[https://geonode.com/use-cases](https://geonode.com/use-cases)*
***
## Troubleshooting Tips
* **Timeout Errors:** Check if the proxy is active or switch to another proxy.
* **Authentication Issues:** Double-check your Geonode API credentials.
* **Incompatible ChromeDriver:** Ensure ChromeDriver matches your Chrome version.
***
## FAQs
The authentication details are embedded in the proxy URL, so no pop-up should appear.
{" "}
Yes, but you need to adjust the proxy settings using Firefox profiles.
Verify the proxy server details and ensure Geonode proxies are correctly configured.
# macOS (/docs/proxies/getting-started/setup_and_configuration/desktop-OS/macOS)
import BrowsersFaqs from "../../../../../snippets/browsers-faqs.mdx";
import ProxyInfoFromGeonode from "../../../../../snippets/get-proxy-info-from-geonode.mdx";
import VerifyProxyConnectionComponent from "../../../../../snippets/verify-proxy-connection-component.mdx";
import SupportParagraph from "../../../../../snippets/support-paragraph.mdx";
import PopupAuth from "../../../../../snippets/popup-auth.mdx";
***
***
## Step-by-step setup for macOS
Follow these steps to configure your proxy manually.
### Step 1 — Open System Preferences
1. Click the **Apple menu** icon in the top-left corner.
2. Select **System Preferences** from the dropdown.
***
### Step 2 — Open Network Settings
1. In **System Preferences**, click **Network**.
2. Choose your active Wi-Fi network and click **Details** (or the “i” icon).
***
### Step 3 — Select Your Active Network
Make sure the selected network is the one you’re currently connected to.
***
### Step 4 — Open the Proxies Tab
1. Click **Advanced** in the Network window.
2. Open the **Proxies** tab.
***
### Step 5 — Configure Proxy Settings
1. In the **Proxies** tab, you’ll see several proxy types:
* Web Proxy (HTTP)
* Secure Web Proxy (HTTPS)
* SOCKS Proxy
2. Check the box next to the proxy type you’re setting up (for example, **Web Proxy (HTTP)**).
3. Enter your **Proxy Server** address and **Port** — both can be copied from your Geonode Dashboard (click the copy icon next to *Host*).
***
### Step 6 — Apply and Save Changes
1. Click **OK** to close Advanced settings.
2. Click **Apply** in the main Network window to save your configuration.
***
***
***
# Windows (/docs/proxies/getting-started/setup_and_configuration/desktop-OS/windows)
import BrowsersFaqs from "../../../../../snippets/browsers-faqs.mdx";
import ProxyInfoFromGeonode from "../../../../../snippets/get-proxy-info-from-geonode.mdx";
import VerifyProxyConnectionComponent from "../../../../../snippets/verify-proxy-connection-component.mdx";
import SupportParagraph from "../../../../../snippets/support-paragraph.mdx";
import PopupAuth from "../../../../../snippets/popup-auth.mdx";
***
***
## Step-by-step setup for Windows 10/11
Follow these steps to configure your proxy manually:
### Step 1 — Open Proxy Settings
1. Click the **Windows Search Bar** and type:\
`"Change proxy settings"`
2. In the settings window, scroll to **Manual Proxy Setup**.
***
### Step 2 — Enter Proxy Details
1. A configuration window will open:
2. Enter the **Proxy IP** and **Port** you copied from your Geonode Dashboard.
3. Click **Save** to apply the settings.
***
***
***
# AdsPower (/docs/proxies/getting-started/setup_and_configuration/browsers/adspower)
import { Steps } from "fumadocs-ui/components/steps";
import ProxyInfoFromGeonode from "../../../../../snippets/get-proxy-info-from-geonode.mdx";
import SupportParagraph from "../../../../../snippets/support-paragraph.mdx";
import BrowsersFaqs from "../../../../../snippets/browsers-faqs.mdx";
This guide explains how to configure a Geonode proxy in the **AdsPower** browser.
## Setting Up a Proxy in AdsPower
### Step 1: Install AdsPower
1. Go to the [AdsPower Download Page](https://www.adspower.com/download).
2. Download the version for your operating system.
3. Install AdsPower and create an account.
***
### Step 2: Create a New Profile
1. Open AdsPower and click **New Profile**.\
2. Fill in the required details:
* Profile name
* Operating system
* Browser version
3. (Optional) Randomize the fingerprint for extra security.
4. Review your browser details on the right side.\
***
### Step 3: Add a Proxy
1. Go to the **Proxy** tab.\
2. Select the connection type — for this guide, choose **HTTPS**.\
3. Enter your proxy details:
* Proxy (IP:Port)
* Username
* Password
4. Set **Geonode** as the main proxy provider.
***
### Step 4: Create and Launch the Profile
1. Click **Create Profile** to save your settings.
2. Once created, the profile will appear in the list.
* If a location appears, the proxy is active.\
3. Click **Start** to launch the browser with your configured proxy.\
Your **Geonode** proxy is now successfully configured in AdsPower.
# Brave (/docs/proxies/getting-started/setup_and_configuration/browsers/brave)
import ProxyOs from "../../../../../snippets/proxy-os.mdx";
import BrowsersFaqs from "../../../../../snippets/browsers-faqs.mdx";
import ProxyInfoFromGeonode from "../../../../../snippets/get-proxy-info-from-geonode.mdx";
import VerifyProxyConnectionComponent from "../../../../../snippets/verify-proxy-connection-component.mdx";
import SupportParagraph from "../../../../../snippets/support-paragraph.mdx";
This guide will help you configure a Geonode proxy in the Brave browser.
## Setting Up a Proxy in Brave
### Open Brave Settings
1. Click the three-dot menu in the top-right corner of Brave.
2. Select **Settings** from the dropdown menu.
***
### Access System Proxy Settings
1. In the sidebar, click **System**.
2. Then click **Open your computer’s proxy settings**.
***
### Configure the Proxy on Your Operating System
Brave will now open your system proxy configuration window.\
Follow the appropriate setup guide for your OS below:
Once configured, Brave will automatically route its traffic through the assigned proxy.
# Chrome (/docs/proxies/getting-started/setup_and_configuration/browsers/chrome)
import ProxyOs from "../../../../../snippets/proxy-os.mdx";
import BrowsersFaqs from "../../../../../snippets/browsers-faqs.mdx";
import ProxyInfoFromGeonode from "../../../../../snippets/get-proxy-info-from-geonode.mdx";
import VerifyProxyConnectionComponent from "../../../../../snippets/verify-proxy-connection-component.mdx";
import SupportParagraph from "../../../../../snippets/support-paragraph.mdx";
This guide will help you configure a Geonode proxy in Chrome using two methods:
* **Using the Geonode Proxy Manager Extension (recommended)**
* **Manual setup through Chrome system settings**
## Method 1: Using the Geonode Proxy Manager Extension
The easiest and most flexible way to configure a proxy in Chrome is by using the **Geonode Proxy Manager** extension.\
It lets you switch proxies quickly without changing system-wide settings.
### Install and Configure the Extension
1. Install **Geonode Proxy Manager** from the [Chrome Web Store](https://chromewebstore.google.com/detail/geonode-proxy-manager/ippioaknloonmaibmmepkemhmhinohge).
2. Follow this guide to complete the setup:\
[How to Use the Geonode Chrome Extension for Proxy Management](/docs/proxies/getting-started/setup_and_configuration/extensions/geonode-proxy-manager).
Once installed, Chrome will route all traffic through the selected proxy.
## Method 2: Manual Proxy Setup in Chrome
If you prefer to configure the proxy manually, follow these steps:
### Open Chrome Settings
1. Click the three-dot menu in the top-right corner of Chrome.
2. Select **Settings** from the dropdown menu.
***
### Access System Proxy Settings
1. In the left sidebar, click **System**.
2. Then click **Open your computer’s proxy settings**.
***
### Configure the Proxy on Your Operating System
Chrome will now open your system proxy configuration window.\
Follow the appropriate guide for your OS below:
Once configured, Chrome will automatically route its traffic through the assigned proxy.
# ClonBrowser (/docs/proxies/getting-started/setup_and_configuration/browsers/clonbrowser)
import BrowsersFaqs from "../../../../../snippets/browsers-faqs.mdx";
import ProxyInfoFromGeonode from "../../../../../snippets/get-proxy-info-from-geonode.mdx";
import SupportParagraph from "../../../../../snippets/support-paragraph.mdx";
This guide will help you configure a Geonode proxy in ClonBrowser.
## Setting Up a Proxy in ClonBrowser
### Install ClonBrowser
1. Go to the [ClonBrowser Download Page](https://www.clonbrowser.com/download).
2. Download the version for your operating system.
3. Install the browser and create an account.
***
### Create a New Profile
1. Open ClonBrowser and click **Add Profile**.\
2. Fill in the required details:
* Profile name
* Operating system
3. (Optional) Randomize the fingerprint for extra security.
4. Review your browser details on the right panel.\
***
### Add Proxy Details
1. Open the **Proxy** tab.
2. Choose a connection type — for this guide, select **HTTP**.
3. Enter your proxy information:
* Proxy (IP:Port)
* Username
* Password
4. Click **Create Profile** to save your configuration.
***
### Test the Proxy Connection
1. Click **Connect Test** to verify the connection.
2. If successful, the proxy will appear as active.
***
### Launch the Profile
1. Once the profile is created, it will appear in your list.
* If the proxy shows a location, it means it’s active.\
2. Click **Start** to open the browser with your configured proxy.\
Your Geonode proxy is now successfully configured in ClonBrowser.
# Dolphin Anty (/docs/proxies/getting-started/setup_and_configuration/browsers/dolphin-anty)
import BrowsersFaqs from "../../../../../snippets/browsers-faqs.mdx";
import ProxyInfoFromGeonode from "../../../../../snippets/get-proxy-info-from-geonode.mdx";
import SupportParagraph from "../../../../../snippets/support-paragraph.mdx";
This guide will help you configure a Geonode proxy in the Dolphin Anty browser.
## Setting Up a Proxy in Dolphin Anty
### Install Dolphin Anty
1. Go to the [Dolphin Anty Download Page](https://dolphin-anty.com/download/).
2. Download the version for your operating system.
3. Install the browser and create an account.
***
### Create a New Profile
1. Open Dolphin Anty and click **Add Profile**.\
2. Fill in the required details:
* Profile name
* Operating system
3. (Optional) Randomize the fingerprint for extra security.
4. Review your browser details on the right.\
***
### Add Proxy Details
1. Open the **Proxy** tab.
2. Choose a connection type — for this guide, select **HTTP**.
3. Enter your proxy information:
* Proxy (IP:Port)
* Username
* Password
4. Click **Create Profile** to save the configuration.
5. A green check mark indicates the proxy is connected.
***
### Launch and Verify the Profile
1. Once created, your profile will appear in the list.
* If the proxy shows a location, it’s active.\
2. Click **Start** to launch the browser with the configured proxy.
3. Visit [`http://ip-api.com/json`](http://ip-api.com/json) to confirm your IP.
Your Geonode proxy is now successfully configured in Dolphin Anty.
# Edge (/docs/proxies/getting-started/setup_and_configuration/browsers/edge)
import BrowsersFaqs from "../../../../../snippets/browsers-faqs.mdx";
import ProxyInfoFromGeonode from "../../../../../snippets/get-proxy-info-from-geonode.mdx";
import VerifyProxyConnectionComponent from "../../../../../snippets/verify-proxy-connection-component.mdx";
import ProxyOs from "../../../../../snippets/proxy-os.mdx";
import SupportParagraph from "../../../../../snippets/support-paragraph.mdx";
This guide will help you configure a Geonode proxy in the Microsoft Edge browser.
## Setting Up a Proxy in Edge
### Open Edge Settings
1. Click the three-dot menu in the top-right corner of Edge.
2. Select **Settings** from the dropdown menu.
***
### Access System Proxy Settings
1. In the left sidebar, click **System**.
2. Then click **Open your computer’s proxy settings**.
***
### Configure the Proxy on Your Operating System
Edge will now open your system proxy configuration window.\
Follow the appropriate setup guide for your OS below:
Once configured, Edge will automatically route its traffic through the assigned proxy.
# Firefox (/docs/proxies/getting-started/setup_and_configuration/browsers/firefox)
import BrowsersFaqs from "../../../../../snippets/browsers-faqs.mdx";
import ProxyInfoFromGeonode from "../../../../../snippets/get-proxy-info-from-geonode.mdx";
import VerifyProxyConnectionComponent from "../../../../../snippets/verify-proxy-connection-component.mdx";
import SupportParagraph from "../../../../../snippets/support-paragraph.mdx";
This guide will help you configure a Geonode proxy in Firefox for secure and anonymous browsing.
## Setting Up a Proxy in Firefox
### Open Firefox Settings
1. Click the **three-line menu (☰)** in the top-right corner of Firefox.
2. Select **Settings** from the dropdown menu.
***
### Access Network Settings
1. Scroll down to **Network Settings**.
2. Click **Settings** to open the proxy configuration panel.
***
### Configure Proxy Settings
1. In the **Connection Settings** popup, enter the following details:
* **Manual proxy configuration**: select this option.
* **HTTP Proxy**: enter your proxy IP.
* **Port**: enter your proxy port.
* **SOCKS Proxy** (optional): use SOCKS5 if applicable.
2. Click **OK** to save changes.
Firefox will now route all traffic through the configured proxy.
# GeeLark (/docs/proxies/getting-started/setup_and_configuration/browsers/geelark)
import BrowsersFaqs from "../../../../../snippets/browsers-faqs.mdx";
import ProxyInfoFromGeonode from "../../../../../snippets/get-proxy-info-from-geonode.mdx";
import SupportParagraph from "../../../../../snippets/support-paragraph.mdx";
This guide will help you configure a Geonode proxy in the GeeLark browser for both **mobile** and **desktop** profiles.
You’ll learn how to:
* Add and manage Geonode proxies in GeeLark
* Create mobile and desktop profiles
* Set up device environments
* Launch virtual profiles and test proxy connections
## Setting Up a Proxy in GeeLark
### Install GeeLark
1. Go to the [GeeLark Download Page](https://www.geelark.com/).
2. Download the version for your operating system.
3. Install the software and create an account.
***
### For Mobile Profiles
#### Create a New Profile
1. Open GeeLark and click **New Profile**.\
2. Enter the required details:
* Profile name
* Operating system
3. Review your browser details on the right side.\
***
#### Add a Proxy
You can add proxies in two ways:
##### Option A: Add Proxies First (Recommended)
1. Go to the **Proxies** tab from the sidebar.\
2. Click **Add Proxy**.\
3. Enter your proxy in one of the following formats:
```
proxy.geonode.io:9000:geonode_username:password
```
You can also use:
```
username:password@host:port
```
Or
```
http://username:password@proxy.geonode.io:9000
```
4. Select:
* **Type:** HTTP
* **Proxy group:** e.g., “Geonode Proxies”
* **IP Query Channel:** use `ip-api` for geolocation checks
5. Add one proxy per line (up to 100).
6. Test them using **Proxy Tests** to confirm connectivity — green icons indicate success.\
##### Option B: Add Proxy During Profile Creation
1. While creating a profile, go to the **Proxy** tab.
2. Choose the connection type (e.g., HTTP).
3. Enter:
* Proxy (IP:Port)
* Username
* Password
4. Set IP Query Channel (e.g., `ip-api`) and click **Check Proxy**.\
***
#### Configure Profile Settings
1. Under **Profile Settings**, set:
* Profile name
* Operating system (Android or iOS)
* Group, tags, remark (optional)\
2. Add the proxy:
* **Custom Proxy:** enter host, port, username, password
* **Saved Proxy:** choose from a pre-added list\
***
#### Configure Device Information
Under **Device Information**, you can simulate mobile hardware and network behavior:
* Charging Method: pay per minute or monthly
* Android version: e.g., 12–15
* Network: Wi-Fi or Cellular
* Phone number: auto or custom
* Area, Device Brand, Language: auto or manual
***
#### Create and Launch the Profile
1. Click **Create** to save the profile.
2. The profile appears in your list — showing OS, proxy region, tags, etc.
3. Click the **Action (▶️)** button to launch.
4. A new window opens — visit [ip-api.com](https://ip-api.com) to verify your IP and proxy location.\
***
### For Desktop Profiles
#### Create a New Desktop Profile
1. In GeeLark, click **New Profile**, then switch to the desktop icon (🖥️).
2. Under **Profile Settings**, configure:
* Profile name
* Group / tags / remark
* Operating system (Windows or macOS)
* Browser (e.g., Kiwi)
* User-Agent (auto or custom)
* Optional cookies for session import\
***
#### Set Proxy for Desktop Profile
You can either:
* Use a **Custom Proxy**, or
* Select a **Saved Proxy** (from the “Geonode Proxies” group).
***
#### Optional: Account Settings
Configure:
* Platform credentials (if needed)
* Startup behavior (e.g., reopen tabs, open custom page)\
***
#### Optional: Advanced Settings
Fine-tune fingerprinting and device behavior:
* Time zone, language, geolocation
* WebRTC, Canvas, WebGL, AudioContext
* Resolution, fonts, storage, noise controls
* Device hardware and network simulation\
***
#### Review Device Information
The right-hand panel summarizes fingerprint and environment settings:
* Browser, OS, User-Agent
* Time zone, WebRTC, language, Canvas, WebGL
* Fonts, resolution, and more
You can click **Generate New Fingerprint** to randomize your identity.\
***
#### Launch the Desktop Profile
1. Click **Create** to save.
2. The new profile appears in your list.
3. Click the **Action (▶️)** button to launch.
4. Visit [ip-api.com](https://ip-api.com) to confirm your Geonode proxy is active.\
# Ghost Browser (/docs/proxies/getting-started/setup_and_configuration/browsers/ghost)
import BrowsersFaqs from "../../../../../snippets/browsers-faqs.mdx";
import ProxyOs from "../../../../../snippets/proxy-os.mdx";
import ProxyInfoFromGeonode from "../../../../../snippets/get-proxy-info-from-geonode.mdx";
import SupportParagraph from "../../../../../snippets/support-paragraph.mdx";
This guide will help you configure a Geonode proxy in Ghost Browser.
## Setting Up a Proxy in Ghost Browser
### Install Ghost Browser
1. Go to the [Ghost Browser Download Page](https://ghostbrowser.com/download/).
2. Download the version for your operating system.
3. Install the software.
***
### Add a Proxy
1. Open **Proxy Control** and click **Add/Edit Proxies**.\
2. In the new window, select **Add a Single Proxy**.\
3. Enter your proxy details:
* Proxy (IP:Port)
* Username
* Password\
***
### Test the Proxy Connection
1. Confirm that your proxy appears in the **Proxy Management Table**.\
2. Click the **Test Proxy** tab.
3. Enter a website to verify your connection.\
4. If successful, a confirmation popup will appear.\
Your Geonode proxy is now successfully configured in Ghost Browser.
# GoLogin (/docs/proxies/getting-started/setup_and_configuration/browsers/gologin)
import BrowsersFaqs from "../../../../../snippets/browsers-faqs.mdx";
import ProxyInfoFromGeonode from "../../../../../snippets/get-proxy-info-from-geonode.mdx";
import SupportParagraph from "../../../../../snippets/support-paragraph.mdx";
This guide will help you configure a Geonode proxy in the GoLogin browser.
## Setting Up a Proxy in GoLogin
### Install GoLogin
1. Go to the [GoLogin Download Page](https://gologin.com/download-started/).
2. Download the version for your operating system.
3. Install the software and create an account.
***
### Create a New Profile
1. Open GoLogin and click **Add Profile**.\
2. Enter the required details:
* Profile name
* Operating system
3. (Optional) Randomize the fingerprint for extra security.
4. Review your browser details on the right.\
***
### Add Proxy Details
1. Go to the **Proxy** tab.
2. Choose the connection type — for this guide, select **HTTP**.
3. Enter your proxy credentials:
* Proxy (IP:Port)
* Username
* Password
4. Click **Create Profile** to save the settings.\
***
### Test the Proxy Connection
1. Click **Check Proxy** to verify the connection.
2. If successful, the proxy will show as active.\
***
### Launch the Profile
1. Once created, the profile will appear in your list.
* If the proxy displays a location, it means it’s active.\
2. Click **Start** to launch the browser with your configured proxy.
3. Visit [`http://ip-api.com/json`](http://ip-api.com/json) to verify your IP and proxy location.\
Your Geonode proxy is now successfully configured in GoLogin.
# Incognito Mode (/docs/proxies/getting-started/setup_and_configuration/browsers/incognito-mode)
import ProxyOs from "../../../../../snippets/proxy-os.mdx";
import BrowsersFaqs from "../../../../../snippets/browsers-faqs.mdx";
import ProxyInfoFromGeonode from "../../../../../snippets/get-proxy-info-from-geonode.mdx";
import VerifyProxyConnectionComponent from "../../../../../snippets/verify-proxy-connection-component.mdx";
import SupportParagraph from "../../../../../snippets/support-paragraph.mdx";
This guide will help you configure a Geonode proxy in **Chrome’s Incognito Mode**.
## Setting Up a Proxy in Incognito Mode
### Open Chrome Settings
1. Click the **three-dot menu** in the top-right corner of Chrome.
2. Select **Settings** from the dropdown menu.
***
### Access System Proxy Settings
1. In the left sidebar, click **System**.
2. Then click **Open your computer’s proxy settings**.
***
### Configure the Proxy on Your Operating System
Chrome will now open your system’s proxy configuration window.\
Follow the appropriate setup guide for your OS below:
Once configured, Chrome will automatically route all traffic—including Incognito Mode—through the assigned proxy.
# Incogniton (/docs/proxies/getting-started/setup_and_configuration/browsers/incogniton)
import BrowsersFaqs from "../../../../../snippets/browsers-faqs.mdx";
import ProxyInfoFromGeonode from "../../../../../snippets/get-proxy-info-from-geonode.mdx";
import SupportParagraph from "../../../../../snippets/support-paragraph.mdx";
This guide will help you configure a Geonode proxy in the Incogniton browser.
## Setting Up a Proxy in Incogniton
### Install Incogniton
1. Go to the [Incogniton Download Page](https://incogniton.com/download-incogniton/).
2. Download the version for your operating system.
3. Install the software and create an account.
***
### Create a New Profile
1. Open Incogniton and click **New Profile**.\
2. Fill in the required details:
* Profile name
* Operating system
* Browser version
3. (Optional) Randomize the fingerprint for added security.
4. Review your browser details on the right panel.\
***
### Add Proxy Settings
1. Go to the **Proxy Management** page.\
2. Choose the connection type — for this guide, select **HTTP**.\
3. Enter your proxy credentials:
* Proxy (IP:Port)
* Username
* Password\
***
### Check the Proxy Connection
1. Click **Check Proxy** to verify the connection.\
2. If successful, the proxy will show as active.
***
### Create and Launch the Profile
1. Click **Create Profile** to save your settings.\
2. Once created, you’ll see the profile in your list.
* If the proxy shows a green tick, it means it’s active.\
3. Click **Start** to launch the browser with your configured proxy.\
Your Geonode proxy is now successfully configured in Incogniton.
# MoreLogin (/docs/proxies/getting-started/setup_and_configuration/browsers/morelogin)
import BrowsersFaqs from "../../../../../snippets/browsers-faqs.mdx";
import ProxyInfoFromGeonode from "../../../../../snippets/get-proxy-info-from-geonode.mdx";
import SupportParagraph from "../../../../../snippets/support-paragraph.mdx";
This guide will help you configure a Geonode proxy in the MoreLogin browser.
## Setting Up a Proxy in MoreLogin
### Install MoreLogin
1. Go to the [MoreLogin Download Page](https://www.morelogin.com/).
2. Download the version for your operating system.
3. Install the software and create an account.
***
### Create a New Profile
1. Open MoreLogin and click **Add Profile**.\
2. Enter the required details:
* Profile name
* Operating system
3. (Optional) Randomize the fingerprint for added security.
4. Review your browser details on the right side.\
***
### Add Proxy Details
1. Go to the **Proxy** tab.
2. Choose the connection type — for this guide, select **HTTP**.
3. Enter your proxy credentials:
* Proxy (IP:Port)
* Username
* Password
4. Click **Create Profile** to save the configuration.\
***
### Test the Proxy Connection
1. Click **Proxy Detection** to verify the connection.
2. If successful, the proxy will show as active.\
***
### Launch the Profile
1. Click **Create Profile** to save your settings.
2. Once created, your profile will appear in the list.
* If the proxy shows a location, it means it’s active.\
3. Click **Start** to launch the browser with your configured proxy.\
Your Geonode proxy is now successfully configured in MoreLogin.
# MultiLogin (/docs/proxies/getting-started/setup_and_configuration/browsers/multilogin)
import BrowsersFaqs from "../../../../../snippets/browsers-faqs.mdx";
import ProxyInfoFromGeonode from "../../../../../snippets/get-proxy-info-from-geonode.mdx";
import SupportParagraph from "../../../../../snippets/support-paragraph.mdx";
This guide will help you configure a Geonode proxy in the MultiLogin browser.
## Setting Up a Proxy in MultiLogin
### Install MultiLogin
1. Go to the [MultiLogin Download Page](https://multilogin.com/).
2. Download the version for your operating system.
3. Install the software and create an account.
***
### Create a New Profile
1. Open MultiLogin and click **Add Profile**.\
2. Enter the required details:
* Profile name
* Operating system
3. (Optional) Randomize the fingerprint for extra security.
4. Review your browser details on the right side.\
***
### Add Proxy Details
1. Open the **Proxy** tab.
2. Choose the connection type — for this guide, select **HTTP**.
3. Enter your proxy credentials:
* Proxy (IP:Port)
* Username
* Password
4. Click **Create Profile** to save the configuration.\
***
### Test the Proxy Connection
1. Click **Proxy Detection** to verify the connection.
2. If successful, the proxy will show as active.\
***
### Launch the Profile
1. Click **Create Profile** to save your settings.
2. Once created, your profile will appear in the list.
* If the proxy shows a location, it means it’s active.\
3. Click **Start** to launch the browser with your configured proxy.\
Your Geonode proxy is now successfully configured in MultiLogin.
# Octo Browser (/docs/proxies/getting-started/setup_and_configuration/browsers/octo)
import BrowsersFaqs from "../../../../../snippets/browsers-faqs.mdx";
import ProxyInfoFromGeonode from "../../../../../snippets/get-proxy-info-from-geonode.mdx";
import SupportParagraph from "../../../../../snippets/support-paragraph.mdx";
This guide will help you configure a Geonode proxy in the Octo Browser.
## Setting Up a Proxy in Octo Browser
### Install Octo Browser
1. Go to the [Octo Browser Download Page](https://octobrowser.org/download/).
2. Download the version for your operating system.
3. Install the software and create an account.
***
### Create a New Profile
1. Open Octo Browser and click **Add Profile**.\
2. Enter the required details:
* Profile name
* Operating system
3. (Optional) Randomize the fingerprint for extra security.
4. Review your browser details.\
***
### Add Proxy Details
1. Click the **Proxy** button.\
2. Choose a connection type — for this guide, select **HTTP**.
3. Enter your proxy credentials:
* Proxy (IP:Port)
* Username
* Password\
4. Click **Check Proxy Connection** to test it.
5. Click **Confirm** to save your proxy settings.
***
### Launch the Profile
1. Once created, the profile will appear in your list.
* If the proxy shows a location, it means it’s active.\
2. Click **Start** to launch the browser with your configured proxy.
3. Visit [`http://ip-api.com`](http://ip-api.com) to verify your IP and proxy location.\
Your Geonode proxy is now successfully configured in Octo Browser.
# Safari (/docs/proxies/getting-started/setup_and_configuration/browsers/safari)
import BrowsersFaqs from "../../../../../snippets/browsers-faqs.mdx";
import ProxyInfoFromGeonode from "../../../../../snippets/get-proxy-info-from-geonode.mdx";
import VerifyProxyConnectionComponent from "../../../../../snippets/verify-proxy-connection-component.mdx";
import SupportParagraph from "../../../../../snippets/support-paragraph.mdx";
This guide will help you configure a Geonode proxy in Safari.
## Setting Up a Proxy in Safari
### Open Safari Settings
1. In the top menu bar, click **Safari → Settings** (or **Preferences** on older macOS versions).
2. Select the **Advanced** tab.
3. Click **Change Settings** next to **Proxies**.
***
### Configure the Proxy on macOS
1. The **Network** window will open in System Preferences.
2. Select your active network connection (Wi-Fi or Ethernet).
3. Click **Advanced → Proxies**.
4. Choose the protocol you want to configure (e.g., HTTP or HTTPS).
5. Enter your Geonode proxy credentials:
* **Proxy server:** IP and Port
* **Username / Password** if required
6. Click **OK**, then **Apply** to save your settings.
***
### Verify the Proxy Connection
Once configured, Safari will automatically route traffic through your Geonode proxy.\
You can verify that it’s working correctly below.
# FoxyProxy (/docs/proxies/getting-started/setup_and_configuration/extensions/foxyproxy)
import ProxyInfoFromGeonode from "../../../../../snippets/get-proxy-info-from-geonode.mdx";
import VerifyProxyConnectionComponent from "../../../../../snippets/verify-proxy-connection-component.mdx";
import SupportParagraph from "../../../../../snippets/support-paragraph.mdx";
import ExtensionFAQs from "../../../../../snippets/extensions-faqs.mdx";
This guide explains how to install, configure, and use **Geonode proxies** in the **FoxyProxy** browser extension.
***
## What Is FoxyProxy
FoxyProxy is a browser extension that lets you easily manage and switch between multiple proxy configurations.\
Instead of manually entering proxy details each time, you can save and control proxies directly within your browser.
***
***
## Steps to Set Up Geonode Proxy in FoxyProxy
Follow these steps to configure your Geonode proxy using the FoxyProxy extension.
***
### Step 1: Install the FoxyProxy Extension
1. Open the [FoxyProxy extension page](https://chromewebstore.google.com/detail/foxyproxy/gcknhkkoolaabfmlnjonogaaifnjlfnp) in the Chrome Web Store.
2. Click **Add to Chrome** to install the extension.
3. Confirm the installation by clicking **Add Extension** when prompted.
***
### Step 2: Pin the Extension for Quick Access
After installation, pin the extension to your Chrome toolbar for faster access.
1. Click the **Extensions** icon (puzzle piece) in Chrome.
2. Find **FoxyProxy** and click the **Pin** icon.
***
### Step 3: Open FoxyProxy Settings
1. Click the **FoxyProxy** icon in your Chrome toolbar.
2. Select **Options** to open the settings panel.
3. You’ll be redirected to the **Proxy Manager** page.\
Click on the **Proxies** tab.
***
### Step 4: Add a New Proxy
1. Click **Add New Proxy**.
2. Fill in your proxy details in the popup window:
* **Name**: Any label for easy identification
* **Host**: Your proxy server (e.g., `proxy.geonode.io`)
* **Port**: Usually `9000`
* **Username** and **Password**: Your Geonode credentials
3. Click **Add Proxy** to save the configuration.
4. Return to the **Proxy Manager** to confirm that your new proxy appears in the list.
***
### Step 5: Connect to the Proxy
1. Click the **FoxyProxy** icon in the Chrome toolbar.
2. Select the proxy you added from the list.
3. Click **Connect** to activate the proxy.\
Once connected, all your browser traffic will route through the selected proxy.
***
***
***
# Geonode Proxy Manager (/docs/proxies/getting-started/setup_and_configuration/extensions/geonode-proxy-manager)
import ProxyInfoFromGeonode from "../../../../../snippets/get-proxy-info-from-geonode.mdx";
import VerifyProxyConnectionComponent from "../../../../../snippets/verify-proxy-connection-component.mdx";
import SupportParagraph from "../../../../../snippets/support-paragraph.mdx";
This guide explains how to install, configure, and use the **Geonode Proxy Manager** Chrome extension.
***
## What Is Geonode Proxy Manager
Geonode Proxy Manager is a Chrome extension that simplifies proxy management.\
You can add, save, and switch between multiple proxies directly from your browser without manually entering details each time.
***
***
## Setting Up a Proxy in Geonode Proxy Manager
Follow these steps to configure your proxy using the Geonode Proxy Manager extension.
***
### Step 1: Install the Geonode Proxy Manager Extension
1. Open the [Geonode Proxy Manager extension page](https://chromewebstore.google.com/detail/geonode-proxy-manager/ippioaknloonmaibmmepkemhmhinohge).
2. Click **Add to Chrome** to install the extension.
3. Confirm installation by selecting **Add Extension** when prompted.
***
### Step 2: Pin the Extension for Quick Access
After installation, pin the extension to your Chrome toolbar for easy access.
1. Click the **Extensions** icon (puzzle piece) in Chrome.
2. Find **Geonode Proxy Manager** and click the **Pin** icon.
***
### Step 3: Open the Extension
1. Click the **Geonode Proxy Manager** icon in your Chrome toolbar.
2. If no proxies have been added yet, click **Add New Proxy**.
3. You’ll be redirected to the Proxy Manager page.
***
### Step 4: Add a New Proxy
1. Click **Add New Proxy**.
2. Enter your proxy details in the popup:
* **Name** — any recognizable label for this proxy
* **Host** — e.g., `proxy.geonode.io`
* **Port** — usually `9000`
* **Username** and **Password** — your Geonode credentials
3. Click **Add Proxy** to save your configuration.
4. Once saved, go back to the Proxy Manager to confirm that your new proxy appears in the list.
***
### Step 5: Connect to the Proxy
1. Click the **Geonode Proxy Manager** icon again in your Chrome toolbar.
2. Select the proxy you want to use from the list.
3. Click **Connect** to activate the proxy.\
Once connected, all browser traffic will be routed through the selected proxy.
***
***
***
## FAQs
Yes, you can add multiple proxies and switch between them by selecting the
desired one from the list and clicking **Connect**.
It saves time by storing proxy details for quick switching, enhances privacy
by routing browser traffic through different servers, and supports multiple
proxies for varied use cases.
Check your IP address using an online tool like [IP API](https://ip-api.com/)
or follow this guide: [Verify Proxy
Connection](/docs/proxies/getting-started/setup_and_configuration/verify-proxy-connection).
Make sure you entered the correct proxy details (host, port, username,
password), try reconnecting, and check that your proxy plan is active in the
Geonode Dashboard.
Yes, but some sites may block proxies. If that happens, switch to a different
server or location.
Yes, but using a VPN and a proxy simultaneously may cause slower speeds or
connection conflicts.
Yes, it’s free to install, but you’ll need an active Geonode Proxy Plan to use
proxy services.
# Android (/docs/proxies/getting-started/setup_and_configuration/mobile-OS/android)
import ProxyInfoFromGeonode from "../../../../../snippets/get-proxy-info-from-geonode.mdx";
import VerifyProxyConnectionComponent from "../../../../../snippets/verify-proxy-connection-component.mdx";
import BrowsersFaqs from "../../../../../snippets/browsers-faqs.mdx";
import SupportParagraph from "../../../../../snippets/support-paragraph.mdx";
import PopupAuth from "../../../../../snippets/popup-auth.mdx";
Android settings may look slightly different depending on the device model and
Android version, but the overall process remains the same.
***
***
## Step-by-step setup for Android
### Step 1 — Open Settings
1. Tap the **Settings** app (gear icon).
***
### Step 2 — Go to Wi-Fi Settings
1. Scroll down and tap **Network & Internet** or **Wi-Fi**, depending on your device.
2. Tap and hold the Wi-Fi network you’re connected to.
3. Select **Modify network** or tap the gear (⚙) or “i” icon next to the Wi-Fi name.
Ensure Wi-Fi is turned on and connected to the network you want to configure
the proxy for.
***
### Step 3 — Access Wi-Fi List
You will now see a list of available Wi-Fi networks.
***
### Step 4 — Find Proxy Settings
1. Scroll down until you find **Proxy settings**.
2. Tap **Proxy** to expand available options.
You’ll also see your current IP address — write it down for future reference.
***
### Step 5 — Choose Manual Configuration
Select **Manual** from the proxy settings.
***
### Step 6 — Enter Proxy Details
1. Enter the **Proxy IP Address** and **Port** from your Geonode Dashboard.
***
### Step 7 — Save and Exit
1. Scroll down and tap **Save** to apply the settings.
2. If your device doesn’t have a Save button, simply exit — the configuration will be applied automatically.
***
***
***
# iOS (/docs/proxies/getting-started/setup_and_configuration/mobile-OS/ios)
import ProxyInfoFromGeonode from "../../../../../snippets/get-proxy-info-from-geonode.mdx";
import VerifyProxyConnectionComponent from "../../../../../snippets/verify-proxy-connection-component.mdx";
import BrowsersFaqs from "../../../../../snippets/browsers-faqs.mdx";
import SupportParagraph from "../../../../../snippets/support-paragraph.mdx";
import PopupAuth from "../../../../../snippets/popup-auth.mdx";
***
***
## Step-by-step setup for iOS
Follow these steps to configure your proxy manually:
### Step 1 — Open iPhone Settings
1. Open the **Settings** app.
2. Tap **Wi-Fi** to view available networks.
***
### Step 2 — Access Wi-Fi Network Settings
1. Connect to the Wi-Fi network you want to configure.
2. Tap the **“i” icon** next to the connected network.
***
### Step 3 — Navigate to Proxy Settings
1. Scroll down to the **HTTP Proxy** section.
2. You will see three options:
* **Off** — disables proxy
* **Manual** — enter proxy details manually
* **Automatic** — configure using a PAC file
Select **Manual** for your Geonode proxy setup.
***
### Step 4 — Enter Proxy Details
1. Under **Manual Proxy Configuration**, fill in:
* **Server:** Proxy IP address
* **Port:** Port number from your Geonode Dashboard
2. If authentication is required, enable **Authentication** and enter your **username** and **password**.
3. Tap **Save** to apply the settings.
***
***
***