# Change Log (/docs/changelog) After months of work, we’re excited to introduce our **updated dashboard** — a major release focused on simplicity, speed, and flexibility.\ Below is an overview of the key improvements. *** ## Redesigned Dashboard The dashboard has been completely reworked for better usability and performance. * Clean, minimal layout * Faster navigation * Intuitive access to essential tools Everything is now more accessible and built with your feedback in mind. Dashboard Interface *** ## Proxy Membership We’ve replaced the old **Proxy Residential Membership** with the new **Proxy Membership** — a unified and flexible way to manage all your proxies. * Manage both **Residential** and **Datacenter** IPs * Filter or combine IP types effortlessly * Simplified structure for all users For API integration details, see the [API documentation](https://docs.geonode.com/). IP Type Filtering *** ## Endpoint Generator Our most requested tool is here — the **Endpoint Generator**.
### Create endpoints in multiple formats Multiple formats are supported for easy integration. ### Copy or export instantly Copy endpoints or export them to your preferred format in one click. ### Integrate with your favorite tools No manual setup — just plug it in and start using it.
Endpoint Generator *** ## Updated Host URL We’ve updated our host structure to improve consistency. **Old host:** `premium-residential.geonode.com` **New host:** `proxy.geonode.io` The old hostname will continue to function, but only the new one will appear in the dashboard moving forward. *** ## Feature Request System Your feedback drives our development.\ With the new **Feature Request Tool**, you can: * Submit new feature ideas directly from your dashboard * Vote for the improvements you want most You’ll find it under **Profile → Feature Request**. Feature Request Tool *** ## Changelog Access You’re reading it!\ From now on, every platform update and feature release will appear here, so you can always stay informed. *** ## Revamped API Documentation The [API documentation](https://docs.geonode.com/) has been fully updated for better clarity and ease of use. * Improved structure and examples * Updated endpoints * More guides coming soon *** ## Under-the-Hood Improvements We’ve made significant backend enhancements to boost speed, reliability, and stability. | Area | Improvement | | --------------- | --------------------------------------------- | | **Speed** | Faster performance across dashboard and APIs | | **Latency** | Reduced response times for smoother operation | | **Reliability** | Higher success rates for proxy connections | | **Stability** | Optimized infrastructure and bug fixes | These upgrades work silently to ensure your experience is faster and more dependable. *** ## Try It Out [Log in now →](https://app.geonode.com/residential-proxies)\ Explore the new dashboard and let us know what you think. *** *Thank you for being part of our journey — more great updates are on the way!* # May 2025 (/docs/changelog/may2025) After months of hard work, we're thrilled to introduce our **updated dashboard** and a major platform upgrade.\ This release focuses on speed, reliability, and a smoother user experience. *** ## New Website & Branding The new **Geonode website and visual identity** mark a major milestone in our journey.\ This redesign reflects our growth and focuses on developers — with a cleaner layout, faster load times, and improved accessibility. *** ## Proxy Infrastructure 2.0 We’ve launched **Proxy Infrastructure 2.0** — a full rebuild of our global proxy network. * Significantly reduced latency * 99 % + success rate across all regions * Major speed boost for **SOCKS5** users * Expanded IP pool in key markets like the United States This upgrade delivers faster, more stable connections and improved reliability worldwide. *** ## Upgraded SOCKS5 Infrastructure Our SOCKS5 network has been completely re-engineered for improved throughput and connection stability.\ You’ll notice faster response times and fewer dropped sessions across all endpoints. *** ## New Pricing Plans We’ve introduced three new plans to better fit different use cases: | Plan | Monthly Price | Included GB | Rate per GB | | ------------ | ------------- | ----------- | ----------- | | **Starter** | $50 / month | 50 GB | $1.00 | | **Growth** | $200 / month | 267 GB | $0.75 | | **Business** | $500 / month | 1000 GB | $0.50 | *** ## 1 TB Free Trial in Dashboard Qualified business users can now apply for the **1 TB free trial** directly inside the dashboard — no sales call required.\ Get access, test at scale, and start evaluating performance in minutes. *** ## User-Friendly Billing & Grace Period Billing and subscription handling have been redesigned for flexibility: * Failed payment? You now have **5 days** to resolve it before cancellation. * Cancelled subscription? You can reactivate anytime before the billing period ends. * Finished plan? You still have **5 days** to restore it and keep unused bandwidth. *** ## Improved Usage Graphs Usage analytics now include: * Bandwidth by **hour, day, week, or month** * Look-back period of up to 31 days * Clearer trend visualization and comparison These updates make monitoring and optimization much easier. *** ## Error Handling and Service Codes Error handling has been simplified and clarified: * Cleaner error messages * Fewer, more actionable error codes See the [updated Error Handling Guide](https://docs.geonode.com/docs/proxies/api-reference/error-handling). *** ## Geonode SDK Program Launch We’ve officially launched the **Geonode SDK Program** — enabling developers to monetize apps across iOS, Android, desktop, smart TVs, and more.\ [Learn more and apply here](https://geonode.com/app-sdk). *** ## Expanded Event-Based Email Notifications You’ll now receive clear, event-based notifications for: * Purchases and upgrades * Downgrades and usage warnings * Trial updates and billing issues Stay informed about every important account event. *** ## Shape the Future of Geonode Want to help define what comes next?\ [Vote on and suggest new features](https://feedback.geonode.com/) directly through our feedback portal. # Feature Requests (/docs/getting-started/feature-request) Have an idea for a new Geonode feature or an improvement to an existing one?\ You can easily submit your request and explore what other users are suggesting. *** ## Submit Your Feature Request To request a new feature or enhancement, visit:\ ➡️ [Geonode Feature Requests](https://feedback.geonode.com/) Provide as much detail as possible about the feature, including what problem it solves and how it would improve your workflow. *** ## Why Submit a Feature Request? * 🧠 **Influence future updates** – Help shape Geonode’s roadmap based on real user needs. * ⚙️ **Improve functionality** – Suggest enhancements that boost performance, security, or usability. * 💬 **Collaborate with the community** – Upvote and comment on ideas from other users. *** ## How It Works 1. **Browse existing requests**\ Someone may have already suggested your idea — check first before creating a new one. 2. **Submit a new request**\ Provide a clear title, description, and (if possible) your use case or expected benefit. 3. **Vote and comment**\ Upvote ideas you support and share feedback to help prioritize development. 4. **Track progress**\ Follow the status of your requests as the Geonode team reviews and implements new features. *** ## Final Notes Geonode values community input — every feature request helps make the platform better.\ Your ideas directly contribute to improving **performance, reliability, and user experience** for everyone. # Getting Started (/docs/getting-started) Welcome to Geonode. Use this section to learn the product basics and set up your account. ## Start here * [What is Geonode?](/docs/getting-started/what-is-geonode/what-is-geonode) — how Geonode works and who it is for * [Account setting and billing](/docs/getting-started/account-setting-and-billing/profile-setting) — profile, subscriptions, payments, and wallet ## Next steps * [Proxies](/docs/proxies) — set up proxies and use the Proxy API * [Web Data](/docs/scraper-api) — extract, crawl, map, and search web content # Service Status (/docs/getting-started/service-status) Stay informed about the current operational status of **Geonode services**, including proxy availability, system uptime, and scheduled maintenance. *** ## Service Status Overview 🟢 **Coming Soon**\ A real-time **Service Status Dashboard** is in development.\ It will let you: * Monitor **system health** in real time * View **proxy uptime and latency metrics** * Track **maintenance windows and past incidents** * Subscribe to **status updates and alerts** *** ## Need Help Right Now? If you’re currently experiencing connection issues or service interruptions, please reach out to our support team: ➡️ [Contact Geonode Support](https://geonode.com/contact) Our team is available to help diagnose and resolve issues as quickly as possible. *** ## Stay Updated Check back soon for the live **status.geonode.com** page —\ your one-stop hub for Geonode uptime, maintenance, and performance insights. # Subscription, Cancellations, Reactivation and Grace Period (/docs/getting-started/subscription-cancellations-reactivation-and-grace-period) This guide explains what happens when you **cancel your Geonode subscription**, how you can **reactivate it**, and what **grace periods** apply.\ Understanding these timelines helps prevent loss of access or unused bandwidth. *** ## Canceling During Your Billing Period When you cancel your subscription, it remains active until the **end of your current billing cycle**. * You can continue using the service (including proxy access and any remaining bandwidth) until the period ends. * You can **reactivate anytime before the billing cycle ends** — this cancels your cancellation and keeps your subscription active without interruption. *** ## After Your Billing Period Ends Once your billing cycle ends and the subscription is fully canceled: * You automatically enter a **5-day grace period**. * During this time, you can **reactivate your subscription** without losing data. ### If you reactivate within the grace period: * Your subscription is **restored immediately**. * Any **unused bandwidth** is reinstated. * Proxy access resumes without delay. ### If you do **not** reactivate within 5 days: * The subscription remains permanently canceled. * Any **unused bandwidth is lost** and cannot be recovered. *** ## Failed Payments If your payment fails: * Access to proxies is **temporarily paused**. * You have **5 days** to update your payment method and fix the issue. * If resolved within 5 days, your subscription continues as normal. * If not resolved, your subscription may be **canceled** and remaining bandwidth **forfeited**. *** ## Summary | Scenario | Grace Period | Can Reactivate? | Bandwidth Restored? | Access Status | | ------------------------------- | ------------ | --------------- | ------------------- | ------------------ | | **Manual Cancellation** | ✅ 5 days | ✅ Yes | ✅ Yes | Temporarily Paused | | **End of Billing Period** | ✅ 5 days | ✅ Yes | ✅ Yes | Inactive | | **Failed Payment (Unresolved)** | ❌ | ❌ No | ❌ No | Canceled | *** If you need help managing your subscription or reactivating access, please contact:\ ➡️ [Geonode Support](https://geonode.com/contact) # Support (/docs/getting-started/support) ## Geonode Community and Support Join the Geonode Community to learn, share, and connect with proxy users at all experience levels.\ Get access to helpful guides, expert tips, and support from the Geonode team. 💬 [Join us on Discord](https://discord.com/invite/32RXgzgeAf) to get started, meet other users, and stay up to date with the latest in proxy services. *** ## Contact Methods If you need help or want to reach our support team, you can contact us using one of the methods below. ### Email Reach out to us directly at [hello@geonode.com](mailto:hello@geonode.com).\ Use this for general inquiries, technical issues, or billing questions. *** ### Contact Form Submit your request using our [Contact Form](https://geonode.com/contact).\ Recommended for non-urgent questions, suggestions, or feedback. *** ### Book a Call Schedule a 30-minute support call with our team through [Calendly](https://calendly.com/maria-geonode/30min).\ Perfect for detailed troubleshooting or guided onboarding. *** Our support team is available around the clock to help with technical issues, billing, or account management. # Useful Links (/docs/getting-started/useful-links) This page provides a quick overview of key Geonode resources — manage your proxies, stay updated, and get support whenever needed. *** ## Geonode Dashboard Your main control panel for managing proxies, monitoring usage, and configuring account settings.\ ➡️ [Access the Geonode Dashboard](https://app.geonode.com/proxies) *** ## Discord Community Join the Geonode community to connect with other users, get real-time help, and receive product updates.\ ➡️ [Join the Geonode Discord](https://discord.com/invite/32RXgzgeAf) *** ## Geonode Blog Read articles, tutorials, and insights from the Geonode team to enhance your proxy experience.\ ➡️ [Visit the Blog](https://geonode.com/blog) *** ## Contact Support Need help? Reach out to the Geonode support team for personalized assistance.\ ➡️ [Contact Support](https://geonode.com/contact) *** ## Subscription Plans Compare available proxy plans and find the best option for your needs.\ ➡️ [View Subscription Plans](https://geonode.com/1TB-program) *** > \[!NOTE]\ > Keep this page bookmarked for quick access to key Geonode resources.\ > Our support team is available 24/7 to help with any questions or issues. # Proxies (/docs/proxies) Use this section to set up Geonode proxies and work with the Proxy API. ## Getting Started Start here if you are new to Geonode proxies. * [Quick Start Guide](/docs/proxies/getting-started/quick-start) — connect your first proxy * [Prerequisites](/docs/proxies/getting-started/prerequisites/access-credentials) — credentials and proxy server details * [Setup and Configuration](/docs/proxies/getting-started/setup_and_configuration/desktop-OS/windows) — desktop, mobile, browsers, and tools * [Knowledge Base](/docs/proxies/getting-started/knowledge-base/overview) — proxy concepts and service basics ## Products Choose the proxy product that matches your use case. * [Residential Proxies](/docs/proxies/guides/residential-proxies/overview) — real residential IPs for geo targeting, sticky sessions, and rotating traffic * [Unlimited Residential Proxies](/docs/proxies/guides/unlimited-residential-proxies/00_unlimited_residential_proxies) — speed-based residential plans with unlimited traffic * [ISP Proxies](/docs/proxies/guides/isp-proxies/isp-proxies) — ISP-assigned proxy IPs you can manage, organize, and monitor from the dashboard * Rotating Datacenter Proxies — high-speed datacenter IPs for large-scale rotating traffic ## API Reference Use the Proxy API to target locations, manage sticky sessions, and track usage. * [Introduction](/docs/proxies/api-reference) — authentication and basic workflow * [Geo Targeting](/docs/proxies/api-reference/geo-targeting/geo-targeting-options) — country, state, city, ISP, and OS targeting * [Sticky Sessions](/docs/proxies/api-reference/sticky-session/get-sticky-session/get-create) — create and release sticky sessions * [Error Handling](/docs/proxies/api-reference/error-handling) — common API errors and responses ## Next Steps * Start with the [Quick Start Guide](/docs/proxies/getting-started/quick-start) * Pick a product guide above for your use case # Overview (/docs/scraper-api) Use this section to extract web content, discover URLs, run search jobs, and connect Geonode to AI tools with MCP. ## Getting Started Start here if you are new to the Scraper API. * [Quick Start Guide](/docs/scraper-api/quick-start) — get your API key and authenticate requests * [Before You Start](/docs/scraper-api/getting-started/00_before_you_start) — prerequisites and setup basics ## Products Choose the API that matches your workflow. * [Extraction](/docs/scraper-api/guides/extraction/01_understanding_extraction) — extract Markdown or HTML from a webpage * [Batch](/docs/scraper-api/guides/batch/00_understanding_batch) — process multiple URLs in one job * [Crawl](/docs/scraper-api/guides/crawl/00_overview) — discover and extract pages across a website * [Map](/docs/scraper-api/guides/map/00_understanding_map) — discover URLs under a base URL * [Search](/docs/scraper-api/guides/search/01_search_overview) — submit search queries and retrieve results ## Dashboard Guides Use the Geonode Dashboard when you want to run jobs without writing code. * [Dashboard Overview](/docs/scraper-api/dashboard-guides/overview) — navigate Scraper, Map, and Search in the UI ## MCP You can also use Geonode through MCP to connect the Scraper API to AI assistants and IDEs. See the [MCP Guides](/docs/scraper-api/guides/mcp/00_overview) to get started. ## API Reference Use the API reference when you need endpoint details, parameters, and responses. * [API Overview](/docs/scraper-api/v1) — Scraper API reference entry point * [Extraction](/docs/scraper-api/v1/extraction/extract-content) — extract content endpoints * [Batch](/docs/scraper-api/v1/batch/start-batch-job) — batch job endpoints * [Crawl](/docs/scraper-api/v1/crawl/start-crawl-job) — crawl job endpoints * [Map](/docs/scraper-api/v1/map/map-urls) — map job endpoints * [Search](/docs/scraper-api/v1/search/start-search-job) — search job endpoints ## Next Steps * Start with the [Quick Start Guide](/docs/scraper-api/quick-start) to authenticate * Pick a product guide above for your use case # Quick Start Guide (/docs/scraper-api/quick-start) This guide helps you get started with the Geonode Scraper API in a few minutes. ## Get Your API Key 1. Sign in to your Geonode account. 2. Open the Dashboard. 3. Navigate to the API Keys section. 4. Create or copy an existing API key. API & Integrations Modal ## Authentication All Scraper API requests require the `X-Api-Key` header. ```bash title="request.sh" curl -H "X-Api-Key: YOUR_API_KEY" ``` Replace `YOUR_API_KEY` with your actual API key. ## Choose an API | API | Use When | | ---------- | ------------------------------------------------------- | | Extraction | You want to extract content from one or more webpages | | Batch | You have multiple URLs to process in a single job | | Crawl | You want to discover and extract pages across a website | | Webhooks | You want notifications when jobs complete | ## Next Steps Choose the API that matches your use case and follow the guides in that section. # Active Subscriptions (/docs/getting-started/account-setting-and-billing/active-subscriptions) ## Step 1 — Open Profile Settings 1. Log in to your Geonode account. 2. Go to **Profile Settings** in the left-hand sidebar. Profile Settings *** ## Step 2 — View your active subscription 1. In the Profile Settings page, select **Active Subscriptions**. 2. You’ll see details of your current plan, including billing cycle and renewal date. Active Subscriptions *** ## Tips * Keep your subscription active to maintain uninterrupted access. * Review renewal dates regularly to avoid service pauses. * To upgrade or cancel, use the **Billing** section in your dashboard. * For any payment or renewal issues, contact [Geonode Support](https://geonode.com/contact). Managing subscriptions in Geonode is quick and transparent — everything you need is in your profile settings. # Payment Methods (/docs/getting-started/account-setting-and-billing/payments-method) Geonode lets you add, update, and manage your payment methods for smooth and secure billing. ## Step 1 — Access the Payments section 1. Log in to your Geonode account. 2. From the left sidebar, open **Payments**. Payments Section *** ## Step 2 — Add a new payment method 1. Click **Add Payment Method**. 2. A pop-up window will appear for card details. ### Required information * **Card Number** (Visa, Mastercard, or AMEX) * **Expiration Date** (MM/YY) * **Security Code (CVC)** — 3-digit code on the back of your card * **Billing Country** 3. Click **Add** to save your payment method. Add New Card *** ## Step 3 — Update billing information 1. In the **Billing Information** section, click **Update Information**. 2. Update fields such as: * Name * Company (optional) * Billing address * Phone number 3. Click **Save Changes** to confirm updates. *** ## Step 4 — View payment history In the **Payment History** section, you can track: * Transaction date * Transaction details * Amount charged * Invoice records If no payments have been made, you’ll see *No Data Yet*. *** ## Tips * Keep billing details accurate to avoid failed payments. * Supported payment types: **Visa**, **Mastercard**, **AMEX**. * For payment or billing issues, contact [Geonode Support](https://geonode.com/contact). Managing payment methods in Geonode ensures secure and uninterrupted service access. # Profile Settings (/docs/getting-started/account-setting-and-billing/profile-setting) Your profile settings let you manage personal information, change passwords, and, if needed, delete your account. ## Step 1 — Access profile settings 1. Log in to your Geonode account. 2. Open **Profile** from the left-hand sidebar. Profile Settings 3. View detailed profile information. Profile Details *** ## Step 2 — Update personal details In **Personal Details**, you can: * Edit your first and last name. * Update your phone number. * Note: your email address cannot be changed. Click **Save changes** after updating. *** ## Step 3 — Change your password 1. Scroll to the **Password** section. 2. Enter and confirm a new password. 3. Click **Update password**. Change Password *** ## Step 4 — Delete your account If you wish to permanently remove your account: 1. Scroll to **Delete Account**. 2. Click **Delete**. Delete button 3. Confirm the action — all data will be removed. Confirm Delete Deleting your account does **not** automatically cancel any active services. Cancel subscriptions or contact Geonode Support before deletion to stop future charges. *** ## Final tips * Keep profile information current to avoid issues. * Use a strong password for better security. * Account deletion is irreversible. * For billing or payment concerns, contact [Geonode Support](https://geonode.com/contact). Managing your profile properly helps maintain a smooth, secure experience in Geonode. # Referral Program (/docs/getting-started/account-setting-and-billing/referral-program) The Geonode Referral Program lets you earn a 10% commission every time a new user signs up and makes a purchase through your unique affiliate link. ## How to get your referral link Follow these steps to access and start sharing your link. ### Step 1 — Open Profile Settings 1. Log in to your Geonode account. 2. Go to **Profile Settings** in the left sidebar. Profile Settings *** ### Step 2 — Copy your referral link 1. In the **Profile Settings** page, find the **Referral Program** section. 2. Copy your unique affiliate link. Referral Link *** ## How the referral program works * Share your referral link anywhere: social media, websites, or direct messages. * When a user registers and makes a purchase using your link, you earn **10% commission**. * Commissions are tracked and displayed in your referral dashboard. *** ## Tips * Share your link across different platforms for better reach. * More referrals mean higher earnings. * Monitor your performance and payouts in the referral dashboard. * For any payment or tracking issues, contact [Geonode Support](https://geonode.com/contact). Start earning today with Geonode’s Referral Program. # Reset API Password (/docs/getting-started/account-setting-and-billing/reset-api-password) If you need to reset your API password, follow these steps to generate a new one securely.\ Resetting your API password immediately deactivates the old one — don’t forget to update your integrations afterward. ## Step 1 — Access API credentials 1. Log in to your Geonode account. 2. Open the **Proxy Configuration** tab in the dashboard. 3. Locate the **API Credentials** section. API Credentials *** ## Step 2 — Generate a new API password 1. Click the **reset icon** next to your current API password. 2. A confirmation prompt will appear. 3. Click **Confirm** to generate a new password. Reset API Password Confirmation *** ## Tips * Resetting your API password deactivates the old one instantly. * Update all connected tools, scripts, and automations with the new password. * You can reset it anytime if you forget or lose access. * For any billing or payment issues, contact [Geonode Support](https://geonode.com/contact). Managing your API credentials properly ensures secure and uninterrupted access to Geonode services. # Wallet (/docs/getting-started/account-setting-and-billing/wallet-ocerview) The Geonode Wallet lets you view your balance, review transactions, and add funds for seamless proxy usage. ## Step 1 — Access the Wallet 1. Log in to your Geonode account. 2. Open **Wallet** from the left sidebar. Wallet Navigation *** ## Step 2 — Check your balance * Your current wallet balance appears at the top. * Previous transactions are shown under **Transactions from wallet**. If you haven’t made any payments yet, the transaction list will be empty. *** ## Step 3 — Add funds 1. In the **Add Funds** section, enter the desired amount. 2. Click **Top Up Wallet** to start the payment process. 3. Once the payment is confirmed, the amount will appear in your wallet balance. Add Funds *** ## Tips * Keep your wallet funded to prevent service interruptions. * Review your transaction history regularly for clarity and tracking. * For any billing or payment issues, contact [Geonode Support](https://geonode.com/contact). Using the Geonode Wallet makes managing payments simple and ensures uninterrupted access to all services. # Billing & Payments (/docs/getting-started/faqs/billing-and-payments) *** {" "} All invoices are in USD, we are unable to offer other currencies. {" "} If you fail to update your payment information before the next billing cycle, you may encounter service interruptions. {" "} After completing the payment update process, your next billing cycle will be automatically processed. {" "} To view your billing history and invoices: 1. Click on the **Profile** tab. 2. Navigate to **Account Settings**. 3. Go to the **Payment and Wallet** section. 4\. You will find your **billing history** and invoices listed there. {" "} Yes, you can add your credit card details through the Billing tab. {" "} Yes, but you must ensure that your credit card is not set as the default card. If you need extra support, our team is always available to help. # General FAQs (/docs/getting-started/faqs/general-faqs) *** {" "} We do not currently offer static IPs. We are planning to release this service in the coming months! {" "} No. Any hacking/cracking or illegal activity is strictly forbidden on our platform. {" "} We do not recommend our current services for streaming, you are welcome to try and use them for this purpose. {" "} We do not block any websites, including survey sites. # Subscriptions & Cancellations (/docs/getting-started/faqs/subscriptions-and-cancellations) *** {" "} Yes, you can cancel anytime. You can still use our service until the end of the billing period if you have unused bandwidth. {" "} Unfortunately, we can't extend subscriptions as we will encounter a break in the system. However, we will find the right solution for you to compensate for the downtime, so please reach out to our support team for manual assistance. {" "} Yes, you can add multiple cards, but only one will be set as the default card. {" "} You can easily re-activate your subscription through your dashboard or reach out to our support team for help. # Technical Issues & Support (/docs/getting-started/faqs/technical-issues-and-support) *** {" "} Our team is working on resolving any occurring issues as soon as possible. We will inform our users once everything is functioning normally. {" "} If your issue is not listed or you need further assistance, please [contact our support team](https://geonode.com/contact). We are available 24/7 to help you with any issues. # Proxy Related Queries (/docs/proxies/additional-resources/faqs) *** {" "} Sticky proxies have a maximum timeout of 60 minutes. Static proxies are the type of proxies that do not expire, which we do not currently offer. We recommend setting auto-replace to true and choosing a lengthy enough rotating interval for sticky connections when clients bring up the offline issue. {" "} We do not recommend our current services for streaming, but you are welcome to try. {" "} We do not block any websites, including survey sites. You will see popups everywhere asking for your proxy username and password. {" "} You can whitelist **up to 150 IP addresses**. {" "} A proxy might slightly affect speed, but Geonode's high-speed servers minimize this impact. {" "} Yes, but you must configure it separately for each network type. {" "} There are no restrictions on the number of simultaneous connections. # Core Features of Geonode (/docs/getting-started/what-is-geonode/core-features-of-geonode) When it comes to managing your online activity securely and efficiently, Geonode provides the tools you need.\ Here’s what it offers for both individual users and businesses. *** ## 1. Proxy products Websites constantly track online activity — Geonode’s proxy network gives you control again. With Geonode, you can: * Browse anonymously — hide your IP and location to stay private * Access blocked content — bypass geo-restrictions and visit region-limited websites Geonode acts as a secure, private gateway to the web. *** ## 2. Scraper tools and data collection Collecting data from multiple sources often leads to blocks. Geonode prevents that by rotating IPs automatically. * Collect data efficiently — extract information without interruptions * Avoid blocks — stay under detection limits through IP rotation Ideal for businesses, researchers, and anyone working with large-scale web data. *** ## 3. Advanced tools Managing multiple accounts or troubleshooting network issues can be complex.\ Geonode’s advanced tools simplify these tasks. * Manage multiple accounts — operate several profiles safely at once * Detect and fix issues quickly — identify and resolve connection problems easily These tools are designed for both beginners and professionals. *** ## 4. Global reach Geonode keeps you connected anywhere — from Singapore to New York or Tokyo. * Access global content — open region-specific websites and services * Choose IPs worldwide — select IPs from diverse countries for testing, marketing, or research With Geonode, you don’t just use the internet — you control how you connect to it. # Use Cases and Applications (/docs/getting-started/what-is-geonode/uses-cases-and-applications) Geonode isn’t just a proxy provider — it’s a flexible tool for diverse online workflows.\ Here are a few key ways you can use it effectively: *** ## Web scraping and data collection Extract valuable data from websites for research, business analytics, and market insights.\ Geonode ensures stable access and prevents IP bans, so your data collection runs smoothly. *** ## Digital marketing and SEO Analyze competitor strategies, monitor performance, and gather location-specific results.\ Geonode helps marketers and SEO specialists test campaigns and view search results from any region. *** ## Cybersecurity and network testing Test firewalls, identify vulnerabilities, and verify network configurations securely.\ Geonode assists in assessing your infrastructure without exposing your real IP or internal endpoints. *** These are only a few ways Geonode supports technical and business workflows.\ [Explore more use cases →](https://geonode.com/use-cases) # What is Geonode? (/docs/getting-started/what-is-geonode/what-is-geonode) Geonode is a proxy service that helps you bypass internet restrictions by masking your real IP address. It makes your online activity private, secure, and unrestricted. *** ## How Geonode works Geonode provides scalable proxy solutions to keep your connection stable and private. Main proxy types: * Residential proxies — real IPs from real devices for undetectable browsing * Datacenter proxies — high-speed, cost-efficient options for large data operations *** ## Problems Geonode solves Common challenges users face online: * Access and data collection — scrape or gather data without getting blocked * Privacy and security — protect sensitive information while browsing or automating tasks *** ## Why choose Geonode Geonode stands out through reliability and flexibility. * Unlimited proxy endpoints — generate and manage as many as needed * Simple dashboard — intuitive design for fast configuration * Global coverage — millions of IPs across worldwide locations *** ## Who should use this guide This guide is suitable for all users — beginners and experienced alike.\ It explains how to use Geonode effectively to enhance privacy, access, and productivity online. # Retrieve Available Geo-locations (/docs/proxies/api-reference/available-geo-locations) Retrieve a list of available geo-locations, including countries, cities, states, and ASNs that you can use for geo-targeting your proxy requests. ## Supported Service Types This endpoint supports the following proxy services: | Service Type | Description | | ----------------------- | ----------------------------- | | `RESIDENTIAL-PREMIUM` | Premium residential proxies | | `ROTATING-DATACENTER` | Rotating datacenter proxies | | `UNLIMITED-RESIDENTIAL` | Unlimited residential proxies | Replace the service type in the request URL with the one you want to query. ## Request ### Premium Residential ```bash curl -X GET "https://monitor.geonode.com/services/RESIDENTIAL-PREMIUM/targeting-options" \ -u geonode_username:password ``` ### Rotating Datacenter ```bash curl -X GET "https://monitor.geonode.com/services/ROTATING-DATACENTER/targeting-options" \ -u geonode_username:password ``` ### Unlimited Residential ```bash curl -X GET "https://monitor.geonode.com/services/UNLIMITED-RESIDENTIAL/targeting-options" \ -u geonode_username:password ``` ## Response ### 200 Success Returns the available geo-targeting options for the selected proxy service. ### Response Structure The response is an array of country objects. | Field | Type | Description | | ---------------- | ------ | -------------------------------------------------------- | | `code` | string | ISO 3166-1 alpha-2 country code | | `name` | string | Full country name | | `cities` | object | Available city-level targeting options | | `cities.prefix` | string | Prefix used for city targeting (for example, `-city-`) | | `cities.options` | array | Available cities | | `states` | object | Available state-level targeting options | | `states.prefix` | string | Prefix used for state targeting (for example, `-state-`) | | `states.options` | array | Available states | | `asns` | object | Available ASN-level targeting options | | `asns.prefix` | string | Prefix used for ASN targeting (for example, `-asn-`) | | `asns.options` | array | Available ASNs | ### Example Response ```json [ { "code": "AF", "name": "Afghanistan", "cities": { "prefix": "-city-", "options": [ { "code": "kabul", "name": "Kabul" } ] }, "states": { "prefix": "-state-", "options": [ { "code": "kabul", "name": "Kabul" } ] }, "asns": { "prefix": "-asn-", "options": [ { "code": "131284", "name": "AS131284 Etisalat Afghan" } ] } } ] ``` # Error Handling (/docs/proxies/api-reference/error-handling) When using the Geonode Proxy API, you may encounter various HTTP status codes. Understanding these codes helps you handle errors gracefully and implement robust error handling in your applications. ## HTTP Status Codes | Status Code | Description | Expanded Description | | ----------- | -------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | 403 | Invalid request configuration. . | This error occurs when the request contains invalid or improperly formatted parameters. The placeholder will include details about which field caused the issue. Common causes include incorrect formatting or unsupported values for fields like country, city, state, or ISP. Please verify that all parameters follow the expected structure and use values listed in the API documentation or supported geo-targeting options. | | 407 | Authentication error. Please check your authentication settings. | This error indicates that authentication with the proxy server failed. Authentication is required, but the credentials provided were missing, invalid, or insufficient. Common causes include incorrect or missing proxy username/password, misconfigured authentication headers, or using an IP address that hasn't been whitelisted. Please verify that your credentials are correct and that your IP address is authorized to access the proxy if applicable. Refer to the authentication section of the API documentation for setup instructions and troubleshooting tips. | | 411 | Your account has been blocked. If you think it is a mistake, please contact customer support. | This error means that the user's account has been manually blocked by the system or an administrator. This action is typically taken due to violations of the terms of service, suspicious activity, billing issues, or abuse prevention measures. If you believe this is an error, please contact customer support to review your account status and resolve the issue. | | 464 | Connection to the specified target is not permitted due to security policies. | This error occurs when the requested connection to a specific host, IP address, or port is blocked due to security or access control policies. Common reasons include attempting to connect to restricted destinations, using unsupported ports, or using protocols that are not allowed (e.g., FTP or SMTP). Ensure that the target address, port, and protocol are supported and comply with the platform's usage policies. If you're unsure or believe the restriction is incorrect, please contact support for clarification. | | 465 | No proxies available in the selected location. Please choose a different targeting configuration or try again later. | This error indicates that the proxy server was unable to find any available IP addresses that meet the specific geo-targeting requirements (e.g., country, city, ISP) specified in the request. This can occur when demand exceeds supply in a particular region or when targeting criteria are too narrow. To resolve this issue, try adjusting your targeting configuration to be less restrictive or retry your request later. If consistent access to specific regions is critical, please contact support to explore custom proxy allocations or availability options. | | 466 | You've reached your bandwidth limit. Buy more data or upgrade your plan to continue. | This error indicates that the user has consumed all of the bandwidth allocated in their current plan. When the bandwidth limit is reached, no further requests can be processed until more data is added. To resume usage, you can upgrade your plan or purchase an add-on directly in the user dashboard. | | 467 | You have reached your session bandwidth limit. | Specified bandwidth limit for this session has been reached. Limit is defined by -limit- parameter of the first session request and can not be changed until session is expired or released. | | 468 | Target Forbidden. | The proxy refused this destination because it matches an entry on your customer [Block list](/docs/proxies/guides/residential-proxies/block-list) (domain, IP, or wildcard). Remove the entry if you need the proxy to fetch that host again. This is different from **464**, which is a platform security policy rather than your own Block list. | | 500 | Internal error. Please try again. | This error indicates that something went wrong on the server while processing the request. It is not caused by user input or configuration. Please try again. If the issue persists, contact support with your request details so we can investigate. | | 517 | Internal error. Please try again. | This error indicates that something went wrong on the server while processing the request. It is not caused by user input or configuration. Please try again. If the issue persists, contact support with your request details so we can investigate. | | 518 | Internal error. Please try again. | This error indicates that something went wrong on the server while processing the request. It is not caused by user input or configuration. Please try again. If the issue persists, contact support with your request details so we can investigate. | | 560 | Internal error. Please try again. | This error indicates that something went wrong on the server while processing the request. It is not caused by user input or configuration. Please try again. If the issue persists, contact support with your request details so we can investigate. | | 561 | Internal error. Please try again. | This error indicates that something went wrong on the server while processing the request. It is not caused by user input or configuration. Please try again. If the issue persists, contact support with your request details so we can investigate. | | 562 | Internal error. Please try again. | This error indicates that something went wrong on the server while processing the request. It is not caused by user input or configuration. Please try again. If the issue persists, contact support with your request details so we can investigate. | | 563 | Internal error. Please try again. | This error indicates that something went wrong on the server while processing the request. It is not caused by user input or configuration. Please try again. If the issue persists, contact support with your request details so we can investigate. | | 564 | Internal error. Please try again. | This error indicates that something went wrong on the server while processing the request. It is not caused by user input or configuration. Please try again. If the issue persists, contact support with your request details so we can investigate. | | 565 | Internal error. Please try again. | This error indicates that something went wrong on the server while processing the request. It is not caused by user input or configuration. Please try again. If the issue persists, contact support with your request details so we can investigate. | | 566 | Internal error. Please try again. | This error indicates that something went wrong on the server while processing the request. It is not caused by user input or configuration. Please try again. If the issue persists, contact support with your request details so we can investigate. | | 567 | Internal error. Please try again. | This error indicates that something went wrong on the server while processing the request. It is not caused by user input or configuration. Please try again. If the issue persists, contact support with your request details so we can investigate. | | 569 | Internal error. Please try again. | This error indicates that something went wrong on the server while processing the request. It is not caused by user input or configuration. Please try again. If the issue persists, contact support with your request details so we can investigate. | ## Best Practices 1. **Always Implement Error Handling** * Never assume requests will succeed * Handle both expected and unexpected errors 2. **Use Retry Logic** * Implement exponential backoff for 5xx errors * Respect rate limits and Retry-After headers 3. **Log Errors Appropriately** * Include relevant request details * Don't log sensitive information 4. **User Feedback** * Provide clear error messages to end users * Include actionable steps for resolution ## Support If you encounter persistent errors or need assistance, contact our support team at [Support](https://geonode.com/contact). *** # Introduction (/docs/proxies/api-reference) Welcome to the Geonode Proxy API! This API allows you to: * Target specific geolocations (countries, states, cities, and even ISPs). * Manage sticky sessions to keep a consistent IP across multiple requests. * Track usage statistics (e.g., bandwidth, session counts). * Filter proxy types (residential, data center, mobile).
This guide will help you make your first calls, retrieve essential information, and move on to more advanced features—without overwhelming you. ## Prerequisites: * API Credentials: **Our API uses Basic Authentication, requiring a proxy username and password for access. You must include these credentials in every request using the Authorization header.** *(Base64-encoded string)* * Service Name: Indicate which specific service plan or tier you're using. > Note: *Some sections provide a static cURL example rather than a live, testable endpoint—so you won't be able to send requests directly from this page. If you want to try it out, simply copy the cURL command into your terminal (or HTTP client) and insert your real credentials or parameters.* Please refer to our [user dashboard](https://app.geonode.com/) to find the information mentioned above.
## Basic Workflow: * Authenticate: Include your proxy username and password in each request using Basic Authentication (via the Authorization header). * Specify Your Service: Use the appropriate service name in each call so the API applies the correct proxy settings. * Send Requests: Interact using standard HTTP methods (GET, POST, PUT, etc.). * Parse JSON Responses: You'll generally receive JSON objects containing the requested data or error details.
For detailed information about handling API errors and implementing robust error handling, please refer to our [Error Handling Guide](/docs/proxies/api-reference/error-handling). ***
## Next Steps: * Explore the Full Reference: Dive deeper to learn how to configure proxy sessions, target specific regions, or gather usage insights. * Check Plan Limits: Monitor bandwidth and session counts to stay within your plan's capacity. * Stay Informed: Check our Changelog for updates & and new features. Share your opinion on feature requests. * Contact Support: If you have any questions or run into issues, reach out to our support team at [hello@geonode.com](mailto:hello@geonode.com). *** # IP Type Filtering (/docs/proxies/api-reference/ip-filter) Filter proxy IPs by type to match your specific requirements. You can choose between residential IPs, datacenter IPs, or a mix of both. You can filter IPs by type using the `-type-` parameter in your username: * **Residential**: Real IPs from home users * **Datacenter**: IPs from data centers * **Mixed**: Combination of both types ## Request ```bash curl -x "http://proxy.geonode.io:" \ --user "-type--country-:" \ --url "http://ip-api.com/json" \ --header "Accept: application/json" ``` ## Response ### 200 Success Successfully filtered IPs based on type. #### Response Fields | Field | Type | Description | | --------------- | ------- | ------------------------------------------------------ | | `status` | string | The status of the request (e.g., "success") | | `continent` | string | The continent where the IP is located | | `continentCode` | string | The continent code | | `country` | string | The country where the IP is registered | | `countryCode` | string | The country code in ISO 3166-1 alpha-2 format | | `region` | string | The regional subdivision (state/province) | | `regionName` | string | The full name of the region | | `city` | string | The city associated with the IP address | | `district` | string | The district or subdivision of the city | | `zip` | string | The postal or ZIP code of the location | | `lat` | number | Latitude coordinate of the location | | `lon` | number | Longitude coordinate of the location | | `timezone` | string | Time zone in which the IP is located | | `offset` | integer | Time offset from UTC in seconds | | `currency` | string | Local currency used in the country | | `isp` | string | The name of the Internet Service Provider (ISP) | | `org` | string | The name of the organization associated with the IP | | `as` | string | The Autonomous System (AS) number and name | | `asname` | string | The full Autonomous System (AS) name | | `mobile` | boolean | Indicates whether the IP is from a mobile network | | `proxy` | boolean | Indicates whether the IP is being used as a proxy | | `hosting` | boolean | Indicates whether the IP belongs to a hosting provider | | `query` | string | The IP address queried in the request | # Retrieve Usage Statistics (/docs/proxies/api-reference/usage-statistics) Retrieve bandwidth usage statistics for your Geonode proxy service account. This endpoint provides detailed information about your data consumption. ## Request ```bash curl -X GET "https://monitor.geonode.com/monitor-light/proxies" \ -H "Authorization: Basic base64(username:password)" ``` ## Response ### 200 Success Usage statistics retrieved successfully. #### Response Fields | Field | Type | Description | | ------------------------------------------ | ------- | ------------------------------------------------------ | | `data` | object | Container for usage statistics | | `data.bandwidth` | object | Bandwidth usage information | | `data.bandwidth.data` | object | Detailed bandwidth data | | `data.bandwidth.data.default` | integer | Total bandwidth used in bytes | | `data.bandwidth.data.currentFastBandwidth` | integer | Current fast bandwidth usage (unlimited services only) | | `data.bandwidth.data.totalBandwidthInGB` | integer | Total bandwidth used in gigabytes | #### Example Response ```json { "data": { "bandwidth": { "data": { "default": 17930356, "currentFastBandwidth": 0, "totalBandwidthInGB": 2 } } } } ``` # Quick Start Guide (/docs/proxies/getting-started/quick-start) import BrowsersFaqs from "../../../snippets/browsers-faqs.mdx"; import ProxyInfoFromGeonode from "../../../snippets/get-proxy-info-from-geonode.mdx"; import SupportParagraph from "../../../snippets/support-paragraph.mdx"; *** ## 1. Access the User Dashboard Start by accessing your user dashboard — this is where you manage all proxy settings. * Open your Dashboard. * Scroll down to Proxy Configuration. Accessing the User Dashboard *** ## 2. Get Proxy Credentials from Geonode Obtain your authentication credentials directly from the dashboard. * **Username:** Copy your unique API username. * **Password:** Copy your API password. Get Your User Credentials *** ## 3. Configure Proxy Parameters Within the Proxy Configuration section, adjust the parameters as needed. | Parameter | Example | Description | | ------------------- | ------------------------------------------------ | --------------------------------------------------------- | | **Endpoint Format** | `hostname:port:username:password` | Connection string format | | **IP Type** | `-type-datacenter` | Choose between Datacenter or Residential | | **Gateway** | `192.155.103.209` | Proxy gateway | | **Geo-Targeting** | `-country-jp`, `-state-tokyo`,
`-as-2501` | Specify region or ASN (state and city cannot be combined) | | **Protocol** | HTTP / HTTPS | Default port 9000 | | **Session Type** | — | Rotate or Sticky sessions | | **Output Format** | — | Customize response format | Configure Proxy Parameters See the full guide:\ [How to use the Endpoint Generator to configure proxy parameters](/docs/proxies/getting-started/knowledge-base/geo-targeting) *** ## 4. Make Your First API Call Once your endpoint is configured, generate your first API request. 1. Go to **API Code Generator** in the dashboard. 2. Select your configuration settings. 3. Copy the generated code snippet. Access the API Code Generator *** ## 5. Configure Proxy in Different Browsers or OS To integrate your proxy on specific platforms, follow one of the setup guides below. ### Desktop Operating Systems * [Windows 10/11](/docs/proxies/getting-started/setup_and_configuration/desktop-OS/windows) * [macOS](/docs/proxies/getting-started/setup_and_configuration/desktop-OS/macOS) ### Mobile Operating Systems * [Android](/docs/proxies/getting-started/setup_and_configuration/mobile-OS/android) * [iOS](/docs/proxies/getting-started/setup_and_configuration/mobile-OS/ios) ### Browsers * [Chrome](/docs/proxies/getting-started/setup_and_configuration/browsers/chrome) * [Edge](/docs/proxies/getting-started/setup_and_configuration/browsers/edge) * [Brave](/docs/proxies/getting-started/setup_and_configuration/browsers/brave) * [Firefox](/docs/proxies/getting-started/setup_and_configuration/browsers/firefox) * [Safari](/docs/proxies/getting-started/setup_and_configuration/browsers/safari) * [Incognito Mode](/docs/proxies/getting-started/setup_and_configuration/browsers/incognito-mode) * [Incogniton](/docs/proxies/getting-started/setup_and_configuration/browsers/incogniton) * [AdsPower](/docs/proxies/getting-started/setup_and_configuration/browsers/adspower) * [Ghost Browser](/docs/proxies/getting-started/setup_and_configuration/browsers/ghost) * [GoLogin](/docs/proxies/getting-started/setup_and_configuration/browsers/gologin) * [Dolphin Anty](/docs/proxies/getting-started/setup_and_configuration/browsers/dolphin-anty) * [ClonBrowser](/docs/proxies/getting-started/setup_and_configuration/browsers/clonbrowser) * [MoreLogin](/docs/proxies/getting-started/setup_and_configuration/browsers/morelogin) * [MultiLogin](/docs/proxies/getting-started/setup_and_configuration/browsers/multilogin) ### Extensions * [FoxyProxy](/docs/proxies/getting-started/setup_and_configuration/extensions/foxyproxy) * [Geonode Proxy Manager](/docs/proxies/getting-started/setup_and_configuration/extensions/geonode-proxy-manager) *** ## Conclusion You’re all set!\ Your Geonode proxy is now configured and ready for use across browsers, operating systems, and API integrations. *** *** # Quick Start Guide (/docs/proxies/guides) import BrowsersFaqs from "../../../snippets/browsers-faqs.mdx"; import ProxyInfoFromGeonode from "../../../snippets/get-proxy-info-from-geonode.mdx"; import SupportParagraph from "../../../snippets/support-paragraph.mdx"; *** ## 1. Access the User Dashboard Start by accessing your user dashboard — this is where you manage all proxy settings. * Open your Dashboard. * Scroll down to Proxy Configuration. Accessing the User Dashboard *** ## 2. Get Proxy Credentials from Geonode Obtain your authentication credentials directly from the dashboard. * **Username:** Copy your unique API username. * **Password:** Copy your API password. Get Your User Credentials *** ## 3. Configure Proxy Parameters Within the Proxy Configuration section, adjust the parameters as needed. | Parameter | Example | Description | | ------------------- | ------------------------------------------------ | --------------------------------------------------------- | | **Endpoint Format** | `hostname:port:username:password` | Connection string format | | **IP Type** | `-type-datacenter` | Choose between Datacenter or Residential | | **Gateway** | `192.155.103.209` | Proxy gateway | | **Geo-Targeting** | `-country-jp`, `-state-tokyo`,
`-as-2501` | Specify region or ASN (state and city cannot be combined) | | **Protocol** | HTTP / HTTPS | Default port 9000 | | **Session Type** | — | Rotate or Sticky sessions | | **Output Format** | — | Customize response format | Configure Proxy Parameters See the full guide:\ [How to use the Endpoint Generator to configure proxy parameters](/docs/proxies/getting-started/knowledge-base/geo-targeting) *** ## 4. Make Your First API Call Once your endpoint is configured, generate your first API request. 1. Go to **API Code Generator** in the dashboard. 2. Select your configuration settings. 3. Copy the generated code snippet. Access the API Code Generator *** ## 5. Configure Proxy in Different Browsers or OS To integrate your proxy on specific platforms, follow one of the setup guides below. ### Desktop Operating Systems * [Windows 10/11](/docs/proxies/getting-started/setup_and_configuration/desktop-OS/windows) * [macOS](/docs/proxies/getting-started/setup_and_configuration/desktop-OS/macOS) ### Mobile Operating Systems * [Android](/docs/proxies/getting-started/setup_and_configuration/mobile-OS/android) * [iOS](/docs/proxies/getting-started/setup_and_configuration/mobile-OS/ios) ### Browsers * [Chrome](/docs/proxies/getting-started/setup_and_configuration/browsers/chrome) * [Edge](/docs/proxies/getting-started/setup_and_configuration/browsers/edge) * [Brave](/docs/proxies/getting-started/setup_and_configuration/browsers/brave) * [Firefox](/docs/proxies/getting-started/setup_and_configuration/browsers/firefox) * [Safari](/docs/proxies/getting-started/setup_and_configuration/browsers/safari) * [Incognito Mode](/docs/proxies/getting-started/setup_and_configuration/browsers/incognito-mode) * [Incogniton](/docs/proxies/getting-started/setup_and_configuration/browsers/incogniton) * [AdsPower](/docs/proxies/getting-started/setup_and_configuration/browsers/adspower) * [Ghost Browser](/docs/proxies/getting-started/setup_and_configuration/browsers/ghost) * [GoLogin](/docs/proxies/getting-started/setup_and_configuration/browsers/gologin) * [Dolphin Anty](/docs/proxies/getting-started/setup_and_configuration/browsers/dolphin-anty) * [ClonBrowser](/docs/proxies/getting-started/setup_and_configuration/browsers/clonbrowser) * [MoreLogin](/docs/proxies/getting-started/setup_and_configuration/browsers/morelogin) * [MultiLogin](/docs/proxies/getting-started/setup_and_configuration/browsers/multilogin) ### Extensions * [FoxyProxy](/docs/proxies/getting-started/setup_and_configuration/extensions/foxyproxy) * [Geonode Proxy Manager](/docs/proxies/getting-started/setup_and_configuration/extensions/geonode-proxy-manager) *** ## Conclusion You’re all set!\ Your Geonode proxy is now configured and ready for use across browsers, operating systems, and API integrations. *** *** # Choosing the Right Scraper API Plan (/docs/scraper-api/additional-resources/choosing_scraper_api_plan) import Link from "next/link"; Geonode Scraper API offers two pricing models: * **Request-Based** pricing, where successful page extractions use requests from your available balance. * **Unlimited** plans, where you pay a flat monthly price and your plan determines the number of concurrent threads. This guide helps you decide which model is better for your workload. ## Quick Comparison | | Request-Based | Unlimited | | ------------------- | ------------------------------------ | ------------------------- | | Pricing | Based on request volume | Flat monthly price | | Monthly requests | Plan-dependent | Unlimited | | Main pricing factor | Successful extractions | Concurrent threads | | Cost | Based on your request plan and usage | Predictable monthly price | | Best for | Variable or smaller workloads | High-volume workloads | | Batch extraction | Supported | Supported | | Crawl | Supported | Supported | If you already know which pricing model you need, you can go directly to the detailed guide:

Request-Based

Pay based on successful page extractions and choose the request capacity that fits your workload.

View Request-Based Pricing

Unlimited

Get unlimited requests with a fixed monthly price based on concurrent threads.

View Unlimited Pricing
## Choose Request-Based Pricing If Request-Based pricing is a good fit when your extraction volume changes from month to month or you do not need a large amount of concurrent processing. Consider Request-Based pricing if: * Your scraping volume is relatively small. * Your workload is unpredictable. * You only run extraction jobs occasionally. * You want your pricing to be based on request usage. * You are testing or starting a new scraping workflow. * You do not need a high number of concurrent threads. Request-Based pricing also includes a free tier with **1,500 requests per month**, which makes it suitable for getting started without immediately choosing a paid plan. Explore Request-Based Pricing ## Choose Unlimited If Unlimited plans are designed for workloads that require a larger amount of extraction capacity and predictable monthly pricing. Consider an Unlimited plan if: * You regularly process a large number of pages. * You need to run many extractions in parallel. * Your application can take advantage of higher concurrency. * You want a predictable monthly cost. * You regularly run large Batch or Crawl jobs. * You do not want to manage a monthly request balance. Unlimited plans are priced according to concurrent threads. For example, the available plans provide different concurrency levels: | Plan | Concurrent threads | | -------- | -----------------: | | Starter | 2 | | Growth | 10 | | Scale | 25 | | Pro | 50 | | Business | 100 | When all available threads are busy, additional work waits until a thread becomes available. Explore Unlimited Pricing ## Compare Common Workloads Use your expected workload to determine which pricing model is more appropriate. | Workload | Recommended | | ----------------------------------- | ----------------- | | Testing the Scraper API | Request-Based | | Small or occasional extraction jobs | Request-Based | | Variable monthly scraping volume | Request-Based | | Regular production scraping | Depends on volume | | Large Batch jobs | Unlimited | | Large Crawl jobs | Unlimited | | High-volume scraping | Unlimited | | Many parallel extraction tasks | Unlimited | | Predictable monthly scraping costs | Unlimited | These recommendations are based on the difference between the two pricing models. Your actual choice should depend on both your request volume and the amount of parallel processing your application requires. ## Consider Concurrency Concurrency is particularly important when comparing plans. A higher concurrency level allows more extraction tasks to run at the same time. For example: | Concurrent threads | Parallel extraction capacity | | -----------------: | ---------------------------: | | 2 | Up to 2 extractions | | 10 | Up to 10 extractions | | 25 | Up to 25 extractions | | 50 | Up to 50 extractions | | 100 | Up to 100 extractions | If your application processes URLs sequentially, increasing concurrency may not provide much benefit. If your application can submit many extraction tasks at the same time, a higher-concurrency Unlimited plan can provide greater throughput. ## Consider Batch and Crawl Jobs Batch and Crawl jobs have their own job-size limits. These limits apply to an individual job and are separate from the monthly request allowance on request-based plans or the unlimited request volume on Unlimited plans. For Unlimited plans, the job limits are: | Plan | Max URLs per batch | Max pages per crawl | | -------- | -----------------: | ------------------: | | Starter | 50 | 50 | | Growth | 500 | 500 | | Scale | 1,000 | 1,000 | | Pro | 2,000 | 2,000 | | Business | 4,000 | 4,000 | If a workload is larger than the maximum size of a single job, split it into multiple jobs. ## Example Scenarios ### You are testing the API If you are testing a new extraction workflow or only need a small number of requests, start with Request-Based pricing. The free tier provides **1,500 requests per month**. ### You scrape occasionally If you only run extraction jobs when you need specific data and your monthly volume varies, Request-Based pricing can be a better fit. You pay according to the request capacity you need instead of committing to a fixed Unlimited concurrency level. ### You process thousands of pages regularly If your application regularly processes thousands of pages and can run extraction tasks in parallel, an Unlimited plan may be a better fit. The main consideration becomes how many concurrent threads your workload needs. ### You run large Batch or Crawl jobs If your application regularly runs large Batch or Crawl jobs, consider an Unlimited plan with enough concurrency and job capacity for your workload. You can split larger workloads across multiple jobs when required. ## A Simple Decision Use this as a quick starting point:
### Choose Request-Based if * Your volume is small or unpredictable. * You want to pay based on request usage. * You do not need high concurrency. * You are testing or starting a project. ### Choose Unlimited if * Your volume is consistently high. * You need many extractions to run in parallel. * You want predictable monthly pricing. * You regularly process large scraping workloads.
## You Can Change Plans as Your Workload Grows Your initial choice does not need to be permanent. You can start with Request-Based pricing while evaluating your workload. If your extraction volume grows or you need more parallel processing, you can move to an Unlimited plan. Similarly, if your workload does not require high concurrency, a request-based plan may be more appropriate. The best plan is the one that matches your actual workload rather than simply choosing the plan with the highest limits. ## Related Guides
  • Request-Based Pricing
  • Unlimited Scraper API Pricing
  • Extract a Single URL
  • Extract Multiple URLs
# FAQs (/docs/scraper-api/additional-resources/faq) The Scraper API is a hosted extraction API. You send it a URL, and it returns clean Markdown or HTML from the target page. It also handles proxy routing, geo-targeting, JavaScript rendering, and anti-bot handling. One credit means one Scraper API request. In these docs, you'll usually see the word request instead of credit because the dashboard and pricing model are request-based. One successful page extraction uses one request. Batch jobs count each successfully extracted URL as one request, and crawl jobs count each successfully extracted page as one request. There are no extra request multipliers for JavaScript rendering, proxy type, geo-targeting, or requesting both Markdown and HTML. When you use all free requests for the billing month, new extraction requests stop working until you add paid requests, upgrade to a subscription, or wait for the next monthly renewal. The API can return `402` when your account does not have enough available requests. No. Free tier requests renew every billing month and expire at the end of that month. Unused free requests do not roll over. Yes. Unused subscription requests roll over to the next month with no rollover cap while your subscription remains active. No. Pay as you Go requests do not expire. They are a good fit when your scraping workload is occasional or bursty. No. JavaScript rendering does not cost extra requests. A successful extraction with `render_js: true` still uses one request. JavaScript rendering can take longer, so it is best to start with `render_js: false` and enable it only when the returned content is incomplete. With proxies directly, you still have to write the scraper, manage retries, choose when to run a browser, parse HTML, handle noisy pages, and normalize output. The Scraper API sits above that work. You send a URL and get extracted Markdown or HTML back. It is a better fit when you want page content quickly and do not want to maintain browser or proxy orchestration yourself. Direct proxies are still useful when you need full control over the browser, request flow, cookies, sessions, or a custom scraper pipeline. You can use the Scraper API with public webpages that your account is allowed to access. It works well with static pages, documentation pages, articles, product pages, and many JavaScript-rendered pages. Complex sites can still return noisy output. Some sites include large navigation menus, tracking links, ads, images, recommendations, cookie banners, or challenge pages in the extracted content. In those cases, inspect the returned Markdown or HTML and add post-processing if your workflow needs cleaner fields. Always follow the target site's terms, applicable laws, and your own compliance requirements. The most common reason is that the page loads content with JavaScript after the initial HTML response. Try the same request with `render_js: true`. ```json { "url": "https://quotes.toscrape.com/js/", "formats": ["markdown"], "render_js": true } ``` If the result is still incomplete, the site may require a user interaction, a longer wait condition, a login, or a site-specific scraping strategy. The API extracts raw page content. Many modern sites include navigation, filters, image data, tracking links, cookie banners, recommendations, and footer content in the page itself. The API may return some of that content because it exists in the source page. For LLM, search, or analytics workflows, it is a good idea to test a few pages from the same target site and add your own cleanup step if needed. Yes. Pass `proxy.country` with a two-letter ISO country code. ```json { "proxy": { "country": "US", "type": "residential" } } ``` If you omit `proxy`, the API applies default residential proxy routing and tries to infer a useful country from the target URL when possible. The extraction endpoint supports `residential`, `datacenter`, and `mix`. ```json { "proxy": { "country": "US", "type": "residential" } } ``` Yes. Use the `headers` object for headers that should be included in the target extraction request. ```json { "url": "https://example.com", "formats": ["markdown"], "headers": { "Accept-Language": "en-US,en;q=0.9" } } ``` Do not send your Geonode API key inside this object. Your API key belongs in the `X-Api-Key` header sent to the Scraper API. The public extraction schema supports custom headers, but it does not provide a full browser session management interface for logging in, clicking through flows, or maintaining user state across many pages. You can use the `headers` field to send authentication cookies or tokens with each extraction request. If your use case requires authenticated scraping beyond what custom headers can provide, contact support so the team can recommend the right setup. Use `sync` mode for quick pages when you want the result in the same HTTP response. Use `async` mode for slower pages, JavaScript-heavy pages, and workflows where you want to start the extraction and poll for the result later. The map endpoint discovers URLs under a base URL by reading sitemaps and links from the seed page. ```json { "url": "https://quotes.toscrape.com/", "search": "author", "include_subdomains": false, "ignore_query_parameters": true } ``` The `search` field filters URLs the API already discovered. It does not run a Google search or query an external search engine. Open the Geonode dashboard and check your Scraper API request balance. Use the dashboard UI as the source of truth for your current balance and plan details. Read the full Scraper API reference here: [Open the Scraper API Reference](/docs/scraper-api/v1/reference) # Pricing and Requests (/docs/scraper-api/additional-resources/pricing-and-requests) Scraper API supports both **request-based pricing** and **Unlimited plans**. Request-based plans charge per successful page extraction, while Unlimited plans use concurrency (threads) instead of request balances to determine throughput. ## Request Counting The request model is intentionally simple for page extraction work: * One page extraction equals one request. * A batch job counts each successfully extracted URL as one request. * A crawl job counts each successfully extracted page as one request. * JavaScript rendering does not cost extra requests. * Requesting Markdown, HTML, or both does not cost extra requests. * Geo-targeting does not cost extra requests. * Failed extraction requests are not charged. * Authentication errors, validation errors, and requests rejected before extraction are not charged. Some dashboard, API, or billing surfaces may still use older token or credit naming, such as `tokens_charged`, `estimated_tokens`, `tokens_charged_total`, or `tokens_reserved`. For Scraper API billing, read those values as charged, estimated, or reserved requests. *** ## Endpoint Usage Not every API endpoint extracts a page. Use this table to understand how each endpoint category relates to request balance. | Endpoint category | Request balance behavior | | --------------------------------------------------- | ----------------------------------------------------------------------------------------------------------- | | Single-page extraction | A successful `POST /v1/extract` extraction uses one request. | | Async extraction | The completed extraction uses one request when the page is successfully extracted. | | Batch extraction | Each successfully extracted URL in the batch uses one request. | | Crawl jobs | Each successfully extracted crawled page uses one request. | | Map | Discovers URLs only. It does not extract page content. | | Statistics, webhooks, job lookup, and health checks | These endpoints do not extract page content. They are used to manage, inspect, or monitor Scraper API work. | *** ## Free Tier The free tier includes **1,500 requests every month**. You do not need a credit card to start using the free tier. Free tier requests: * Renew every billing month. * Expire at the end of the billing month. * Do not roll over to the next month. * Stop working when the monthly free allowance runs out unless you upgrade or add paid requests. *** ## Subscription Plans Subscription plans include a monthly request allowance. Unused subscription requests roll over to the next month with no rollover cap while your subscription remains active. If your workload requires unlimited requests, Geonode also offers **Unlimited plans**, which use concurrency-based pricing instead of request balances. Use the Geonode dashboard to see the currently available subscription tiers, included monthly requests, and current prices. The dashboard is the source of truth for plan availability and billing details. Subscription requests: * Renew every billing month. * Roll over when unused. * Stay available while the subscription remains active. * Can be combined with Pay as you Go requests if your account has both balances. *** ## Unlimited Plans Unlimited plans charge a flat monthly price with no request counting. Instead of a request balance, each plan sets the number of concurrent threads—how many extractions can run in parallel at any moment. * No request allowance and no top-ups — total monthly volume is not capped. * No multipliers — JavaScript rendering, geo-targeting, and output format never change the price. * Throughput scales with your plan's thread count — more threads means more pages processed in parallel. Use the Geonode dashboard to view the available Unlimited plans and current pricing. *** ## Concurrency and Job Limits **Concurrency (threads)** determines how many extractions your plan can run in parallel. When all threads are busy, additional work waits in a queue until a thread becomes available. Requests are not rejected because all threads are in use. **Job-size limits** define the maximum number of URLs or pages that can be processed in a single Batch or Crawl job. These limits apply per job, not per month, so you can run as many jobs as needed. | Plan | Concurrency | Max URLs per batch job | Max pages per crawl job | | ---------------------- | ----------: | ---------------------: | ----------------------: | | Starter (2 threads) | 2 | 50 | 50 | | Growth (10 threads) | 10 | 500 | 500 | | Scale (25 threads) | 25 | 1,000 | 1,000 | | Pro (50 threads) | 50 | 2,000 | 2,000 | | Business (100 threads) | 100 | 4,000 | 4,000 | To process more URLs than a single job allows, split the work into multiple jobs. *** ## Pay as You Go Pay as You Go lets you buy request top-ups without committing to a subscription. Pay as You Go requests are prepaid and do not expire. Use the Geonode dashboard to see the available top-up sizes and current prices for your account. Pay as You Go requests: * Are prepaid. * Do not expire. * Can be used for bursty workloads. * Remain available even if you do not use the API every month. *** ## What Happens When You Run Out For request-based plans, if you run out of available requests, the API returns `402`. Add more requests from the dashboard, buy a Pay as You Go top-up, or upgrade to a subscription plan. ```json { "code": "PAYMENT_REQUIRED", "message": "Insufficient request balance.", "correlation_id": "req_...", "retryable": false } ``` The exact message may vary, but the response tells you that the request could not be processed because the account does not have enough available requests. *** ## Checking Usage You can monitor your usage in the Geonode dashboard. Depending on your plan, the dashboard shows: * Remaining free requests for the current billing month. * Current subscription allowance and rollover balance. * Pay as You Go request balance. * Thread usage and concurrency information for Unlimited plans. *** ## Billing Examples If you extract one static page as Markdown, it uses one request. ```json { "url": "https://example.com", "formats": ["markdown"], "render_js": false } ``` If you extract the same page as Markdown and HTML, it still uses one request. ```json { "url": "https://example.com", "formats": ["markdown", "html"], "render_js": false } ``` If you render a JavaScript-heavy page before extracting it, it still uses one request. ```json { "url": "https://quotes.toscrape.com/js/", "formats": ["markdown"], "render_js": true } ``` If the extraction fails, the failed extraction is not charged. For Batch and Crawl jobs, check `token_summary.tokens_charged_total` on the job status response to see how many requests have been charged so far. `token_summary.tokens_reserved` shows the amount still reserved for queued or processing items. # Request-Based Pricing (/docs/scraper-api/additional-resources/request_based_pricing) The Geonode Scraper API supports two pricing models: * Request-based pricing, where successful page extractions consume requests. * Unlimited plans, where throughput is based on concurrent threads instead of a request balance. This guide explains the request-based model, including the free tier, subscription request allowances, and Pay As You Go usage. If your workload requires unlimited requests and predictable monthly pricing, see [Unlimited Scraper API Pricing](/docs/scraper-api/additional-resources/unlimited_scraper_api_pricing). ## Request-Based Plans Request-based plans charge based on the number of successful page extractions. | Plan | Requests | Price | Cost per 1K requests | Best for | | ---- | ------------: | ------------: | -------------------: | ------------------------------------------ | | Free | 1,500 / month | $0 | — | Trying the API without a paid subscription | | 10K | 10,000 | $3.50 / month | $0.35 | Smaller recurring workloads | | 50K | 50,000 | $13 / month | $0.26 | Regular scraping workloads | | 250K | 250,000 | $43 / month | $0.17 | Larger scraping workloads | | 1M | 1,000,000 | $126 / month | $0.13 | High-volume request-based workloads | | 3M+ | 3,000,000+ | Custom | Custom | Production-scale workloads | The Free plan renews every billing month and does not require a credit card. For 3M+ plans, [contact sales](https://geonode.com/contact). > Pricing and plan availability can change. Check the [Geonode Scraper API pricing page](https://geonode.com/products/scraper-api) for the latest pricing. ## How requests are counted The request model is based on successful page extraction. | Operation | Request usage | | ---------------------- | ------------------------------------------------- | | Single-page extraction | 1 request per successfully extracted page | | Async extraction | 1 request when the page is successfully extracted | | Batch extraction | 1 request for each successfully extracted URL | | Crawl | 1 request for each successfully extracted page | | Map | Does not extract page content | | Job lookup | Does not extract page content | | Statistics | Does not extract page content | | Webhooks | Does not extract page content | | Health checks | Does not extract page content | ## What does not cost additional requests? The following options do not create additional request charges for a successful extraction: * JavaScript Rendering * Geo-targeting * Markdown output * HTML output * Requesting multiple supported output formats For example, extracting a page as both Markdown and HTML still uses one request. ## Failed requests Failed extraction requests are not charged. Requests rejected before extraction are also not charged, including: * Authentication errors * Validation errors * Requests rejected before extraction starts If an extraction does not successfully produce a page result, it does not consume a request from your request balance. ## Free Tier The free tier includes **1,500 requests every month**. You can start using the Scraper API without a credit card. Free requests: * Renew every billing month. * Do not roll over to the next month. * Expire at the end of the billing period. * Stop being available when the monthly allowance is exhausted. Once your free allowance is used, you can upgrade to a paid request-based plan or choose an Unlimited plan. ## Pay As You Go Pay As You Go is part of the request-based pricing model. It is useful when you need additional request capacity without moving to an Unlimited concurrency plan. Pay As You Go requests are prepaid and can be used for additional scraping workloads. The available top-up sizes and prices can vary. Use the Geonode dashboard to view the options currently available for your account. Pay As You Go requests: * Are prepaid. * Do not expire. * Can be used for bursty workloads. * Can be used alongside applicable request-based balances. ## Subscription requests Paid request-based subscriptions provide a monthly request allowance. Unused subscription requests roll over to the next month while your subscription remains active, according to the applicable subscription terms. Subscription requests: * Renew every billing month. * Can roll over when unused. * Remain available while the subscription is active. * Can be combined with applicable Pay As You Go balances. Use the Geonode dashboard to view the current subscription options and available request balances. ## What happens when you run out? If you have no available request balance, the API returns a `402` response. ```json { "code": "PAYMENT_REQUIRED", "message": "Insufficient request balance.", "correlation_id": "req_...", "retryable": false } ``` # Unlimited Scraper API Pricing (/docs/scraper-api/additional-resources/unlimited_scraper_api_pricing) Geonode's Unlimited Scraper API lets you run unlimited extraction requests on a flat monthly plan. Instead of using a monthly request balance, Unlimited plans are based on **concurrent threads**. Your thread count determines how many extractions can run at the same time. ## Unlimited Plans Choose an Unlimited plan based on the amount of work you need to process in parallel. | Plan | Concurrent threads | Price | Requests | Best for | | ---------- | -----------------: | -------------: | ---------------------- | ----------------------------------- | | Free Trial | 5 | $0 | Unlimited for 48 hours | Trying Unlimited before a paid plan | | Starter | 2 | $47 / month | Unlimited | Smaller production workloads | | Growth | 10 | $199 / month | Unlimited | Regular production workloads | | Scale | 25 | $399 / month | Unlimited | Large scraping workloads | | Pro | 50 | $699 / month | Unlimited | High-volume scraping workloads | | Business | 100 | $1,799 / month | Unlimited | Very high-volume scraping workloads | For Pro and Business plans, [contact sales](https://geonode.com/contact). > Pricing and plan availability can change. Check the [Geonode Scraper API pricing page](https://geonode.com/products/scraper-api) for the latest pricing and available plans. ## How Unlimited Pricing Works Unlimited plans do not use a monthly request allowance. Instead, each plan provides a specific number of **concurrent threads**. This determines how many extractions can run at the same time. For example, a plan with 10 concurrent threads can process up to 10 extraction tasks concurrently. If all available threads are busy, additional work waits in the queue until a thread becomes available. This means that a higher thread count increases your potential throughput when your workload can be processed in parallel. ## Understanding Concurrency Concurrency is the number of extraction tasks that can run at the same time. For example: | Concurrent threads | Maximum parallel extractions | | -----------------: | ---------------------------: | | 2 | 2 | | 10 | 10 | | 25 | 25 | | 50 | 50 | | 100 | 100 | When all threads are occupied, new work is queued rather than rejected because of concurrency. ## Job Limits Unlimited monthly requests do not mean that an individual Batch or Crawl job can contain an unlimited number of URLs or pages. Job-size limits apply separately to each Batch and Crawl job. | Plan | Concurrency | Max URLs per batch job | Max pages per crawl job | | -------- | ----------: | ---------------------: | ----------------------: | | Starter | 2 | 50 | 50 | | Growth | 10 | 500 | 500 | | Scale | 25 | 1,000 | 1,000 | | Pro | 50 | 2,000 | 2,000 | | Business | 100 | 4,000 | 4,000 | If your workload is larger than the maximum size of a single job, split it across multiple jobs. ## Choosing the Right Plan Choose your plan based on how much work you need to process **at the same time**. | Workload | Recommended plan | | ---------------------------- | ---------------- | | Testing and evaluation | Free Trial | | Small production workloads | Starter | | Regular production workloads | Growth | | Large scraping workloads | Scale | | High-volume scraping | Pro | | Very high-volume scraping | Business | If you are unsure which plan you need, start with the free trial and measure your workload before moving to a paid plan. ## What Is Included? Unlimited plans provide access to the Scraper API features used for web extraction. You can use features such as: * HTML and Markdown output * JavaScript rendering * Proxy selection * Geo-targeting * Batch extraction * Crawl jobs These features do not create additional request charges on an Unlimited plan. ## Unlimited Requests Unlimited means that your plan does not have a monthly request allowance. You can continue submitting extraction requests throughout the billing period without tracking a remaining monthly request balance. Your plan still determines how many extractions can run concurrently. For example, with 10 concurrent threads: 1. Up to 10 extractions can run at the same time. 2. Additional work waits when all 10 threads are occupied. 3. A waiting extraction starts when a thread becomes available. ## Unlimited vs Request-Based Pricing Unlimited and request-based pricing use different billing models. | | Unlimited | Request-Based | | ------------------------- | --------------------- | -------------------------------- | | Pricing model | Flat monthly price | Based on request volume | | Monthly request allowance | Unlimited | Plan-dependent | | Main pricing factor | Concurrent threads | Successful extractions | | Best for | High-volume workloads | Variable workloads | | Monthly cost | Predictable | Based on selected plan and usage | If you want to understand how request-based pricing works, including request counting, free requests, subscriptions, and Pay As You Go, see the [Request-Based Pricing](/docs/scraper-api/additional-resources/request_based_pricing) guide. ## Related Guides * [Request-Based Pricing](/docs/scraper-api/additional-resources/request_based_pricing) * [Extraction](/docs/scraper-api/guides/extraction/01_understanding_extraction) * [Batch](/docs/scraper-api/guides/batch/00_understanding_batch) * [Dashboard Overview](/docs/scraper-api/dashboard-guides/overview) ## Get Started Ready to use the Unlimited Scraper API? Open the [Geonode dashboard](https://app.geonode.com/login) to choose a plan. # Dashboard Overview (/docs/scraper-api/dashboard-guides/overview) The GeoNode Dashboard provides a graphical interface for using the **Scraper** and **Map** APIs without writing code. From the dashboard, you can submit extraction and mapping requests, monitor usage, review previous jobs, and manage your account from a single place. After signing in, you'll see the main dashboard with navigation on the left and your selected tool in the main workspace. GeoNode Dashboard Overview ## Available Requests The **Available Requests** card provides a quick overview of your current usage, including: * Remaining requests available in your account. * Your free request allowance. * When your request quota renews. * The number of requests used during the last 24 hours. This allows you to monitor your usage without leaving the dashboard. ## Upgrade Your Plan The **Upgrade Your Plan** section lets you review available pricing plans and upgrade your subscription whenever you need additional requests or higher usage limits. ## Dashboard Navigation Use the navigation menu on the left to switch between the available tools. * **Extraction** - Extract structured content from one or more web pages. * **Mapping** - Discover URLs from websites before extracting their content. * **Search** - Search the web for specific content. * **Crawl** - Crawl a website and extract its content. ## Settings Both the Scraper and Map dashboards include a **Settings** panel where you can configure your requests before starting a job. Detailed explanations of each setting are available in the corresponding API Guides and are not repeated in the Dashboard documentation. ## Recent Jobs The lower section of each dashboard displays your recent activity. Depending on the selected tool, you can: * View recent extraction or mapping jobs. * Check the current job status. * Search previous jobs. * Open completed results. * Review execution times. Dedicated guides explain how to work with jobs in more detail. ## Statistics The **Statistics** tab provides an overview of your API usage and request activity, helping you monitor your overall usage over time. ## Next Steps Now that you're familiar with the dashboard layout, continue with one of the following guides: * **Scraper Overview** – Learn how to extract content from web pages using the Dashboard. * **Map Overview** – Learn how to discover URLs from websites using the Dashboard. * **Crawl Overview** – Learn how to crawl a website and download results using the Dashboard. * **Search Overview** – Learn how to search the web using the Dashboard. # CrewAI (/docs/scraper-api/developer-guides/crewai) {/* DRAFT: do not publish / do not add to meta.json until approved */} `geonode-scraper-crewai` builds CrewAI tools you can call with `tool.run(...)` or attach to an `Agent`. Each tool returns a JSON-friendly dict. Responses below are from live runs against `https://scraper.geonode.io`. ## Setup CrewAI requires Python 3.10 through 3.13. ```bash pip install geonode-scraper-crewai python-dotenv ``` Create a `.env` file: ```bash GEONODE_SCRAPER_API_KEY=your_geonode_key SCRAPER_API_BASE_URL=https://scraper.geonode.io ``` ```python import os from dotenv import load_dotenv from geonode_scraper_crewai import build_crewai_tools from geonode_scraper_tools_core import ScraperToolSettings load_dotenv() settings = ScraperToolSettings( host=os.environ.get("SCRAPER_API_BASE_URL", "https://scraper.geonode.io"), api_key=os.environ["GEONODE_SCRAPER_API_KEY"], ) tools = build_crewai_tools(settings=settings) by_name = {tool.name: tool for tool in tools} ``` Use `https://scraper.geonode.io` for production. ### Response shape Every tool returns a dict like: ```python { "ok": True, "operation": "extract", "attempts": 1, "result": { ... }, } ``` Read the payload from `response["result"]`. ## Tools overview | Group | Tool names | | -------------- | ------------------------------------------------------------------------------------------------------- | | **Extraction** | `scraper_extract_content`, `scraper_get_job_result`, `scraper_wait_for_job`, `scraper_list_jobs` | | **Batch** | `scraper_create_batch`, `scraper_get_batch_status`, `scraper_wait_for_batch`, `scraper_list_batch_jobs` | | **Crawl** | `scraper_create_crawl`, `scraper_get_crawl_status`, `scraper_wait_for_crawl`, `scraper_list_crawl_jobs` | | **Map** | `scraper_map_urls`, `scraper_list_map_jobs`, `scraper_get_map_job` | | **Search** | `scraper_search`, `scraper_list_search_jobs`, `scraper_get_search_job` | | **Account** | `scraper_get_statistics`, `scraper_get_concurrency_usage`, `scraper_check_health` | *** ## Extraction ### scraper\_extract\_content (sync) ```python response = by_name["scraper_extract_content"].run( url="https://docs.geonode.com/docs/scraper-api/quick-start", formats=["markdown"], processing_mode="sync", ) result = response["result"] markdown = (result.get("data") or {}).get("markdown") or "" print("ok:", response["ok"]) print("tokens:", result.get("tokens_charged")) print("markdown_len:", len(markdown)) print("preview:", markdown[:200]) ``` Response (live run): ```text ok: True tokens: 1 markdown_len: 15871 preview: --- canonical: https://docs.geonode.com/docs/scraper-api/quick-start meta-description: Get your API key, authenticate requests, and choose the right API for your use case. ... ``` ### scraper\_extract\_content (async) ```python response = by_name["scraper_extract_content"].run( url="https://docs.geonode.com/docs/scraper-api/quick-start", formats=["markdown"], processing_mode="async", ) result = response["result"] print("job_id:", result["job_id"]) print("status:", result["status"]) ``` Response (live run): ```text job_id: e35c4d0e-9df3-43b9-9446-33c517b1cfc4 status: queued ``` ### scraper\_wait\_for\_job ```python response = by_name["scraper_wait_for_job"].run( job_id="e35c4d0e-9df3-43b9-9446-33c517b1cfc4", timeout_seconds=120, ) result = response["result"] markdown = (result.get("data") or {}).get("markdown") or "" print("status:", result["status"]) print("markdown_len:", len(markdown)) print("poll_attempts:", response.get("poll_attempts")) ``` Response (live run): ```text status: completed markdown_len: 15871 poll_attempts: 4 ``` ### scraper\_get\_job\_result ```python response = by_name["scraper_get_job_result"].run( job_id="e35c4d0e-9df3-43b9-9446-33c517b1cfc4", ) result = response["result"] print("status:", result["status"]) print("tokens:", result.get("tokens_charged")) ``` Response (live run): ```text status: completed tokens: 1 ``` ### scraper\_list\_jobs ```python response = by_name["scraper_list_jobs"].run(page=1, page_size=3) result = response["result"] print("page:", result["page"], "page_size:", result["page_size"]) for job in result.get("jobs") or []: print(job["job_id"], job["status"], job.get("url")) ``` Response (live run): ```text page: 1 page_size: 3 752b8599-5915-441c-bc1f-b9fb0d938f72 completed ... 0bcd424e-b5ab-4b85-9cbd-5b44f7ed433c completed http://example.com/ 0dbb126c-4331-4b4b-9d55-0c380ba69ae7 completed http://example.com/ ``` *** ## Batch ### scraper\_create\_batch ```python response = by_name["scraper_create_batch"].run( urls=[ "https://docs.geonode.com/docs/scraper-api/quick-start", "https://docs.geonode.com/docs/scraper-api", ], formats=["markdown"], ) result = response["result"] print("job_id:", result["job_id"]) print("accepted_urls:", result["accepted_urls"]) print("status:", result["status"]) ``` Response (live run): ```text job_id: 037be6d5-06e8-480f-9e22-4eac6cfb2966 accepted_urls: 2 status: queued ``` ### scraper\_get\_batch\_status ```python response = by_name["scraper_get_batch_status"].run( job_id="037be6d5-06e8-480f-9e22-4eac6cfb2966", page=1, page_size=10, ) result = response["result"] print(result["status"], result["completed_urls"], "/", result["total_urls"]) ``` Response (live run, mid-job): ```text processing 0 / 2 ``` ### scraper\_wait\_for\_batch ```python response = by_name["scraper_wait_for_batch"].run( job_id="037be6d5-06e8-480f-9e22-4eac6cfb2966", timeout_seconds=120, ) result = response["result"] print(result["status"], result["completed_urls"], "/", result["total_urls"]) print("poll_attempts:", response.get("poll_attempts")) ``` Response (live run): ```text completed 2 / 2 poll_attempts: 3 ``` ### scraper\_list\_batch\_jobs ```python response = by_name["scraper_list_batch_jobs"].run(page=1, page_size=3) for job in (response["result"].get("jobs") or []): print( job["job_id"], job["status"], job["completed_urls"], "/", job["accepted_urls"], ) ``` Response (live run): ```text 037be6d5-06e8-480f-9e22-4eac6cfb2966 completed 2 / 2 600bb35a-d6fc-4e05-be49-816b8f3ad5d7 completed 2 / 2 55b11790-5c7f-4945-bc13-bd9a365a1835 completed 2 / 2 ``` *** ## Crawl ### scraper\_create\_crawl ```python response = by_name["scraper_create_crawl"].run( url="https://docs.geonode.com/docs/scraper-api", depth=2, limit=3, formats=["markdown"], same_domain_only=True, ) result = response["result"] print("job_id:", result["job_id"]) print("estimated_pages:", result["estimated_pages"]) print("status:", result["status"]) ``` Response (live run): ```text job_id: dbc0fa1a-1fea-496c-abe4-a33c2b91efb7 estimated_pages: 3 status: queued ``` ### scraper\_get\_crawl\_status ```python response = by_name["scraper_get_crawl_status"].run( job_id="dbc0fa1a-1fea-496c-abe4-a33c2b91efb7", page=1, page_size=10, ) result = response["result"] print( result["status"], result.get("completed_pages"), "/", result.get("total_pages"), ) ``` Response (live run, early poll): ```text processing 0 / 1 ``` ### scraper\_wait\_for\_crawl ```python response = by_name["scraper_wait_for_crawl"].run( job_id="dbc0fa1a-1fea-496c-abe4-a33c2b91efb7", timeout_seconds=180, ) result = response["result"] print(result["status"], result["completed_pages"], "/", result["total_pages"]) print("poll_attempts:", response.get("poll_attempts")) ``` Response (live run): ```text completed 3 / 3 poll_attempts: 4 ``` ### scraper\_list\_crawl\_jobs ```python response = by_name["scraper_list_crawl_jobs"].run(page=1, page_size=2) for job in (response["result"].get("jobs") or []): print( job["job_id"], job["status"], job["completed_pages"], "/", job["total_pages"], ) ``` Response (live run): ```text dbc0fa1a-1fea-496c-abe4-a33c2b91efb7 completed 3 / 3 ab83f113-5086-46d3-a045-3a42263b5c87 completed 3 / 3 ``` *** ## Map ### scraper\_map\_urls ```python response = by_name["scraper_map_urls"].run( url="https://docs.geonode.com/docs/scraper-api", ) result = response["result"] links = result.get("links") or [] print("link_count:", result.get("links_count") or len(links)) for link in links[:5]: print(link.get("source"), link.get("url")) ``` Response (live run): ```text link_count: 112 sitemap https://docs.geonode.com/docs/scraper-api sitemap https://docs.geonode.com/docs/scraper-api/quick-start sitemap https://docs.geonode.com/docs/scraper-api/additional-resources/choosing_scraper_api_plan sitemap https://docs.geonode.com/docs/scraper-api/additional-resources/faq sitemap https://docs.geonode.com/docs/scraper-api/additional-resources/pricing-and-requests ``` ### scraper\_list\_map\_jobs ```python response = by_name["scraper_list_map_jobs"].run(page=1, page_size=2) for job in (response["result"].get("jobs") or []): print(job["job_id"], job["status"], job.get("links_count"), job.get("url")) ``` Response (live run): ```text 88b13fab-8d1b-421d-aa0e-256dd705b3aa completed 112 https://docs.geonode.com/docs/scraper-api 3abfc672-2aff-4daf-9f68-f033978ddfbd completed 112 https://docs.geonode.com/docs/scraper-api ``` ### scraper\_get\_map\_job ```python response = by_name["scraper_get_map_job"].run( job_id="88b13fab-8d1b-421d-aa0e-256dd705b3aa", ) result = response["result"] print("status:", result["status"]) print("link_count:", result.get("links_count")) for link in (result.get("links") or [])[:3]: print(link.get("source"), link.get("url")) ``` Response (live run): ```text status: completed link_count: 112 sitemap https://docs.geonode.com/docs/scraper-api sitemap https://docs.geonode.com/docs/scraper-api/quick-start sitemap https://docs.geonode.com/docs/scraper-api/additional-resources/choosing_scraper_api_plan ``` *** ## Search ### scraper\_search ```python response = by_name["scraper_search"].run(query="geonode scraper api") result = response["result"] print("job_id:", result["job_id"]) print("hit_count:", result.get("results_count") or len(result.get("results") or [])) for hit in (result.get("results") or [])[:5]: print(hit["position"], hit["title"], hit["url"]) ``` Response (live run): ```text job_id: 7eeb9599-c39a-4eba-8e0d-2eaed5a8deab hit_count: 15 1 Wholesale proxies & web data infrastructure | Geonode https://geonode.com/ 2 Geonode - PyPI https://pypi.org/user/Geonode/ 3 pavel.s - PyPI https://pypi.org/user/pavel.s/ 4 Geonode Documentation | Geonode https://docs.geonode.com/ 5 GeoNode https://geonode.org/ ``` ### scraper\_list\_search\_jobs ```python response = by_name["scraper_list_search_jobs"].run(page=1, page_size=2) for job in (response["result"].get("jobs") or []): print(job["job_id"], job.get("query"), job["status"], job.get("results_count")) ``` Response (live run): ```text 7eeb9599-c39a-4eba-8e0d-2eaed5a8deab geonode scraper api completed 15 f80269e1-1911-4fd4-b1a1-fc0f7bbae7f9 geonode scraper api completed 15 ``` ### scraper\_get\_search\_job ```python response = by_name["scraper_get_search_job"].run( job_id="7eeb9599-c39a-4eba-8e0d-2eaed5a8deab", ) result = response["result"] print("status:", result["status"]) print("hit_count:", result.get("results_count")) for hit in (result.get("results") or [])[:3]: print(hit["position"], hit["title"], hit["url"]) ``` Response (live run): ```text status: completed hit_count: 15 1 Wholesale proxies & web data infrastructure | Geonode https://geonode.com/ 2 Geonode - PyPI https://pypi.org/user/Geonode/ 3 pavel.s - PyPI https://pypi.org/user/pavel.s/ ``` *** ## Statistics, usage, and health ### scraper\_get\_statistics ```python response = by_name["scraper_get_statistics"].run() result = response["result"] print("extraction_count:", result.get("extraction_count")) print("success_rate:", result.get("success_rate")) ``` Response (live run): ```text extraction_count: 4401 success_rate: 0.95137 ``` ### scraper\_get\_concurrency\_usage ```python response = by_name["scraper_get_concurrency_usage"].run() result = response["result"] print( result["work_concurrency_in_use"], "/", result["work_concurrency_limit"], ) ``` Response (live run): ```text 0 / 50 ``` ### scraper\_check\_health ```python response = by_name["scraper_check_health"].run() result = response["result"] print(result.get("service"), result.get("status"), result.get("version")) ``` Response (live run): ```text Scraper API ok 0.1.0 ``` *** ## Selecting a subset of tools ```python tools = build_crewai_tools( settings=settings, operations=["extract", "map_urls", "create_crawl", "wait_for_crawl"], ) print([tool.name for tool in tools]) ``` Response (live run): ```text ['scraper_extract_content', 'scraper_map_urls', 'scraper_create_crawl', 'scraper_wait_for_crawl'] ``` *** ## Use tools with an Agent Attach the tools to a CrewAI agent: ```python from crewai import Agent from geonode_scraper_crewai import build_crewai_tools from geonode_scraper_tools_core import ScraperToolSettings settings = ScraperToolSettings( host="https://scraper.geonode.io", api_key=os.environ["GEONODE_SCRAPER_API_KEY"], ) agent = Agent( role="Web Researcher", goal="Extract and inspect web content.", backstory="Focused on pulling structured data from URLs.", tools=build_crewai_tools(settings=settings), ) ``` The agent can call the same tools shown above. Direct checks still use `tool.run(...)`. *** ## Toolkit helper ```python from geonode_scraper_crewai import ScraperCrewAIToolkit from geonode_scraper_tools_core import ScraperToolSettings settings = ScraperToolSettings( host="https://scraper.geonode.io", api_key=os.environ["GEONODE_SCRAPER_API_KEY"], ) toolkit = ScraperCrewAIToolkit.from_settings(settings) tools = toolkit.get_tools() ``` For package versions and changelog, see [geonode-scraper-crewai on PyPI](https://pypi.org/project/geonode-scraper-crewai/). # LangChain (/docs/scraper-api/developer-guides/langchain) {/* DRAFT: do not publish / do not add to meta.json until approved */} `geonode-scraper-langchain` builds LangChain `StructuredTool` objects you can call with `tool.invoke(...)` or pass into an agent. Each tool returns a JSON-friendly dict. Responses below are from live runs against `https://scraper.geonode.io`. ## Setup ```bash pip install geonode-scraper-langchain python-dotenv ``` Create a `.env` file: ```bash GEONODE_SCRAPER_API_KEY=your_geonode_key SCRAPER_API_BASE_URL=https://scraper.geonode.io ``` ```python import os from dotenv import load_dotenv from geonode_scraper_langchain import build_langchain_tools from geonode_scraper_tools_core import ScraperToolSettings load_dotenv() settings = ScraperToolSettings( host=os.environ.get("SCRAPER_API_BASE_URL", "https://scraper.geonode.io"), api_key=os.environ["GEONODE_SCRAPER_API_KEY"], ) tools = build_langchain_tools(settings=settings) by_name = {tool.name: tool for tool in tools} ``` Use `https://scraper.geonode.io` for production. ### Response shape Every tool returns a dict like: ```python { "ok": True, "operation": "extract", "attempts": 1, "result": { ... }, } ``` Read the payload from `response["result"]`. ## Tools overview | Group | Tool names | | -------------- | ------------------------------------------------------------------------------------------------------- | | **Extraction** | `scraper_extract_content`, `scraper_get_job_result`, `scraper_wait_for_job`, `scraper_list_jobs` | | **Batch** | `scraper_create_batch`, `scraper_get_batch_status`, `scraper_wait_for_batch`, `scraper_list_batch_jobs` | | **Crawl** | `scraper_create_crawl`, `scraper_get_crawl_status`, `scraper_wait_for_crawl`, `scraper_list_crawl_jobs` | | **Map** | `scraper_map_urls`, `scraper_list_map_jobs`, `scraper_get_map_job` | | **Search** | `scraper_search`, `scraper_list_search_jobs`, `scraper_get_search_job` | | **Account** | `scraper_get_statistics`, `scraper_get_concurrency_usage`, `scraper_check_health` | *** ## Extraction ### scraper\_extract\_content (sync) ```python response = by_name["scraper_extract_content"].invoke( { "url": "https://docs.geonode.com/docs/scraper-api/quick-start", "formats": ["markdown"], "processing_mode": "sync", } ) result = response["result"] markdown = (result.get("data") or {}).get("markdown") or "" print("ok:", response["ok"]) print("tokens:", result.get("tokens_charged")) print("markdown_len:", len(markdown)) print("preview:", markdown[:200]) ``` Response (live run): ```text ok: True tokens: 1 markdown_len: 15871 preview: --- canonical: https://docs.geonode.com/docs/scraper-api/quick-start meta-description: Get your API key, authenticate requests, and choose the right API for your use case. ... ``` ### scraper\_extract\_content (async) ```python response = by_name["scraper_extract_content"].invoke( { "url": "https://docs.geonode.com/docs/scraper-api/quick-start", "formats": ["markdown"], "processing_mode": "async", } ) result = response["result"] print("job_id:", result["job_id"]) print("status:", result["status"]) ``` Response (live run): ```text job_id: 2ac66637-905c-4892-a6b9-5ace58b2ffe9 status: queued ``` ### scraper\_wait\_for\_job ```python response = by_name["scraper_wait_for_job"].invoke( { "job_id": "2ac66637-905c-4892-a6b9-5ace58b2ffe9", "timeout_seconds": 120, } ) result = response["result"] markdown = (result.get("data") or {}).get("markdown") or "" print("status:", result["status"]) print("markdown_len:", len(markdown)) print("poll_attempts:", response.get("poll_attempts")) ``` Response (live run): ```text status: completed markdown_len: 15871 poll_attempts: 3 ``` ### scraper\_get\_job\_result ```python response = by_name["scraper_get_job_result"].invoke( {"job_id": "2ac66637-905c-4892-a6b9-5ace58b2ffe9"} ) result = response["result"] print("status:", result["status"]) print("tokens:", result.get("tokens_charged")) ``` Response (live run): ```text status: completed tokens: 1 ``` ### scraper\_list\_jobs ```python response = by_name["scraper_list_jobs"].invoke({"page": 1, "page_size": 3}) result = response["result"] print("page:", result["page"], "page_size:", result["page_size"]) for job in result.get("jobs") or []: print(job["job_id"], job["status"], job.get("url")) ``` Response (live run): ```text page: 1 page_size: 3 752b8599-5915-441c-bc1f-b9fb0d938f72 completed ... 0bcd424e-b5ab-4b85-9cbd-5b44f7ed433c completed http://example.com/ 0dbb126c-4331-4b4b-9d55-0c380ba69ae7 completed http://example.com/ ``` *** ## Batch ### scraper\_create\_batch ```python response = by_name["scraper_create_batch"].invoke( { "urls": [ "https://docs.geonode.com/docs/scraper-api/quick-start", "https://docs.geonode.com/docs/scraper-api", ], "formats": ["markdown"], } ) result = response["result"] print("job_id:", result["job_id"]) print("accepted_urls:", result["accepted_urls"]) print("status:", result["status"]) ``` Response (live run): ```text job_id: 600bb35a-d6fc-4e05-be49-816b8f3ad5d7 accepted_urls: 2 status: queued ``` ### scraper\_get\_batch\_status ```python response = by_name["scraper_get_batch_status"].invoke( { "job_id": "600bb35a-d6fc-4e05-be49-816b8f3ad5d7", "page": 1, "page_size": 10, } ) result = response["result"] print(result["status"], result["completed_urls"], "/", result["total_urls"]) ``` Response (live run, mid-job): ```text processing 0 / 2 ``` ### scraper\_wait\_for\_batch ```python response = by_name["scraper_wait_for_batch"].invoke( { "job_id": "600bb35a-d6fc-4e05-be49-816b8f3ad5d7", "timeout_seconds": 120, } ) result = response["result"] print(result["status"], result["completed_urls"], "/", result["total_urls"]) print("poll_attempts:", response.get("poll_attempts")) ``` Response (live run): ```text completed 2 / 2 poll_attempts: 3 ``` ### scraper\_list\_batch\_jobs ```python response = by_name["scraper_list_batch_jobs"].invoke({"page": 1, "page_size": 3}) for job in (response["result"].get("jobs") or []): print( job["job_id"], job["status"], job["completed_urls"], "/", job["accepted_urls"], ) ``` Response (live run): ```text 600bb35a-d6fc-4e05-be49-816b8f3ad5d7 completed 2 / 2 55b11790-5c7f-4945-bc13-bd9a365a1835 completed 2 / 2 d8e92d3a-c939-4ec6-b133-3e04c6de71d1 completed 2 / 2 ``` *** ## Crawl ### scraper\_create\_crawl ```python response = by_name["scraper_create_crawl"].invoke( { "url": "https://docs.geonode.com/docs/scraper-api", "depth": 2, "limit": 3, "formats": ["markdown"], "same_domain_only": True, } ) result = response["result"] print("job_id:", result["job_id"]) print("estimated_pages:", result["estimated_pages"]) print("status:", result["status"]) ``` Response (live run): ```text job_id: ab83f113-5086-46d3-a045-3a42263b5c87 estimated_pages: 3 status: queued ``` ### scraper\_get\_crawl\_status ```python response = by_name["scraper_get_crawl_status"].invoke( { "job_id": "ab83f113-5086-46d3-a045-3a42263b5c87", "page": 1, "page_size": 10, } ) result = response["result"] print( result["status"], result.get("completed_pages"), "/", result.get("total_pages"), ) ``` Response (live run, early poll): ```text processing 0 / 1 ``` ### scraper\_wait\_for\_crawl ```python response = by_name["scraper_wait_for_crawl"].invoke( { "job_id": "ab83f113-5086-46d3-a045-3a42263b5c87", "timeout_seconds": 180, } ) result = response["result"] print(result["status"], result["completed_pages"], "/", result["total_pages"]) ``` Response (live run): ```text completed 3 / 3 ``` ### scraper\_list\_crawl\_jobs ```python response = by_name["scraper_list_crawl_jobs"].invoke({"page": 1, "page_size": 2}) for job in (response["result"].get("jobs") or []): print( job["job_id"], job["status"], job["completed_pages"], "/", job["total_pages"], ) ``` Response (live run): ```text ab83f113-5086-46d3-a045-3a42263b5c87 completed 3 / 3 9f447ced-4aa3-45a3-b9af-421f751b8cec completed 3 / 3 ``` *** ## Map ### scraper\_map\_urls ```python response = by_name["scraper_map_urls"].invoke( {"url": "https://docs.geonode.com/docs/scraper-api"} ) result = response["result"] links = result.get("links") or [] print("link_count:", result.get("links_count") or len(links)) for link in links[:5]: print(link.get("source"), link.get("url")) ``` Response (live run): ```text link_count: 112 sitemap https://docs.geonode.com/docs/scraper-api sitemap https://docs.geonode.com/docs/scraper-api/quick-start sitemap https://docs.geonode.com/docs/scraper-api/additional-resources/choosing_scraper_api_plan sitemap https://docs.geonode.com/docs/scraper-api/additional-resources/faq sitemap https://docs.geonode.com/docs/scraper-api/additional-resources/pricing-and-requests ``` ### scraper\_list\_map\_jobs ```python response = by_name["scraper_list_map_jobs"].invoke({"page": 1, "page_size": 2}) for job in (response["result"].get("jobs") or []): print(job["job_id"], job["status"], job.get("links_count"), job.get("url")) ``` Response (live run): ```text 3abfc672-2aff-4daf-9f68-f033978ddfbd completed 112 https://docs.geonode.com/docs/scraper-api c7cb6e40-ddfa-416a-b255-00dfdfd717bc completed 112 https://docs.geonode.com/docs/scraper-api ``` ### scraper\_get\_map\_job ```python response = by_name["scraper_get_map_job"].invoke( {"job_id": "3abfc672-2aff-4daf-9f68-f033978ddfbd"} ) result = response["result"] print("status:", result["status"]) print("link_count:", result.get("links_count")) for link in (result.get("links") or [])[:3]: print(link.get("source"), link.get("url")) ``` Response (live run): ```text status: completed link_count: 112 sitemap https://docs.geonode.com/docs/scraper-api sitemap https://docs.geonode.com/docs/scraper-api/quick-start sitemap https://docs.geonode.com/docs/scraper-api/additional-resources/choosing_scraper_api_plan ``` *** ## Search ### scraper\_search ```python response = by_name["scraper_search"].invoke({"query": "geonode scraper api"}) result = response["result"] print("job_id:", result["job_id"]) print("hit_count:", result.get("results_count") or len(result.get("results") or [])) for hit in (result.get("results") or [])[:5]: print(hit["position"], hit["title"], hit["url"]) ``` Response (live run): ```text job_id: f80269e1-1911-4fd4-b1a1-fc0f7bbae7f9 hit_count: 15 1 Wholesale proxies & web data infrastructure | Geonode https://geonode.com/ 2 Geonode - PyPI https://pypi.org/user/Geonode/ 3 pavel.s - PyPI https://pypi.org/user/pavel.s/ 4 Geonode Documentation | Geonode https://docs.geonode.com/ 5 GeoNode https://geonode.org/ ``` ### scraper\_list\_search\_jobs ```python response = by_name["scraper_list_search_jobs"].invoke({"page": 1, "page_size": 2}) for job in (response["result"].get("jobs") or []): print(job["job_id"], job.get("query"), job["status"], job.get("results_count")) ``` Response (live run): ```text f80269e1-1911-4fd4-b1a1-fc0f7bbae7f9 geonode scraper api completed 15 29809a61-49b6-4bf6-925b-aa32dd91e441 geonode scraper api completed 15 ``` ### scraper\_get\_search\_job ```python response = by_name["scraper_get_search_job"].invoke( {"job_id": "f80269e1-1911-4fd4-b1a1-fc0f7bbae7f9"} ) result = response["result"] print("status:", result["status"]) print("hit_count:", result.get("results_count")) for hit in (result.get("results") or [])[:3]: print(hit["position"], hit["title"], hit["url"]) ``` Response (live run): ```text status: completed hit_count: 15 1 Wholesale proxies & web data infrastructure | Geonode https://geonode.com/ 2 Geonode - PyPI https://pypi.org/user/Geonode/ 3 pavel.s - PyPI https://pypi.org/user/pavel.s/ ``` *** ## Statistics, usage, and health ### scraper\_get\_statistics ```python response = by_name["scraper_get_statistics"].invoke({}) result = response["result"] print("extraction_count:", result.get("extraction_count")) print("success_rate:", result.get("success_rate")) ``` Response (live run): ```text extraction_count: 4392 success_rate: 0.95128 ``` ### scraper\_get\_concurrency\_usage ```python response = by_name["scraper_get_concurrency_usage"].invoke({}) result = response["result"] print( result["work_concurrency_in_use"], "/", result["work_concurrency_limit"], ) ``` Response (live run): ```text 0 / 50 ``` ### scraper\_check\_health ```python response = by_name["scraper_check_health"].invoke({}) result = response["result"] print(result.get("service"), result.get("status"), result.get("version")) ``` Response (live run): ```text Scraper API ok 0.1.0 ``` *** ## Selecting a subset of tools ```python tools = build_langchain_tools( settings=settings, operations=["extract", "map_urls", "create_batch", "wait_for_batch"], ) print([tool.name for tool in tools]) ``` Response (live run): ```text ['scraper_extract_content', 'scraper_map_urls', 'scraper_create_batch', 'scraper_wait_for_batch'] ``` *** ## Use tools with an agent Pass the tool list into your LangChain agent or bind them to a chat model. Example with `bind_tools`: ```python from langchain_openai import ChatOpenAI llm = ChatOpenAI(model="gpt-4o-mini") llm_with_tools = llm.bind_tools(tools) # The model can request tool calls; your agent loop then runs tool.invoke(...) ``` Or build an agent that owns the tools (exact API depends on your LangChain version): ```python # Pseudocode: wire `tools` into your preferred LangChain agent helper # agent = create_agent(model=llm, tools=tools) # agent.invoke({"messages": [("user", "Extract markdown from https://example.com")]}) ``` The tool calls themselves match the `tool.invoke({...})` examples above. *** ## Toolkit helper You can also use the toolkit class: ```python from geonode_scraper_langchain import ScraperLangChainToolkit from geonode_scraper_tools_core import ScraperToolSettings settings = ScraperToolSettings( host="https://scraper.geonode.io", api_key=os.environ["GEONODE_SCRAPER_API_KEY"], ) toolkit = ScraperLangChainToolkit.from_settings(settings) tools = toolkit.get_tools() ``` For package versions and changelog, see [geonode-scraper-langchain on PyPI](https://pypi.org/project/geonode-scraper-langchain/). # Python SDK (/docs/scraper-api/developer-guides/python-sdk) {/* DRAFT: do not publish / do not add to meta.json until approved */} The Geonode Scraper Python SDK wraps every public Scraper API endpoint. This guide: 1. **API overview:** what each API group does and when to use it 2. **Shared request options:** enums used across APIs (`formats`, `processing_mode`, proxy, wait) 3. **Each API section:** that API’s methods, request fields, then code + live response per method Responses below are from live runs against `https://scraper.geonode.io`. ## Setup ```bash pip install geonode-scraper-sdk python-dotenv ``` Create a `.env` file: ```bash GEONODE_SCRAPER_API_KEY=your_geonode_key SCRAPER_API_BASE_URL=https://scraper.geonode.io ``` ```python import os from dotenv import load_dotenv from geonode_scraper_sdk import Configuration, ApiClient load_dotenv() configuration = Configuration( host=os.environ.get("SCRAPER_API_BASE_URL", "https://scraper.geonode.io"), api_key={"ApiKeyAuth": os.environ["GEONODE_SCRAPER_API_KEY"]}, ) ``` Use `https://scraper.geonode.io` for production. If `host` is omitted, the client defaults to `http://localhost`. ## API overview | API | SDK class | Endpoint | When to use | | -------------- | --------------- | ---------------- | ---------------------------------------- | | **Extract** | `ExtractionApi` | `/v1/extract` | One URL → HTML/Markdown. Sync or async. | | **Batch** | `BatchApi` | `/v1/batch` | Many known URLs in one job. | | **Crawl** | `CrawlApi` | `/v1/crawl` | Seed URL + follow links (depth/limit). | | **Map** | `MapApi` | `/v1/map` | Discover URLs without scraping content. | | **Search** | `SearchApi` | `/v1/search` | Web search → ranked URLs. | | **Usage** | `UsageApi` | `/v1/usage` | Live concurrency vs plan limit. | | **Statistics** | `StatisticsApi` | `/v1/statistics` | Historical counts, tokens, success rate. | | **System** | `SystemApi` | `/health` | Service health check. | | **Webhooks** | `WebhooksApi` | `/v1/webhooks` | Callbacks when async jobs finish. | **Typical flows:** single page → Extract sync · known URL list → Batch · whole site → Map then Crawl/Batch · discovery → Search then Extract · production async → ASYNC + Webhooks. Each API section below lists that API’s methods and request fields, then walks every method with code and a live response. *** ## Shared request options Enums and helpers reused by Extract, Batch, and Crawl. Per-API field tables live in each API section. ### Output formats ```python from geonode_scraper_sdk import OutputFormat # Available values: OutputFormat.HTML # "html": raw page HTML OutputFormat.MARKDOWN # "markdown": cleaned Markdown ``` Pass as a list: `formats=[OutputFormat.MARKDOWN]` or `[OutputFormat.HTML, OutputFormat.MARKDOWN]`. Extract defaults to `[HTML]`; Batch/Crawl default to server-side defaults if omitted. ### Processing mode (Extract only) ```python from geonode_scraper_sdk import ProcessingMode ProcessingMode.SYNC # "sync": block until content is ready (default) ProcessingMode.ASYNC # "async": return job_id immediately; poll get_job_result ``` Batch, Crawl, Map, and Search are always async jobs (create → poll status / get job). ### Proxy settings ```python from geonode_scraper_sdk import ProxySettings, ProxyType ProxySettings( country="US", # ISO 3166-1 alpha-2 (optional) type=ProxyType.RESIDENTIAL, # residential | datacenter | mix ) ``` ### Wait config (JS rendering) Used with `render_js=True` to control headless browser timing: ```python from geonode_scraper_sdk import WaitConfig, WaitUntil WaitConfig( wait_until=WaitUntil.NETWORKIDLE, # commit | domcontentloaded | load | networkidle wait_for="#content", # CSS selector (optional) wait_timeout=10000, # ms, 0-30000 ) ``` ### Custom headers Pass a dict on Extract/Batch requests: `headers={"User-Agent": "my-bot/1.0"}`. ### Job status values Async jobs move through: `queued` → `processing` → `completed` | `failed` | `cancelled`. *** ## Extraction API `ExtractionApi` (`/v1/extract`) **Methods:** * `extract_v1_extract_post(extract_request)`: sync or async extract * `get_job_result_v1_extract_job_id_get(job_id)`: fetch async job result * `list_jobs_v1_extract_jobs_get(...)`: paginate past extract jobs **`ExtractRequest` fields:** | Field | Type | Notes | | ----------------- | ---------------- | ---------------------------------- | | `url` | `str` | Required. Target URL. | | `formats` | `[OutputFormat]` | Default `[HTML]`. | | `processing_mode` | `ProcessingMode` | `SYNC` (default) or `ASYNC`. | | `render_js` | `bool` | Headless browser. Default `False`. | | `proxy` | `ProxySettings` | Optional. | | `headers` | `dict[str, str]` | Optional request headers. | | `wait_config` | `WaitConfig` | Optional browser wait policy. | ### extract\_v1\_extract\_post (sync) Scrape one URL and return content in the same response. ```python from geonode_scraper_sdk import ( ApiClient, ExtractRequest, ExtractionApi, OutputFormat, ProcessingMode, ) with ApiClient(configuration) as api_client: api = ExtractionApi(api_client) response = api.extract_v1_extract_post( ExtractRequest( url="https://docs.geonode.com/docs/scraper-api/quick-start", formats=[OutputFormat.MARKDOWN], processing_mode=ProcessingMode.SYNC, ) ) markdown = response.data.markdown if response.data else "" print("Scraped content length:", len(markdown or "")) print("Tokens charged:", response.tokens_charged) print("Preview:", (markdown or "")[:300]) ``` Response (live run): ```text Scraped content length: 15871 Tokens charged: 1 Preview: --- canonical: https://docs.geonode.com/docs/scraper-api/quick-start meta-description: Get your API key, authenticate requests, and choose the right API for your use case. ... ``` ### extract\_v1\_extract\_post (async) Submit a job and receive a `job_id` immediately. Poll with `get_job_result_v1_extract_job_id_get`. ```python import time from geonode_scraper_sdk import ( ApiClient, ExtractRequest, ExtractionApi, OutputFormat, ProcessingMode, ) with ApiClient(configuration) as api_client: api = ExtractionApi(api_client) submit = api.extract_v1_extract_post( ExtractRequest( url="https://docs.geonode.com/docs/scraper-api/quick-start", formats=[OutputFormat.MARKDOWN], processing_mode=ProcessingMode.ASYNC, ) ) print("Job ID:", submit.job_id) while True: job = api.get_job_result_v1_extract_job_id_get(str(submit.job_id)) print("Status:", job.status) if str(job.status).lower().endswith("completed"): if job.data and job.data.markdown: print("Scraped content length:", len(job.data.markdown)) break time.sleep(2) ``` Response (live run): ```text Job ID: 6d92b4d5-c9c2-4b66-aa0a-98c07c3a31da Status: queued Status: completed Scraped content length: 15871 ``` ### get\_job\_result\_v1\_extract\_job\_id\_get Fetch status and content for a single extract job (used after async submit or to re-read a past job). ```python from geonode_scraper_sdk import ApiClient, ExtractionApi JOB_ID = "6d92b4d5-c9c2-4b66-aa0a-98c07c3a31da" with ApiClient(configuration) as api_client: api = ExtractionApi(api_client) job = api.get_job_result_v1_extract_job_id_get(JOB_ID) md_len = len(job.data.markdown) if job.data and job.data.markdown else 0 print("Job ID:", JOB_ID) print("Status:", job.status) print("Markdown length:", md_len) ``` Response (live run): ```text Job ID: 6d92b4d5-c9c2-4b66-aa0a-98c07c3a31da Status: JobStatus.COMPLETED Markdown length: 15871 ``` ### list\_jobs\_v1\_extract\_jobs\_get Paginate past extract jobs. Optional filters: `status`, `start_date`, `end_date`, `page`, `page_size`. ```python from geonode_scraper_sdk import ApiClient, ExtractionApi with ApiClient(configuration) as api_client: api = ExtractionApi(api_client) page = api.list_jobs_v1_extract_jobs_get(page=1, page_size=3) jobs = getattr(page, "items", None) or getattr(page, "jobs", None) or [] for job in jobs: print(job.job_id, job.status) ``` Response (live run): ```text 752b8599-5915-441c-bc1f-b9fb0d938f72 JobStatus.COMPLETED 0bcd424e-b5ab-4b85-9cbd-5b44f7ed433c JobStatus.COMPLETED 0706f42e-d9fd-4076-bfae-3cb365f4b634 JobStatus.COMPLETED ``` ### extract\_v1\_extract\_post (JS + proxy) Same method with `render_js`, residential proxy, and wait config for dynamic pages. ```python from geonode_scraper_sdk import ( ApiClient, ExtractRequest, ExtractionApi, OutputFormat, ProcessingMode, ProxySettings, ProxyType, WaitConfig, WaitUntil, ) with ApiClient(configuration) as api_client: api = ExtractionApi(api_client) response = api.extract_v1_extract_post( ExtractRequest( url="https://example.com", formats=[OutputFormat.MARKDOWN], processing_mode=ProcessingMode.SYNC, render_js=True, proxy=ProxySettings(country="US", type=ProxyType.RESIDENTIAL), wait_config=WaitConfig( wait_until=WaitUntil.NETWORKIDLE, wait_timeout=10000, ), ) ) markdown = response.data.markdown if response.data else "" print("Scraped content length:", len(markdown or "")) print("Tokens charged:", response.tokens_charged) print("Preview:", (markdown or "")[:250]) ``` Response (live run): ```text Scraped content length: 250 Tokens charged: 1 Preview: --- meta-viewport: width=device-width, initial-scale=1 title: Example Domain --- # Example Domain This domain is for use in documentation examples without needing permission. Avoid use in operations. [Learn more](https://iana.org/domains/example) ``` *** ## Batch API `BatchApi` (`/v1/batch`) **Methods:** * `create_batch_v1_batch_post(batch_request)`: submit URLs as one job * `get_batch_status_v1_batch_job_id_get(job_id, page, page_size)`: poll status / results * `list_batch_jobs_v1_batch_jobs_get(...)`: list past batch jobs * `cancel_batch_v1_batch_job_id_delete(job_id)`: cancel a running batch **`BatchRequest` fields:** | Field | Type | Notes | | --------------------- | ---------------- | ------------------------------------------------- | | `urls` | `[str]` | Required. 1-1000 URLs. | | `formats` | `[OutputFormat]` | Optional. | | `render_js` | `bool` | Apply to every URL. | | `proxy` | `ProxySettings` | Optional. | | `headers` | `dict[str, str]` | Optional. | | `wait_config` | `WaitConfig` | Optional. | | `ignore_invalid_urls` | `bool` | Default `True`. Skip bad URLs instead of failing. | ### create\_batch\_v1\_batch\_post Submit many URLs as one batch job. ```python from geonode_scraper_sdk import ApiClient, BatchApi, BatchRequest, OutputFormat with ApiClient(configuration) as api_client: api = BatchApi(api_client) accepted = api.create_batch_v1_batch_post( BatchRequest( urls=[ "https://docs.geonode.com/docs/scraper-api/quick-start", "https://docs.geonode.com/docs/scraper-api", ], formats=[OutputFormat.MARKDOWN], ) ) print("Batch job:", accepted.job_id) print("Accepted URLs:", accepted.accepted_urls) ``` Response (live run): ```text Batch job: 863e95fa-24c0-4638-aa09-a70235fdade1 Accepted URLs: 2 ``` ### get\_batch\_status\_v1\_batch\_job\_id\_get Poll progress and read per-URL results (paginated with `page`, `page_size`). ```python import time from geonode_scraper_sdk import ApiClient, BatchApi, BatchRequest, OutputFormat with ApiClient(configuration) as api_client: api = BatchApi(api_client) accepted = api.create_batch_v1_batch_post( BatchRequest( urls=["https://docs.geonode.com/docs/scraper-api/quick-start"], formats=[OutputFormat.MARKDOWN], ) ) while True: status = api.get_batch_status_v1_batch_job_id_get( job_id=accepted.job_id, page=1, page_size=10 ) print( status.status, status.completed_urls, "/", status.total_urls, ) if str(status.status).lower().endswith("completed"): break time.sleep(3) ``` Response (live run): ```text queued 0 / 1 processing 0 / 1 completed 1 / 1 ``` ### list\_batch\_jobs\_v1\_batch\_jobs\_get List batch jobs with optional `status`, `start_date`, `end_date`, pagination. ```python from geonode_scraper_sdk import ApiClient, BatchApi with ApiClient(configuration) as api_client: api = BatchApi(api_client) page = api.list_batch_jobs_v1_batch_jobs_get(page=1, page_size=3) for job in (page.jobs or [])[:3]: print(job.job_id, job.status, job.completed_urls, "/", job.accepted_urls) ``` Response (live run): ```text 1255853d-eef7-45c4-9abe-7e3863535584 JobStatus.COMPLETED 1 / 1 863e95fa-24c0-4638-aa09-a70235fdade1 JobStatus.COMPLETED 2 / 2 d403ad87-66ce-4a48-823b-21e8ebbfb899 JobStatus.COMPLETED 2 / 2 ``` ### cancel\_batch\_v1\_batch\_job\_id\_delete Stop scheduling new batch items. In-flight extractions drain. ```python from geonode_scraper_sdk import ApiClient, BatchApi, BatchRequest, OutputFormat with ApiClient(configuration) as api_client: api = BatchApi(api_client) accepted = api.create_batch_v1_batch_post( BatchRequest( urls=["https://docs.geonode.com/docs/scraper-api/quick-start"], formats=[OutputFormat.MARKDOWN], ) ) cancelled = api.cancel_batch_v1_batch_job_id_delete(accepted.job_id) print("Cancelled:", cancelled.job_id, cancelled.status) ``` Response (live run): ```text Cancelled: 883bd4ad-cb9c-49bc-99f5-d007fce5a217 JobStatus.CANCELLED ``` *** ## Crawl API `CrawlApi` (`/v1/crawl`) **Methods:** * `create_crawl_v1_crawl_post(crawl_request)`: start crawl from seed URL * `get_crawl_status_v1_crawl_job_id_get(job_id, page, page_size)`: poll status / pages * `list_crawl_jobs_v1_crawl_jobs_get(...)`: list past crawl jobs * `cancel_crawl_v1_crawl_job_id_delete(job_id)`: cancel a running crawl **`CrawlRequest` fields:** | Field | Type | Notes | | -------------------- | ---------------- | --------------------------------- | | `url` | `str` | Required seed URL. | | `depth` | `int` | BFS depth, 1-10. Default `2`. | | `limit` | `int` | Max pages, 1-10000. Default `50`. | | `same_domain_only` | `bool` | Default `True`. | | `include_subdomains` | `bool` | Default `False`. | | `formats` | `[OutputFormat]` | Optional per-page formats. | | `render_js` | `bool` | Optional. | | `proxy` | `ProxySettings` | Optional. | | `wait_config` | `WaitConfig` | Optional. | ### create\_crawl\_v1\_crawl\_post Start a crawl from a seed URL. ```python from geonode_scraper_sdk import ApiClient, CrawlApi, CrawlRequest, OutputFormat with ApiClient(configuration) as api_client: api = CrawlApi(api_client) accepted = api.create_crawl_v1_crawl_post( CrawlRequest( url="https://docs.geonode.com/docs/scraper-api", depth=2, limit=5, formats=[OutputFormat.MARKDOWN], same_domain_only=True, ) ) print("Crawl job:", accepted.job_id) print("Estimated pages:", accepted.estimated_pages) ``` Response (live run): ```text Crawl job: efd86d85-f843-4b71-ad58-d5fec23a0ec3 Estimated pages: 5 ``` ### get\_crawl\_status\_v1\_crawl\_job\_id\_get Poll crawl progress and read scraped pages (paginated). ```python import time from geonode_scraper_sdk import ApiClient, CrawlApi, CrawlRequest, OutputFormat with ApiClient(configuration) as api_client: api = CrawlApi(api_client) accepted = api.create_crawl_v1_crawl_post( CrawlRequest( url="https://docs.geonode.com/docs/scraper-api", depth=2, limit=5, formats=[OutputFormat.MARKDOWN], same_domain_only=True, ) ) while True: status = api.get_crawl_status_v1_crawl_job_id_get( job_id=accepted.job_id, page=1, page_size=10 ) print( status.status, status.completed_pages, "/", status.total_pages, ) if str(status.status).lower().endswith("completed"): break time.sleep(4) ``` Response (live run): ```text queued 0 / 5 processing 0 / 5 ... completed 5 / 5 ``` ### list\_crawl\_jobs\_v1\_crawl\_jobs\_get List crawl jobs. Optional filters: `url`, `status`, `start_date`, `end_date`. ```python from geonode_scraper_sdk import ApiClient, CrawlApi with ApiClient(configuration) as api_client: api = CrawlApi(api_client) page = api.list_crawl_jobs_v1_crawl_jobs_get(page=1, page_size=2) for job in (page.jobs or [])[:2]: print(job.job_id, job.status, job.completed_pages, "/", job.total_pages) ``` Response (live run): ```text 34ebec40-5353-4933-ac11-6be09b85c448 JobStatus.COMPLETED 1 / 1 efd86d85-f843-4b71-ad58-d5fec23a0ec3 JobStatus.COMPLETED 5 / 5 ``` ### cancel\_crawl\_v1\_crawl\_job\_id\_delete Cancel a queued or processing crawl. ```python from geonode_scraper_sdk import ApiClient, CrawlApi, CrawlRequest, OutputFormat with ApiClient(configuration) as api_client: api = CrawlApi(api_client) accepted = api.create_crawl_v1_crawl_post( CrawlRequest( url="https://docs.geonode.com/docs/scraper-api", depth=1, limit=50, formats=[OutputFormat.MARKDOWN], ) ) cancelled = api.cancel_crawl_v1_crawl_job_id_delete(accepted.job_id) print("Cancelled:", cancelled.job_id, cancelled.status) ``` Response (live run): ```text Cancelled: d8e5dec2-7450-4023-8fba-858243ac5339 JobStatus.CANCELLED ``` *** ## Map API `MapApi` (`/v1/map`) **Methods:** * `map_urls_v1_map_post(map_request)`: discover URLs under a base URL * `list_map_jobs_v1_map_jobs_get(...)`: list past map jobs * `get_map_job_v1_map_job_id_get(job_id)`: fetch a completed map job **`MapRequest` fields:** | Field | Type | Notes | | ------------------------- | ------ | --------------------------------------- | | `url` | `str` | Required base URL. | | `include_subdomains` | `bool` | Widen discovery scope. Default `False`. | | `ignore_query_parameters` | `bool` | Normalize URLs. Default `True`. | | `search` | `str` | Optional path/url filter. | ### map\_urls\_v1\_map\_post Discover URLs synchronously (returns links inline). Large sites may also create a persisted map job you can fetch later. ```python from geonode_scraper_sdk import ApiClient, MapApi, MapRequest with ApiClient(configuration) as api_client: api = MapApi(api_client) result = api.map_urls_v1_map_post( MapRequest(url="https://docs.geonode.com/docs/scraper-api") ) print("Link count:", len(result.links or [])) for link in (result.links or [])[:5]: print(link.source, link.url) ``` Response (live run): ```text Link count: 112 sitemap https://docs.geonode.com/docs/scraper-api sitemap https://docs.geonode.com/docs/scraper-api/quick-start sitemap https://docs.geonode.com/docs/scraper-api/additional-resources/choosing_scraper_api_plan sitemap https://docs.geonode.com/docs/scraper-api/additional-resources/faq sitemap https://docs.geonode.com/docs/scraper-api/additional-resources/pricing-and-requests ``` ### list\_map\_jobs\_v1\_map\_jobs\_get List past map jobs. Optional filters: `url`, `status`, `start_date`, `end_date`. ```python from geonode_scraper_sdk import ApiClient, MapApi with ApiClient(configuration) as api_client: api = MapApi(api_client) page = api.list_map_jobs_v1_map_jobs_get(page=1, page_size=2) for job in (page.jobs or [])[:2]: print(job.job_id, job.status, job.url) ``` Response (live run): ```text b38bdc1c-2b53-4ccb-8947-6673cf5b53e0 JobStatus.COMPLETED https://docs.geonode.com/docs/scraper-api bc074160-7d94-4a74-bb77-1e4f96f9c5cb JobStatus.COMPLETED https://docs.geonode.com/docs/scraper-api ``` ### get\_map\_job\_v1\_map\_job\_id\_get Retrieve full link list for a completed map job. ```python from geonode_scraper_sdk import ApiClient, MapApi JOB_ID = "b38bdc1c-2b53-4ccb-8947-6673cf5b53e0" with ApiClient(configuration) as api_client: api = MapApi(api_client) detail = api.get_map_job_v1_map_job_id_get(JOB_ID) links = detail.links or [] print("Job ID:", JOB_ID) print("Link count:", len(links)) for link in links[:3]: print(link.source, link.url) ``` Response (live run): ```text Job ID: b38bdc1c-2b53-4ccb-8947-6673cf5b53e0 Link count: 112 sitemap https://docs.geonode.com/docs/scraper-api sitemap https://docs.geonode.com/docs/scraper-api/quick-start sitemap https://docs.geonode.com/docs/scraper-api/additional-resources/choosing_scraper_api_plan ``` *** ## Search API `SearchApi` (`/v1/search`) **Methods:** * `search_v1_search_post(search_request)`: run a search query * `list_search_jobs_v1_search_jobs_get(...)`: list past search jobs * `get_search_job_v1_search_job_id_get(job_id)`: fetch a completed search job **`SearchRequest` fields:** | Field | Type | Notes | | ------------ | ----- | ----------------------------------------------- | | `query` | `str` | Required search string. | | `page` | `int` | Result page 1-20. Default `1`. | | `safe` | `str` | `off` \| `moderate` \| `strict`. Default `off`. | | `time_range` | `str` | Optional: `day`, `week`, `month`, `year`. | | `locale` | `str` | Optional locale hint. | ### search\_v1\_search\_post Run a search query. Returns results inline and a `job_id` for later lookup. ```python from geonode_scraper_sdk import ApiClient, SearchApi, SearchRequest with ApiClient(configuration) as api_client: api = SearchApi(api_client) result = api.search_v1_search_post( SearchRequest(query="geonode scraper api") ) print("Job ID:", result.job_id) print("Hit count:", len(result.results or [])) for hit in (result.results or [])[:5]: print(hit.position, hit.title, hit.url) ``` Response (live run): ```text Job ID: 6aef3e93-f714-4eb5-ab7a-c0b00fee5dea Hit count: 15 1 Wholesale proxies & web data infrastructure | Geonode https://geonode.com/ 2 Geonode - PyPI https://pypi.org/user/Geonode/ 3 pavel.s - PyPI https://pypi.org/user/pavel.s/ 4 Geonode Documentation | Geonode https://docs.geonode.com/ 5 GeoNode https://geonode.org/ ``` ### list\_search\_jobs\_v1\_search\_jobs\_get List past search jobs. Optional filters: `query`, `status`, `start_date`, `end_date`. ```python from geonode_scraper_sdk import ApiClient, SearchApi with ApiClient(configuration) as api_client: api = SearchApi(api_client) page = api.list_search_jobs_v1_search_jobs_get(page=1, page_size=2) for job in (page.jobs or [])[:2]: print(job.job_id, job.query, job.status) ``` Response (live run): ```text 6aef3e93-f714-4eb5-ab7a-c0b00fee5dea geonode scraper api JobStatus.COMPLETED 0a0375f1-a36a-4c51-929d-5fa4b2425877 geonode scraper api JobStatus.COMPLETED ``` ### get\_search\_job\_v1\_search\_job\_id\_get Re-fetch full results for a search job. ```python from geonode_scraper_sdk import ApiClient, SearchApi JOB_ID = "6aef3e93-f714-4eb5-ab7a-c0b00fee5dea" with ApiClient(configuration) as api_client: api = SearchApi(api_client) detail = api.get_search_job_v1_search_job_id_get(JOB_ID) hits = detail.results or [] print("Job ID:", JOB_ID) print("Hit count:", len(hits)) for hit in hits[:3]: print(hit.position, hit.title, hit.url) ``` Response (live run): ```text Job ID: 6aef3e93-f714-4eb5-ab7a-c0b00fee5dea Hit count: 15 1 Wholesale proxies & web data infrastructure | Geonode https://geonode.com/ 2 Geonode - PyPI https://pypi.org/user/Geonode/ 3 pavel.s - PyPI https://pypi.org/user/pavel.s/ ``` *** ## Usage API `UsageApi` (`/v1/usage`) **Methods:** * `get_concurrency_usage_v1_usage_concurrency_get()`: live concurrency in use vs plan limit ### get\_concurrency\_usage\_v1\_usage\_concurrency\_get Check live work-concurrency slots against your plan limit. ```python from geonode_scraper_sdk import ApiClient, UsageApi with ApiClient(configuration) as api_client: api = UsageApi(api_client) usage = api.get_concurrency_usage_v1_usage_concurrency_get() print(usage.work_concurrency_in_use, usage.work_concurrency_limit) ``` Response (live run): ```text 0 50 ``` *** ## Statistics API `StatisticsApi` (`/v1/statistics`) **Methods:** * `get_statistics_v1_statistics_get(...)`: extraction counts, tokens, success rate ### get\_statistics\_v1\_statistics\_get Historical extraction counts, token usage, success rate. ```python from geonode_scraper_sdk import ApiClient, StatisticsApi with ApiClient(configuration) as api_client: api = StatisticsApi(api_client) stats = api.get_statistics_v1_statistics_get() print("Extraction count:", stats.extraction_count) print("Success rate:", stats.success_rate) print("Recent token days:", len(stats.tokens_used or [])) ``` Response (live run): ```text Extraction count: 4378 Success rate: 0.95112 Recent token days: 7 ``` *** ## System API `SystemApi` (`/health`) **Methods:** * `health_check_health_get()`: service health check ### health\_check\_health\_get Confirm the Scraper API is up. ```python from geonode_scraper_sdk import ApiClient, SystemApi with ApiClient(configuration) as api_client: api = SystemApi(api_client) health = api.health_check_health_get() print(health.service, health.status, health.version) ``` Response (live run): ```text Scraper API HealthStatus.OK 0.1.0 ``` *** ## Webhooks API `WebhooksApi` (`/v1/webhooks`) Event types: `extract_completed`, `batch_completed`, `crawl_completed`. **Methods:** * `create_webhook_v1_webhooks_post(webhook_create)`: register a webhook * `list_webhooks_v1_webhooks_get(...)`: list webhooks * `get_webhook_v1_webhooks_webhook_id_get(webhook_id)`: get one webhook * `update_webhook_v1_webhooks_webhook_id_patch(webhook_id, webhook_update)`: update * `delete_webhook_v1_webhooks_webhook_id_delete(webhook_id)`: delete * `list_deliveries_v1_webhooks_webhook_id_deliveries_get(...)`: delivery history * `rotate_secret_v1_webhooks_webhook_id_rotate_secret_post(webhook_id)`: rotate signing secret **`WebhookCreate` fields:** | Field | Type | Notes | | ------------- | ------------------ | ------------------------------------------------------------- | | `url` | `str` | Required callback URL. | | `event_type` | `WebhookEventType` | `extract_completed`, `batch_completed`, or `crawl_completed`. | | `description` | `str` | Optional label. | ### create\_webhook\_v1\_webhooks\_post ```python from geonode_scraper_sdk import ( ApiClient, WebhookCreate, WebhookEventType, WebhooksApi, ) with ApiClient(configuration) as api_client: api = WebhooksApi(api_client) created = api.create_webhook_v1_webhooks_post( WebhookCreate( url="https://example.com/webhook", event_type=WebhookEventType.EXTRACT_COMPLETED, description="sdk guide demo", ) ) print("Created:", created.id, created.url, created.event_type) ``` Response (live run): ```text Created: b9ea9375-90a0-47cb-bb37-48c4c0804925 https://example.com/webhook WebhookEventType.EXTRACT_COMPLETED ``` ### list\_webhooks\_v1\_webhooks\_get ```python from geonode_scraper_sdk import ApiClient, WebhooksApi with ApiClient(configuration) as api_client: api = WebhooksApi(api_client) page = api.list_webhooks_v1_webhooks_get(page=1, page_size=5) items = getattr(page, "items", None) or [] print("Webhook count:", len(items)) for wh in items: print(wh.id, wh.url, wh.event_type) ``` Response (live run): ```text Webhook count: 0 ``` ### get\_webhook\_v1\_webhooks\_webhook\_id\_get ```python from geonode_scraper_sdk import ApiClient, WebhooksApi WEBHOOK_ID = "b9ea9375-90a0-47cb-bb37-48c4c0804925" with ApiClient(configuration) as api_client: api = WebhooksApi(api_client) wh = api.get_webhook_v1_webhooks_webhook_id_get(WEBHOOK_ID) print(wh.url, wh.event_type, wh.is_active) ``` Response (live run): ```text https://example.com/webhook WebhookEventType.EXTRACT_COMPLETED True ``` ### update\_webhook\_v1\_webhooks\_webhook\_id\_patch ```python from geonode_scraper_sdk import ApiClient, WebhookUpdate, WebhooksApi WEBHOOK_ID = "b9ea9375-90a0-47cb-bb37-48c4c0804925" with ApiClient(configuration) as api_client: api = WebhooksApi(api_client) updated = api.update_webhook_v1_webhooks_webhook_id_patch( WEBHOOK_ID, WebhookUpdate(description="updated by sdk guide"), ) print("Description:", updated.description) ``` Response (live run): ```text Description: updated by sdk guide ``` ### rotate\_secret\_v1\_webhooks\_webhook\_id\_rotate\_secret\_post ```python from geonode_scraper_sdk import ApiClient, WebhooksApi WEBHOOK_ID = "b9ea9375-90a0-47cb-bb37-48c4c0804925" with ApiClient(configuration) as api_client: api = WebhooksApi(api_client) rotated = api.rotate_secret_v1_webhooks_webhook_id_rotate_secret_post(WEBHOOK_ID) print("New secret length:", len(rotated.secret or "")) ``` Response (live run): ```text New secret length: 64 ``` ### list\_deliveries\_v1\_webhooks\_webhook\_id\_deliveries\_get ```python from geonode_scraper_sdk import ApiClient, WebhooksApi WEBHOOK_ID = "b9ea9375-90a0-47cb-bb37-48c4c0804925" with ApiClient(configuration) as api_client: api = WebhooksApi(api_client) page = api.list_deliveries_v1_webhooks_webhook_id_deliveries_get( WEBHOOK_ID, page=1, page_size=5 ) items = getattr(page, "items", None) or [] print("Delivery count:", len(items)) ``` Response (live run): ```text Delivery count: 0 ``` ### delete\_webhook\_v1\_webhooks\_webhook\_id\_delete ```python from geonode_scraper_sdk import ApiClient, WebhooksApi WEBHOOK_ID = "b9ea9375-90a0-47cb-bb37-48c4c0804925" with ApiClient(configuration) as api_client: api = WebhooksApi(api_client) api.delete_webhook_v1_webhooks_webhook_id_delete(WEBHOOK_ID) print("Deleted:", WEBHOOK_ID) ``` Response (live run): ```text Deleted: b9ea9375-90a0-47cb-bb37-48c4c0804925 ``` *** ## Error handling Non-2xx HTTP responses raise `ApiException` with `status`, `body`, and parsed `data`. Invalid models can fail before the HTTP call (Pydantic `ValidationError`). **Validation error** (empty URL): ```python from geonode_scraper_sdk import ApiClient, ExtractRequest, ExtractionApi with ApiClient(configuration) as api_client: api = ExtractionApi(api_client) try: api.extract_v1_extract_post(ExtractRequest(url="")) except Exception as exc: print(type(exc).__name__, exc) ``` Response (live run): ```text ValidationError 1 validation error for ExtractRequest url String should have at least 1 character [type=string_too_short, input_value='', input_type=str] ``` **API error** (fake job ID): ```python from geonode_scraper_sdk import ApiClient, ApiException, ExtractionApi with ApiClient(configuration) as api_client: api = ExtractionApi(api_client) try: api.get_job_result_v1_extract_job_id_get( "00000000-0000-0000-0000-000000000000" ) except ApiException as exc: print(exc.status) print(exc.body) ``` Response (live run): ```text 404 {"code":"NOT_FOUND","message":"Job 00000000-0000-0000-0000-000000000000 not found","correlation_id":"f430c05f-5aea-4e98-bcf2-34fabb743f8c","retryable":false,"details":null} ``` For package versions and `*_with_http_info()` variants, see [geonode-scraper-sdk on PyPI](https://pypi.org/project/geonode-scraper-sdk/). # Tools Core (/docs/scraper-api/developer-guides/tools-core) {/* DRAFT: do not publish / do not add to meta.json until approved */} `geonode-scraper-tools-core` gives you named scraper operations and JSON-friendly dictionaries. Call them via `ScraperToolService`, or select a subset with `get_operations()`. LangChain and CrewAI packages use this layer under the hood. Responses below are from live runs against `https://scraper.geonode.io`. ## Setup ```bash pip install geonode-scraper-tools-core python-dotenv ``` Create a `.env` file: ```bash GEONODE_SCRAPER_API_KEY=your_geonode_key SCRAPER_API_BASE_URL=https://scraper.geonode.io ``` ```python import os from dotenv import load_dotenv from geonode_scraper_tools_core import ScraperToolSettings, ScraperToolService load_dotenv() settings = ScraperToolSettings( host=os.environ.get("SCRAPER_API_BASE_URL", "https://scraper.geonode.io"), api_key=os.environ["GEONODE_SCRAPER_API_KEY"], ) service = ScraperToolService(settings) ``` Use `https://scraper.geonode.io` for production. ### ScraperToolSettings fields | Field | Type | Notes | | ----------------------- | ------- | -------------------------------------- | | `host` | `str` | Required. API base URL. | | `api_key` | `str` | Required. Geonode Scraper API key. | | `verify_ssl` | `bool` | Default `True`. | | `request_timeout` | timeout | Optional HTTP timeout. | | `max_retries` | `int` | Default `0`. | | `retry_backoff_seconds` | `float` | Default `1.0`. | | `poll_interval_seconds` | `float` | Default `3.0` (used by wait helpers). | | `poll_timeout_seconds` | `float` | Default `60.0` (used by wait helpers). | ### Response shape Every service method returns a dict like: ```python { "ok": True, "operation": "extract", "attempts": 1, "result": { ... }, # real payload } ``` Wait helpers also include `poll_attempts`. Read the payload from `response["result"]`. ## Operations overview | Group | Operations | | -------------- | ----------------------------------------------------------------------- | | **Extraction** | `extract`, `get_job_result`, `wait_for_job`, `list_jobs` | | **Batch** | `create_batch`, `get_batch_status`, `wait_for_batch`, `list_batch_jobs` | | **Crawl** | `create_crawl`, `get_crawl_status`, `wait_for_crawl`, `list_crawl_jobs` | | **Map** | `map_urls`, `list_map_jobs`, `get_map_job` | | **Search** | `search`, `list_search_jobs`, `get_search_job` | | **Account** | `get_statistics`, `get_concurrency_usage`, `health_check` | `wait_for_job`, `wait_for_batch`, and `wait_for_crawl` poll until a job finishes, so you do not write poll loops yourself. *** ## Extraction **Operations:** `extract`, `get_job_result`, `wait_for_job`, `list_jobs` ### extract (sync) Scrape one URL and wait for content in the same call. **Main inputs:** `url`, `formats` (`markdown` / `html`), `processing_mode` (`sync` / `async`), `render_js`, `proxy_country`, `proxy_type`, `headers`, `wait_for`, `wait_timeout`, `wait_until` ```python response = service.extract( url="https://docs.geonode.com/docs/scraper-api/quick-start", formats=["markdown"], processing_mode="sync", ) result = response["result"] markdown = (result.get("data") or {}).get("markdown") or "" print("ok:", response["ok"]) print("tokens:", result.get("tokens_charged")) print("markdown_len:", len(markdown)) print("preview:", markdown[:200]) ``` Response (live run): ```text ok: True tokens: 1 markdown_len: 15871 preview: --- canonical: https://docs.geonode.com/docs/scraper-api/quick-start meta-description: Get your API key, authenticate requests, and choose the right API for your use case. ... ``` ### extract (async) Same method with `processing_mode="async"`. Returns a `job_id` immediately. ```python response = service.extract( url="https://docs.geonode.com/docs/scraper-api/quick-start", formats=["markdown"], processing_mode="async", ) result = response["result"] print("job_id:", result["job_id"]) print("status:", result["status"]) ``` Response (live run): ```text job_id: 8a2bbbd5-7034-412b-acdf-e30226ae32b6 status: queued ``` ### wait\_for\_job Poll an async extract job until it finishes (or times out). **Inputs:** `job_id`, optional `timeout_seconds`, `poll_interval_seconds` ```python response = service.wait_for_job( job_id="8a2bbbd5-7034-412b-acdf-e30226ae32b6", timeout_seconds=120, ) result = response["result"] markdown = (result.get("data") or {}).get("markdown") or "" print("status:", result["status"]) print("markdown_len:", len(markdown)) print("poll_attempts:", response.get("poll_attempts")) ``` Response (live run): ```text status: completed markdown_len: 15871 poll_attempts: 2 ``` ### get\_job\_result Fetch the current state or final result for one extract job (no waiting). **Inputs:** `job_id` ```python response = service.get_job_result( job_id="8a2bbbd5-7034-412b-acdf-e30226ae32b6", ) result = response["result"] print("status:", result["status"]) print("tokens:", result.get("tokens_charged")) ``` Response (live run): ```text status: completed tokens: 1 ``` ### list\_jobs List past extract jobs. **Inputs:** optional `job_id`, `url`, `status`, `output`, `start_date`, `end_date`, `page`, `page_size` ```python response = service.list_jobs(page=1, page_size=3) result = response["result"] print("page:", result["page"], "page_size:", result["page_size"]) for job in result.get("jobs") or []: print(job["job_id"], job["status"], job.get("url")) ``` Response (live run): ```text page: 1 page_size: 3 752b8599-5915-441c-bc1f-b9fb0d938f72 completed ... 0bcd424e-b5ab-4b85-9cbd-5b44f7ed433c completed http://example.com/ 0dbb126c-4331-4b4b-9d55-0c380ba69ae7 completed http://example.com/ ``` *** ## Batch **Operations:** `create_batch`, `get_batch_status`, `wait_for_batch`, `list_batch_jobs` ### create\_batch Submit many URLs as one job. **Inputs:** `urls`, `formats`, optional `render_js`, `proxy_country`, `proxy_type`, `headers` ```python response = service.create_batch( urls=[ "https://docs.geonode.com/docs/scraper-api/quick-start", "https://docs.geonode.com/docs/scraper-api", ], formats=["markdown"], ) result = response["result"] print("job_id:", result["job_id"]) print("accepted_urls:", result["accepted_urls"]) print("status:", result["status"]) ``` Response (live run): ```text job_id: 55b11790-5c7f-4945-bc13-bd9a365a1835 accepted_urls: 2 status: queued ``` ### get\_batch\_status Poll progress and partial results (paginated). **Inputs:** `job_id`, `page`, `page_size` ```python response = service.get_batch_status( job_id="55b11790-5c7f-4945-bc13-bd9a365a1835", page=1, page_size=10, ) result = response["result"] print( result["status"], result["completed_urls"], "/", result["total_urls"], ) ``` Response (live run, mid-job): ```text processing 0 / 2 ``` ### wait\_for\_batch Poll until the batch finishes. **Inputs:** `job_id`, optional `timeout_seconds`, `poll_interval_seconds` ```python response = service.wait_for_batch( job_id="55b11790-5c7f-4945-bc13-bd9a365a1835", timeout_seconds=120, ) result = response["result"] print( result["status"], result["completed_urls"], "/", result["total_urls"], ) print("poll_attempts:", response.get("poll_attempts")) ``` Response (live run): ```text completed 2 / 2 poll_attempts: 3 ``` ### list\_batch\_jobs **Inputs:** optional `status`, `start_date`, `end_date`, `page`, `page_size` ```python response = service.list_batch_jobs(page=1, page_size=3) for job in (response["result"].get("jobs") or []): print( job["job_id"], job["status"], job["completed_urls"], "/", job["accepted_urls"], ) ``` Response (live run): ```text 55b11790-5c7f-4945-bc13-bd9a365a1835 completed 2 / 2 d8e92d3a-c939-4ec6-b133-3e04c6de71d1 completed 2 / 2 883bd4ad-cb9c-49bc-99f5-d007fce5a217 completed 0 / 1 ``` *** ## Crawl **Operations:** `create_crawl`, `get_crawl_status`, `wait_for_crawl`, `list_crawl_jobs` ### create\_crawl Start from a seed URL. **Inputs:** `url`, `depth`, `limit`, `formats`, `same_domain_only`, `include_subdomains`, optional `render_js`, `proxy_country`, `proxy_type` ```python response = service.create_crawl( url="https://docs.geonode.com/docs/scraper-api", depth=2, limit=3, formats=["markdown"], same_domain_only=True, ) result = response["result"] print("job_id:", result["job_id"]) print("estimated_pages:", result["estimated_pages"]) print("status:", result["status"]) ``` Response (live run): ```text job_id: 9f447ced-4aa3-45a3-b9af-421f751b8cec estimated_pages: 3 status: queued ``` ### get\_crawl\_status **Inputs:** `job_id`, `page`, `page_size` ```python response = service.get_crawl_status( job_id="9f447ced-4aa3-45a3-b9af-421f751b8cec", page=1, page_size=10, ) result = response["result"] print( result["status"], result.get("completed_pages"), "/", result.get("total_pages"), ) ``` Response (live run, early poll): ```text processing 0 / 1 ``` ### wait\_for\_crawl **Inputs:** `job_id`, optional `timeout_seconds`, `poll_interval_seconds` ```python response = service.wait_for_crawl( job_id="9f447ced-4aa3-45a3-b9af-421f751b8cec", timeout_seconds=180, ) result = response["result"] print( result["status"], result["completed_pages"], "/", result["total_pages"], ) print("poll_attempts:", response.get("poll_attempts")) ``` Response (live run): ```text completed 3 / 3 poll_attempts: 19 ``` ### list\_crawl\_jobs **Inputs:** optional `url`, `status`, `start_date`, `end_date`, `page`, `page_size` ```python response = service.list_crawl_jobs(page=1, page_size=2) for job in (response["result"].get("jobs") or []): print( job["job_id"], job["status"], job["completed_pages"], "/", job["total_pages"], ) ``` Response (live run): ```text 9f447ced-4aa3-45a3-b9af-421f751b8cec completed 3 / 3 d8e5dec2-7450-4023-8fba-858243ac5339 completed 0 / 1 ``` *** ## Map **Operations:** `map_urls`, `list_map_jobs`, `get_map_job` ### map\_urls Discover URLs under a base URL (sitemap + HTML links). Does not scrape page content. **Inputs:** `url`, optional `search`, `include_subdomains`, `ignore_query_parameters` ```python response = service.map_urls( url="https://docs.geonode.com/docs/scraper-api", ) result = response["result"] links = result.get("links") or [] print("link_count:", result.get("links_count") or len(links)) for link in links[:5]: print(link.get("source"), link.get("url")) ``` Response (live run): ```text link_count: 112 sitemap https://docs.geonode.com/docs/scraper-api sitemap https://docs.geonode.com/docs/scraper-api/quick-start sitemap https://docs.geonode.com/docs/scraper-api/additional-resources/choosing_scraper_api_plan sitemap https://docs.geonode.com/docs/scraper-api/additional-resources/faq sitemap https://docs.geonode.com/docs/scraper-api/additional-resources/pricing-and-requests ``` ### list\_map\_jobs **Inputs:** optional `url`, `status`, `start_date`, `end_date`, `page`, `page_size` ```python response = service.list_map_jobs(page=1, page_size=2) for job in (response["result"].get("jobs") or []): print(job["job_id"], job["status"], job.get("links_count"), job.get("url")) ``` Response (live run): ```text c7cb6e40-ddfa-416a-b255-00dfdfd717bc completed 112 https://docs.geonode.com/docs/scraper-api b38bdc1c-2b53-4ccb-8947-6673cf5b53e0 completed 112 https://docs.geonode.com/docs/scraper-api ``` ### get\_map\_job **Inputs:** `job_id` ```python response = service.get_map_job( job_id="c7cb6e40-ddfa-416a-b255-00dfdfd717bc", ) result = response["result"] print("status:", result["status"]) print("link_count:", result.get("links_count")) for link in (result.get("links") or [])[:3]: print(link.get("source"), link.get("url")) ``` Response (live run): ```text status: completed link_count: 112 sitemap https://docs.geonode.com/docs/scraper-api sitemap https://docs.geonode.com/docs/scraper-api/quick-start sitemap https://docs.geonode.com/docs/scraper-api/additional-resources/choosing_scraper_api_plan ``` *** ## Search **Operations:** `search`, `list_search_jobs`, `get_search_job` ### search Run a web search and get ranked hits (also returns a `job_id`). **Inputs:** `query`, optional `locale`, `page`, `safe` (`off` / `moderate` / `strict`), `time_range` (`day` / `week` / `month` / `year`) ```python response = service.search(query="geonode scraper api") result = response["result"] print("job_id:", result["job_id"]) print("hit_count:", result.get("results_count") or len(result.get("results") or [])) for hit in (result.get("results") or [])[:5]: print(hit["position"], hit["title"], hit["url"]) ``` Response (live run): ```text job_id: 29809a61-49b6-4bf6-925b-aa32dd91e441 hit_count: 15 1 Wholesale proxies & web data infrastructure | Geonode https://geonode.com/ 2 Geonode - PyPI https://pypi.org/user/Geonode/ 3 pavel.s - PyPI https://pypi.org/user/pavel.s/ 4 Geonode Documentation | Geonode https://docs.geonode.com/ 5 GeoNode https://geonode.org/ ``` ### list\_search\_jobs **Inputs:** optional `query`, `status`, `start_date`, `end_date`, `page`, `page_size` ```python response = service.list_search_jobs(page=1, page_size=2) for job in (response["result"].get("jobs") or []): print(job["job_id"], job.get("query"), job["status"], job.get("results_count")) ``` Response (live run): ```text 29809a61-49b6-4bf6-925b-aa32dd91e441 geonode scraper api completed 15 6aef3e93-f714-4eb5-ab7a-c0b00fee5dea geonode scraper api completed 15 ``` ### get\_search\_job **Inputs:** `job_id` ```python response = service.get_search_job( job_id="29809a61-49b6-4bf6-925b-aa32dd91e441", ) result = response["result"] print("status:", result["status"]) print("hit_count:", result.get("results_count")) for hit in (result.get("results") or [])[:3]: print(hit["position"], hit["title"], hit["url"]) ``` Response (live run): ```text status: completed hit_count: 15 1 Wholesale proxies & web data infrastructure | Geonode https://geonode.com/ 2 Geonode - PyPI https://pypi.org/user/Geonode/ 3 pavel.s - PyPI https://pypi.org/user/pavel.s/ ``` *** ## Statistics, usage, and health **Operations:** `get_statistics`, `get_concurrency_usage`, `health_check` ### get\_statistics **Inputs:** optional `start_date`, `end_date` ```python response = service.get_statistics() result = response["result"] print("extraction_count:", result.get("extraction_count")) print("success_rate:", result.get("success_rate")) ``` Response (live run): ```text extraction_count: 4383 success_rate: 0.95117 ``` ### get\_concurrency\_usage No inputs. Checks live concurrency slots vs your plan limit. ```python response = service.get_concurrency_usage() result = response["result"] print( result["work_concurrency_in_use"], "/", result["work_concurrency_limit"], ) ``` Response (live run): ```text 0 / 50 ``` ### health\_check No inputs. Confirms the Scraper API is up. ```python response = service.health_check() result = response["result"] print(result.get("service"), result.get("status"), result.get("version")) ``` Response (live run): ```text Scraper API ok 0.1.0 ``` *** ## Selecting a subset of operations Use `get_operations()` when you only want some tools (for example before wiring into an agent framework). ```python from geonode_scraper_tools_core import get_operations ops = get_operations(["extract", "map_urls", "create_batch", "wait_for_batch"]) print([op.key for op in ops]) ``` Response: ```text ['extract', 'map_urls', 'create_batch', 'wait_for_batch'] ``` `OPERATIONS` is the full registry (all 21). Each entry has `key`, `tool_name`, `description`, `args_schema`, and `service_method`. *** ## Used by LangChain and CrewAI Most agent users install: * `geonode-scraper-langchain` * `geonode-scraper-crewai` Those packages wrap this core service. Use tools-core directly when you want named operations and plain dicts in your own Python code, without an agent framework. For package versions and changelog, see [geonode-scraper-tools-core on PyPI](https://pypi.org/project/geonode-scraper-tools-core/). # Before You Start (/docs/scraper-api/getting-started/00_before_you_start) import { Step, Steps } from "fumadocs-ui/components/steps"; This guide covers the basic requirements needed to follow the Extraction guides and ensure your account is set up and ready to make requests. ## Prerequisites #### Create a Geonode Account Sign up for a Geonode account if you do not already have one. #### Get Access to the Scraper API Make sure your account has access to the Scraper API. #### Generate an API Key 1. Go to `https://app.geonode.com/scraper-api`. 2. Click `Get Code`. 3. The `API & Integrations` modal will open. 4. Click `Copy API Key` to copy your API key. API & Integrations Modal ## Authentication Include your API key in the `X-Api-Key` request header. ```bash X-Api-Key: YOUR_API_KEY ``` Requests without a valid API key will be rejected. ## Base URL All Extraction API endpoints are available through the following base URL: ```bash https://scraper.geonode.io ``` ## Next Steps Now that your account is set up, continue to **Understanding Extraction** to learn how the Extraction API works and when to use its different features. # Build an Amazon Product Snapshot from Search Results (/docs/scraper-api/real-world/build_amazon_product_snapshot) This guide walks you through a real marketplace workflow with the Geonode Scraper API: **as if we are building it together**. You want a **small competitive product snapshot** on Amazon.com: what shows up for a keyword, at what price, with ratings. You **already know the site**. You do **not** need the Geonode Search API to find Amazon. By the end, you will understand **when to use Extract vs Batch** on a hard, JavaScript-heavy marketplace, why residential geo targeting matters, and how to keep a demo bounded when pages block or fail. Companion example code lives in `geonode-scraper-examples/amazon-product-snapshot/`. This is a **small teaching demo** (one search page, capped ASINs). Amazon frequently blocks or challenges automated traffic. Follow Amazon’s terms and applicable law. This guide does **not** teach bypassing captchas, logging into accounts, or large-scale scraping. For production hardness, contact support. ## Use case Imagine you are on a pricing, marketplace, or assortment team. You need a short list of products for one keyword: * ASIN and product URL * Title * Price (USD when present) * Star rating and review count when present Starting point: a **known Amazon.com search URL**, for example: ```text https://www.amazon.com/s?k=logitech+mx+master ``` ## What we are going to build A master JSON snapshot. Each product looks roughly like this (shortened): ```json { "asin": "B0FB21526X", "url": "https://www.amazon.com/dp/B0FB21526X", "title": "Logitech MX Master 3S Bluetooth Wireless Mouse...", "price": 89.99, "currency": "USD", "rating": 4.6, "reviews_count": 1200 } ``` | Metric | Demo target | | ---------- | -------------: | | SERP pages | 1 | | Max ASINs | 10 | | Proxy | US residential | | JS render | on | Master Amazon snapshot ## The plan (what you should expect) ```text 1. Extract → Amazon search (SERP) Markdown/HTML 2. Parse → ASINs (local; fallback list if SERP empty) 3. Batch → /dp/{ASIN} product pages 4. Parse → title, price, rating (local) 5. Merge → amazon_products.json ``` POST /v1/extract"] --> B["2. Parse ASINs
local"] end subgraph products["Product pages"] direction LR C["3. Batch PDPs
POST /v1/batch"] --> D["4. Parse
local"] D --> E["5. Merge"] E --> F["amazon_products.json"] end discovery --> C classDef api fill:#e8f5f3,stroke:#054D4D,color:#033333 classDef local fill:#f5f5f5,stroke:#666,color:#222 classDef result fill:#fff8e6,stroke:#b8860b,color:#333 class A,C api class B,D,E local class F result `} /> | API | Why we use it here | | -------------------------------------------------------------------------- | ------------------------------------------- | | [Extract](/docs/scraper-api/guides/extraction/01_understanding_extraction) | One known SERP URL; JS + residential + wait | | [Batch](/docs/scraper-api/guides/batch/00_understanding_batch) | Many `/dp/` URLs with the same settings | | Local parse | ASINs and product fields | **Geonode Search is not used.** You already know `amazon.com`. **Map is not used.** You are not inventorying a sitemap; you start from one search URL. **Crawl is not used.** You do not want an unbounded walk of Amazon. Compare with other real-world guides: | Guide | Difference | | ------------------------------------------------------------------------------------------------- | ---------------------------------------------- | | [Retail category price catalog](/docs/scraper-api/real-world/build_retail_category_price_catalog) | Sitemap + PLP pagination on a lighter retailer | | [Docs knowledge corpus](/docs/scraper-api/real-world/build_docs_knowledge_corpus) | Crawl public docs | | [Basic Auth headers](/docs/scraper-api/real-world/build_authenticated_basic_auth) | Authenticated Extract, not marketplace SERP | *** ## Before you start You need: * A Geonode API key * Python 3.9+, `requests`, `python-dotenv` ```bash export GEONODE_SCRAPER_API_KEY="YOUR_API_KEY" ``` Optional: ```bash export AMAZON_SEARCH_URL="https://www.amazon.com/s?k=logitech+mx+master" export AMAZON_MAX_ASINS="10" ``` *** ## Step 1: Extract the search results page ### What we need HTML/Markdown of one Amazon SERP so we can collect ASINs. ### What we will do `POST /v1/extract` with: * `render_js: true` * `proxy: { "country": "US", "type": "residential" }` * `wait_config` with `wait_until: "networkidle"` (Amazon is a heavy SPA) ### What you should expect A large Markdown/HTML blob with `/dp/{ASIN}` links, or a soft block / empty shell. Save raw output either way so you can debug. ```python import os import requests api_key = os.environ["GEONODE_SCRAPER_API_KEY"] url = os.environ.get( "AMAZON_SEARCH_URL", "https://www.amazon.com/s?k=logitech+mx+master", ) response = requests.post( "https://scraper.geonode.io/v1/extract", headers={ "X-Api-Key": api_key, "Content-Type": "application/json", }, json={ "url": url, "formats": ["markdown", "html"], "render_js": True, "processing_mode": "sync", "proxy": {"country": "US", "type": "residential"}, "wait_config": { "wait_until": "networkidle", "wait_timeout": 20000, }, }, ) # Save markdown / html for ASIN parsing ``` SERP extract ### Takeaway **Hard marketplaces need JS + residential geo + patience.** Retry on 429/503/504. If the SERP is blocked, do not pretend Batch will invent ASINs. *** ## Step 2: Parse ASINs locally ### What we need A capped list of product URLs: `https://www.amazon.com/dp/{ASIN}`. ### What we will do Regex over SERP Markdown/HTML for `/dp/`, `/gp/product/`, and similar. Dedupe. Cap at `AMAZON_MAX_ASINS` (demo default **10**). If parsing finds nothing, the companion repo falls back to a checked-in `fallback_asins.json` so later Batch/parse steps still teach the pipeline. ### What you should expect `asins.json` + `product_urls.json`. Parsed ASINs ### Takeaway **Discovery is local once Extract returns page content.** Keep the ASIN cap small for demos and cost control. *** ## Step 3: Batch product detail pages ### What we need Title/price/rating material from each `/dp/` URL. ### What we will do `POST /v1/batch` with the same `render_js`, US residential proxy, and wait settings. Poll until completed. Save per-ASIN Markdown and HTML. ### What you should expect Some URLs succeed, some fail, Amazon demos often show partial completion (for example 6/10). That is normal; parse what you have. ```python # urls = ["https://www.amazon.com/dp/B0…", ...] response = requests.post( "https://scraper.geonode.io/v1/batch", headers={ "X-Api-Key": api_key, "Content-Type": "application/json", }, json={ "urls": urls, "ignore_invalid_urls": True, "formats": ["markdown", "html"], "render_js": True, "proxy": {"country": "US", "type": "residential"}, "wait_config": { "wait_until": "networkidle", "wait_timeout": 20000, }, }, ) job_id = response.json()["job_id"] # Poll GET /v1/batch/{job_id} until completed ``` Batch PDPs ### Takeaway **Batch reuses Extract settings across a URL list.** It does not guarantee every Amazon PDP returns price HTML. *** ## Step 4–5: Parse and merge ### What we need The business artifact: `amazon_products.json`. ### What we will do Locally parse title (including `og:title`), price patterns / `a-offscreen` price HTML, rating, and review counts. Merge into a slim master file with counts and optional SERP meta. ### What you should expect A spreadsheet-ready list. Some products may have `price: null` when the PDP shell loaded without a clear price node. Merged snapshot ### Takeaway Same pattern as other real-world guides: **Geonode fetches; your merge step is the deliverable.** *** ## Why this API mix worked | Tool | Role | | ----------- | ----------------------------------------- | | Extract | One SERP you already know | | Batch | Many PDPs, same browser/proxy config | | Map / Crawl | Wrong shape for a single keyword snapshot | ### When not to use this pattern * You only have known ASINs → skip SERP; start at Batch * You need unbounded site coverage → still not Crawl-on-Amazon for demos * You need account pages → not this guide (and often not feasible via simple headers) *** ## Cost and request usage On [request-based pricing](/docs/scraper-api/additional-resources/request_based_pricing): | Stage | Rough volume | Notes | | ------------- | -----------------------: | ----------------------------------- | | SERP Extract | 1 | JS + residential | | Batch PDPs | up to `AMAZON_MAX_ASINS` | Partial failures still consume work | | Parse / merge | 0 | Local | Keep the ASIN cap small while teaching. *** ## Limitations and good practice * Amazon may return captchas, empty shells, or timeouts, expect flaky runs. * Prefer first-party / permitted use cases; respect robots and terms. * Do not store or publish customer account data. * Multi-page SERP pagination (`page=`) is on you, same idea as retail PLP pagination in the [BBB catalog guide](/docs/scraper-api/real-world/build_retail_category_price_catalog). * Production monitoring usually needs retries, alerting, and support for harder setups. *** ## Recap 1. **Extract** one Amazon search URL with JS + US residential. 2. **Parse** ASINs (cap the list; keep a fallback for demos). 3. **Batch** `/dp/` pages with the same settings. 4. **Parse / merge** into `amazon_products.json`. That is the Amazon teaching story next to the lighter retail catalog: **same Extract → Batch shape, harder site, smaller demo, honest limits.** # Extract a Basic Auth Page with Authorization Headers (/docs/scraper-api/real-world/build_authenticated_basic_auth) This guide walks you through an **authenticated** Extract workflow with the Geonode Scraper API: **as if we are building it together**. Geonode does **not** open a login UI. You send site credentials in request `headers`. Here we use **HTTP Basic Auth** — the same pattern many staging sites and internal tools use. By the end you will understand: * `X-Api-Key` → authenticates you to **Geonode** * `headers.Authorization: Basic …` → authenticates the request to the **target site** Companion example code lives in `geonode-scraper-examples/basic-auth-protected-page/`. ## Use case You need Markdown from a URL that returns **401 Unauthorized** without credentials (staging wiki, internal tool, password-protected doc). For a **reproducible public demo**, this guide uses well-known Basic Auth practice URLs (not a production customer site): ```text https://the-internet.herokuapp.com/basic_auth ``` Public demo credentials: username `admin`, password `admin`. The **API pattern** is what you reuse on your own Basic Auth staging or internal pages. ## What we are going to build A small auth corpus plus before/after meta: ```json { "business_use_case": "authenticated_extract_with_basic_auth_headers", "auth": "headers.Authorization: Basic … — Geonode does not log in", "page_count": 2, "pages": [ { "url": "https://the-internet.herokuapp.com/basic_auth", "authenticated_ok": true, "excerpt": "Congratulations! You must have the proper credentials." } ] } ``` | Metric | Demo target | | ------- | -----------------------------: | | Auth | HTTP Basic (`admin` / `admin`) | | APIs | Extract + Batch | | Outcome | `auth_corpus.json` | Auth corpus sample ## The plan ```text 1. Probe → Extract WITHOUT Authorization 2. Extract → same URL WITH Authorization: Basic 3. Batch → more protected URLs, same header 4. Parse → title + excerpt (local) 5. Merge → auth_corpus.json ``` no Authorization"] --> B["2. Extract
Basic Auth"] B --> C["3. Batch
same header"] end subgraph output["Output"] direction LR D["4. Parse"] --> E["5. Merge"] E --> F["auth_corpus.json"] end auth_flow --> D classDef api fill:#e8f5f3,stroke:#054D4D,color:#033333 classDef local fill:#f5f5f5,stroke:#666,color:#222 classDef result fill:#fff8e6,stroke:#b8860b,color:#333 class A,B,C api class D,E local class F result `} /> | API | Why we use it here | | ---------------------------------------------------------------------------------- | --------------------------------------- | | [Extract](/docs/scraper-api/guides/extraction/01_understanding_extraction) | Clear before/after on one URL | | [Batch](/docs/scraper-api/guides/batch/00_understanding_batch) | Same Basic header on several URLs | | [Custom headers](/docs/scraper-api/guides/making-requests/07_using_custom_headers) | `Authorization` is a target-site header | Compare with other real-world guides (all **public**, no auth): | Guide | Auth | | ------------------------------------------------------------------------------------------------- | -------------------------- | | [B2B agency lead list](/docs/scraper-api/real-world/build_b2b_agency_lead_list) | None | | [Retail category price catalog](/docs/scraper-api/real-world/build_retail_category_price_catalog) | None | | [Docs knowledge corpus](/docs/scraper-api/real-world/build_docs_knowledge_corpus) | None | | **This guide** | **`Authorization: Basic`** | *** ## Before you start * A Geonode API key * Python 3.9+, `requests` ```bash export GEONODE_SCRAPER_API_KEY="YOUR_API_KEY" ``` Heavy SaaS apps that rely on fragile browser cookies often fight proxies. Basic Auth on a simple page is the clearest way to teach authenticated Extract. Apply the same `headers` pattern to **your** staging or internal Basic Auth URLs. *** ## Step 1: Probe without Authorization ### What we need Proof the page is protected. ### What we will do `POST /v1/extract` with **no** `headers.Authorization`. ### What you should expect No “Congratulations” success body (API may surface an error or empty/locked content). ```python import os import requests api_key = os.environ["GEONODE_SCRAPER_API_KEY"] url = "https://the-internet.herokuapp.com/basic_auth" response = requests.post( "https://scraper.geonode.io/v1/extract", headers={ "X-Api-Key": api_key, "Content-Type": "application/json", }, json={ "url": url, "formats": ["markdown"], "render_js": False, "processing_mode": "sync", }, ) # Should NOT contain the protected success copy ``` Anonymous probe ### Takeaway **If anonymous Extract already shows the secret page, you are not testing auth.** *** ## Step 2: Extract with Basic Auth ### What we need The same URL, authorized. ### What we will do Encode `username:password` as Base64 and send: ```text Authorization: Basic YWRtaW46YWRtaW4= ``` (`admin:admin` → `YWRtaW46YWRtaW4=`) Do **not** put the Geonode API key inside `headers`. ```python import base64 import os import requests api_key = os.environ["GEONODE_SCRAPER_API_KEY"] url = "https://the-internet.herokuapp.com/basic_auth" token = base64.b64encode(b"admin:admin").decode("ascii") response = requests.post( "https://scraper.geonode.io/v1/extract", headers={ "X-Api-Key": api_key, "Content-Type": "application/json", }, json={ "url": url, "formats": ["markdown"], "render_js": False, "processing_mode": "sync", "headers": { "Authorization": f"Basic {token}", }, }, ) # Expect Markdown containing "Congratulations" ``` Authenticated extract ### Takeaway **`headers.Authorization` is target-site auth.** Same field works for `Bearer` tokens when a site expects that instead of Basic. *** ## Step 3: Batch with the same header ### What we need Several protected URLs without rewriting Extract. ### What we will do `POST /v1/batch` with `urls` and the same `headers.Authorization`. Demo list: * `https://the-internet.herokuapp.com/basic_auth` * `https://httpbin.org/basic-auth/admin/admin` (same `admin` / `admin`) ### What you should expect One job, one Basic header, multiple Markdown results. Batch authenticated ### Takeaway **Batch reuses headers; it does not invent a login.** *** ## Step 4–5: Parse and merge Locally turn Markdown into `auth_corpus.json`, including optional before/after meta from probe vs authenticated Extract. Merged corpus *** ## Why this API mix worked | Tool | Role | | ------------------ | -------------------------------------------------------------------------------------------------------------- | | Extract | Prove Basic Auth before/after | | Batch | Same header, many URLs | | Cookie / SSO flows | Out of scope here — see FAQ on [custom headers vs full session UI](/docs/scraper-api/additional-resources/faq) | Use only URLs and credentials you are allowed to access. Practice-site credentials are public by design; production secrets stay in `.env`. *** ## Cost and request usage | Stage | Rough volume | | -------------- | -----------: | | Probe Extract | 1 | | Authed Extract | 1 | | Batch | N URLs | | Parse / merge | 0 (local) | `render_js` is off for these simple pages. *** ## Recap 1. **Probe** without `Authorization`. 2. **Extract** with `Authorization: Basic …`. 3. **Batch** more protected URLs with the same header. 4. **Parse / merge** into `auth_corpus.json`. That is the authenticated Extract story: **Geonode key for the API, Basic (or Bearer) for the site.** # Build a B2B Agency Lead List from a Directory (/docs/scraper-api/real-world/build_b2b_agency_lead_list) This guide walks you through a real B2B workflow with the Geonode Scraper API: **as if we are building it together**. You want a list of companies you can use for sales, partnerships, or research. You do **not** already have their websites. You only know the market (Berlin tech / creative agencies). By the end, you will understand **which API to reach for at each stage**, what a good result looks like, and why we skip Crawl (and avoid mapping every company website). ## Use case Imagine you are on a partnerships or outbound team. You need a spreadsheet-ready lead list with: * Company name and website * Location, size, founding year, hourly rate * Services / industries * Public email and phone from the company site Starting point: **market intent only**: for example, “Berlin software / agency companies.” Directory used in this walkthrough: ```text https://techbehemoths.com/companies/software-development/berlin ``` The same pattern works on other public directories. ## What we are going to build A master JSON lead file. Each company looks roughly like this (shortened): ```json { "slug": "andberlin", "name": "&Berlin Creative Agency", "website": "https://www.andberlin.co", "profile_url": "https://techbehemoths.com/company/andberlin", "hourly_rate": "$70-150/h", "founded": 2022, "employees": 10, "verified": true, "locations": ["Berlin"], "services": ["Branding", "Web Design", "Web Development"], "industries": [{ "name": "Business services", "percent": 10 }], "emails": ["example@example.com"], "phones": [], "has_contact": true } ``` In a full run against the Berlin software-development directory slice used for this example: | Metric | Approx. result | | -------------------------------- | -------------: | | Companies from directory pages | \~73 | | Homepages successfully extracted | \~69 | | With at least one email | \~60 | | With at least one phone | \~46 | Exact counts depend on pagination depth and site availability. Master lead file sample ## The plan (what you should expect) We will move through eight stages. At each stage we pick one Geonode product on purpose: ```text 1. Search → find a directory listing URL 2. Map → try link discovery (expect failure on JS listings: teaching moment) 3. Extract → render listing pages + pagination → profile URLs 4. Batch → extract all profile pages → HTML 5. Parse → firmographics + website (local) 6. Batch → extract all company homepages → HTML 7. Parse → emails / phones (local) 8. Merge → master lead file ``` /v1/search"] --> B["Listing URL"] B --> C["2. Map
/v1/map"] C -->|"JS grids → 0 links"| D["3. Extract
/v1/extract + JS"] end subgraph enrich["Enrich"] direction LR subgraph profiles["Profiles"] direction TB E["4. Batch
/v1/batch"] --> F["5. Parse
local"] end subgraph contacts["Contacts"] direction TB G["6. Batch sites
/v1/batch"] --> H["7. Parse
emails / phones"] end end subgraph output["Output"] direction LR I["8. Merge"] --> J["companies_master.json"] end discover --> enrich F --> G F --> I H --> I classDef api fill:#e8f5f3,stroke:#054D4D,color:#033333 classDef local fill:#f5f5f5,stroke:#666,color:#222 classDef result fill:#fff8e6,stroke:#b8860b,color:#333 class A,C,D,E,G api class F,H,I local class B,J result `} /> | API | Why we use it here | | -------------------------------------------------------------------------- | ------------------------------------------------- | | [Search](/docs/scraper-api/guides/search/01_search_overview) | You do not know the best listing URL yet | | [Map](/docs/scraper-api/guides/map/00_understanding_map) | Fast URL inventory check: often fails on JS grids | | [Extract](/docs/scraper-api/guides/extraction/01_understanding_extraction) | Render JS listing pages and collect profile links | | [Batch](/docs/scraper-api/guides/batch/00_understanding_batch) | Many known URLs (profiles, then homepages) | | Local parsing | Turn HTML into structured fields, then merge | **Crawl is not used.** After Extract you already know the profile URLs. After profile parse you already know the websites. You do not need a deep multi-page walk of one domain. *** ## Before you start You need: * A Geonode API key * Python 3.9 or later * The `requests` package ```bash export GEONODE_SCRAPER_API_KEY="YOUR_API_KEY" pip install requests ``` On Windows PowerShell: ```powershell $env:GEONODE_SCRAPER_API_KEY="YOUR_API_KEY" pip install requests ``` Never put your Geonode API key directly in your source code or commit it to your repository. Store it in an environment variable or `.env` file instead. *** ## Step 1: Find a directory with Search ### What we need A **listing URL**: a page that shows many company cards. We do not hardcode one on day one; we discover candidates with Search. ### What we will do Send a natural-language query that matches the market, for example `100 berlin startups`. ### What you should expect Search returns several sources (`url`, `title`, sometimes a snippet). Your job is to **pick one strong directory**. For this demo we keep a **mini list** from a single directory so the walkthrough stays clear. You can add more directories later with the same pipeline. ```python import os import requests api_key = os.environ["GEONODE_SCRAPER_API_KEY"] response = requests.post( "https://scraper.geonode.io/v1/search", headers={ "X-Api-Key": api_key, "Content-Type": "application/json", }, json={ "query": "100 berlin startups", "page": 1, "locale": "en", }, ) response.raise_for_status() results = response.json() print(results) ``` Search results showing a Berlin agency directory From the results, choose one directory listing. In this walkthrough we use TechBehemoths Berlin software-development companies: ```text https://techbehemoths.com/companies/software-development/berlin ``` Selected directory among search results Open that URL in a browser so you know what “success” looks like: company cards, profile links, pagination. Directory listing page in the browser ### Takeaway **Search answers “where do I start?”**: not “give me every company yet.” After this step you should have **one listing URL** saved and ready for the next APIs. *** ## Step 2: Try Map on the listing (expect JS limits) ### What we need A list of company **profile URLs** from the directory (for example `/company/andberlin`). ### What we will do Run Map on the listing. Map reads **sitemaps and static HTML links**. It does **not** execute JavaScript. We try Map **on purpose**. On modern directories the company grid is often rendered client-side. Seeing zero profile links teaches you when to switch tools. ### What you should expect On TechBehemoths-style JS listings, Map often returns **zero** `/company/...` links even though the browser shows dozens of companies. That is normal: not a broken API key. ```python listing_url = "https://techbehemoths.com/companies/software-development/berlin" response = requests.post( "https://scraper.geonode.io/v1/map", headers={ "X-Api-Key": api_key, "Content-Type": "application/json", }, json={ "url": listing_url, "search": "company", }, ) response.raise_for_status() print(response.json()) ``` Map job with zero or few profile links Use Map when you need a fast URL inventory for a site with a sitemap or static navigation: for example, finding `/kontakt`, `/impressum`, or `/about` on a single company domain later. Do not rely on Map alone for JS-heavy listing grids. ### Takeaway **Map failed here as a discovery path: and that is the lesson.** For JS directory grids, move to **Extract with `render_js`**. *** ## Step 3: Extract listing pages with JavaScript rendering ### What we need The profile URLs that Map could not see. ### What we will do Extract the listing with `render_js: true`, wait for the page to settle, then parse `/company/{slug}` links from the Markdown. Repeat for `?page=2`, `?page=3`, … until you have enough companies for your demo. ### What you should expect Each extracted page should contain many profile links in the Markdown. After a few pages you have a clean list of profile URLs ready for Batch. ```python response = requests.post( "https://scraper.geonode.io/v1/extract", headers={ "X-Api-Key": api_key, "Content-Type": "application/json", }, json={ "url": listing_url, "formats": ["markdown"], "render_js": True, "processing_mode": "sync", "proxy": {"country": "DE", "type": "residential"}, "wait_config": { "wait_until": "networkidle", "wait_timeout": 5000, }, }, ) response.raise_for_status() markdown = response.json()["data"]["markdown"] # Parse /company/{slug} links from markdown, then repeat for page=2, page=3, ... ``` Example profile URLs: ```text https://techbehemoths.com/company/andberlin https://techbehemoths.com/company/why-studio ... ``` Extract markdown containing company profile links ### Takeaway **Extract is how you discover URLs on JS listings.** Pagination is usually the same Extract call with `?page=N`: you do not need Crawl for this directory pattern. *** ## Step 4: Batch extract all profile pages ### What we need HTML for every company profile so we can read firmographics and the **website** field. ### What we will do Submit **one Batch job** with all known profile URLs instead of calling Extract in a loop. ### What you should expect You get a `job_id`. Poll `GET /v1/batch/{job_id}` until the job completes. Then save each result’s HTML. Concurrency follows your plan’s thread limits automatically: you do not manage worker threads in your app. ```python profile_urls = [ "https://techbehemoths.com/company/andberlin", "https://techbehemoths.com/company/why-studio", # ... more profile URLs ] response = requests.post( "https://scraper.geonode.io/v1/batch", headers={ "X-Api-Key": api_key, "Content-Type": "application/json", }, json={ "urls": profile_urls, "ignore_invalid_urls": True, "formats": ["html"], "render_js": True, "wait_config": {"wait_until": "domcontentloaded"}, }, ) response.raise_for_status() job = response.json() job_id = job["job_id"] # Poll GET /v1/batch/{job_id} until status is completed, then save each result HTML ``` Batch job for profile pages ### Takeaway **Many known URLs → Batch.** Extract was for discovery on listing pages. Batch is for volume on URLs you already have. *** ## Step 5: Parse profiles locally ### What we need Structured company records: name, website, size, services, and so on. ### What we will do Parse the saved profile HTML locally (no Geonode call). Directory profiles usually expose firmographics and a website: but rarely a public email. ### What you should expect JSON objects with fields such as: * `name`, `website`, `hourly_rate`, `founded`, `employees` * `locations`, `services`, `industries` * `profile_url` You should see a **website** and still see **no email** on most rows. That is why the next steps exist. Parsed company JSON object ### Takeaway **Directory pages give firmographics. Contact data usually lives on the company site.** Keep the website field: it is the bridge to enrichment. *** ## Step 6: Batch extract company homepages ### What we need Emails and phones. Homepages (and footers) are the fastest place to look first. ### What we will do Batch-extract each company’s **homepage** from the website field. We intentionally **do not Map every company domain**. ### Why we skip Map-on-every-website here * Map returns **URLs only**. You still need Extract/Batch afterward. * Agency sites often expose huge sitemaps (hundreds of URLs). Mapping each domain is slow and noisy. * Public emails/phones are often already on the homepage (`mailto:`, `tel:`, footer text). ### What you should expect A second Batch job over homepage URLs. Some sites fail or block; that is fine: those companies stay in the master file with empty contacts. ```python homepage_urls = [ "https://www.andberlin.co", "https://www.why.de", # ... websites from the profile parse step ] response = requests.post( "https://scraper.geonode.io/v1/batch", headers={ "X-Api-Key": api_key, "Content-Type": "application/json", }, json={ "urls": homepage_urls, "ignore_invalid_urls": True, "formats": ["html"], "render_js": True, "wait_config": {"wait_until": "domcontentloaded"}, }, ) response.raise_for_status() # Poll the batch job, then save homepage HTML per company ``` Batch job for company homepages ### Takeaway **Homepage-first enrichment is cheaper and faster than Map fan-out.** If a company still has no contact after this, *then* consider Map on that one domain for `/contact` / `/impressum`. *** ## Step 7: Parse contacts and merge the master file ### What we need The final lead list: firmographics + emails/phones in one record per company. ### What we will do 1. From each homepage HTML, extract `mailto:` / email-like strings and `tel:` / phone-like text. 2. Join profile fields + contacts on a stable key (slug or normalized website). 3. Light-clean placeholders (`example.com`, theme demos), dedupe phones, keep companies even when homepage Batch failed. ### What you should expect A contacts file with `emails` / `phones` arrays, then a master file where many rows have `has_contact: true`. Parsed emails and phones ### Takeaway **Scraper API gets you the pages. Local parse + merge makes the CRM-ready lead list.** Keep parsing simple and filter junk before outreach. *** ## Why this API mix worked ### Search was enough to start You only needed market intent. Search returns candidate sources; you pick one listing and continue. ### Map was useful as a check, not as the main discovery path | Situation | Better tool | | ------------------------------------------ | ---------------------------- | | JS directory grid / infinite scroll cards | **Extract** with `render_js` | | Static site or sitemap-heavy single domain | **Map** | | Deep walk of one large website | **Crawl** | | Many known URLs (profiles or homepages) | **Batch** | ### Extract + Batch covered the whole lead pipeline 1. Extract discovers profile URLs from rendered listing pages (+ pagination). 2. Batch pulls all profile HTML. 3. Batch pulls all homepage HTML. 4. Local code turns HTML into CRM-ready fields. That is usually enough for **directory → firmographics → homepage contacts**. ### When to add Map or Crawl later * **Map a company site** after the homepage has no email: discover `/contact`, `/kontakt`, `/impressum`, then Batch only those few URLs. * **Crawl** a single large corporate domain when you need broad content discovery beyond a short contact URL list. Do not start with Crawl across dozens of unrelated company domains for this use case. *** ## Cost and request usage On [request-based pricing](/docs/scraper-api/additional-resources/request_based_pricing), successful **page extractions** consume requests. Map and job-status polling do not count as page extractions the same way content extraction does. Approximate request shape for a run like this example: | Stage | Rough volume | Notes | | -------------------------- | -----------: | --------------------------------------------- | | Listing Extract | \~3 pages | page 1–3 with JS rendering | | Profile Batch | \~73 URLs | 1 request per successfully extracted profile | | Homepage Batch | \~69 URLs | 1 request per successfully extracted homepage | | **Total page extractions** | **\~145** | Order-of-magnitude for this demo depth | ### Cost contrast: Map every company website first If you Map \~70 company domains, you may discover hundreds of URLs per site. Batching those multiplies spend and runtime. For homepage-first contact enrichment, **Batch homepages first**, then Map selectively only for companies still missing contact data. Pricing and plan details can change: check: * [Request-Based Pricing](/docs/scraper-api/additional-resources/request_based_pricing) * [Unlimited Scraper API Pricing](/docs/scraper-api/additional-resources/unlimited_scraper_api_pricing) * [Choosing a Scraper API Plan](/docs/scraper-api/additional-resources/choosing_scraper_api_plan) Unlimited plans bill by **concurrency (threads)**, not a monthly request balance: useful when you run large Batches often. *** ## Limitations and good practice * Use publicly available pages and respect site terms / robots rules for your jurisdiction and use case. * Directory data can be incomplete or outdated; prefer the company website as the contact source of truth. * Homepage parsing will miss contacts that exist only behind forms, images, or Impressum-only pages: that is when selective Map + Batch helps. * Filter obvious placeholder emails before exporting to a CRM or outreach tool. * This guide produces a **lead list**, not a compliance or consent platform. Apply your own outreach and privacy policies. *** ## Recap 1. **Search** finds the directory. 2. **Map** may fail on JS listings: switch to **Extract**. 3. **Extract** + pagination collects profile URLs. 4. **Batch** extracts all profiles, then all homepages. 5. Local parse + merge produces the master B2B lead file. # Build a Docs Knowledge Corpus with Crawl (/docs/scraper-api/real-world/build_docs_knowledge_corpus) This guide walks you through a real documentation workflow with the Geonode Scraper API: **as if we are building it together**. You want a **knowledge corpus**: many docs pages as clean Markdown for RAG, internal search, or a support bot. You **already know the docs site**. You do **not** have a ready-made list of every guide URL. By the end, you will understand **when to use Crawl** instead of Map, Extract, or Batch, and how a single async job discovers linked pages and returns their content. Companion example code lives in `geonode-scraper-examples/geonode-docs-corpus/`. ## Use case Imagine you are on a docs, support, or AI team. You need: * Many documentation pages as Markdown * Stable URLs and titles for indexing * A bounded crawl (not the whole internet) Starting point: **a known docs root**, for example Geonode’s public docs (the same seed used in the Crawl API guides): ```text https://docs.geonode.com/docs/scraper-api ``` Seeding the Scraper API section (instead of only the docs homepage) walks nested guides under Extraction, Batch, Crawl, Map, and Search. ## What we are going to build A master JSON corpus. Each page looks roughly like this (shortened): ```json { "url": "https://docs.geonode.com/docs/scraper-api/guides/crawl/01_first-crawl", "title": "Your First Crawl", "section": "scraper-api", "path": "/docs/scraper-api/guides/crawl/01_first-crawl", "depth": 2, "excerpt": "In this guide, you'll create your first crawl job...", "markdown_len": 4200 } ``` | Metric | Demo target | | ----------------------- | -----------------------------------: | | Seed | Scraper API docs section | | Crawl `limit` / `depth` | 30 / 3 | | Outcome | `docs_corpus.json` with page records | {/* ![Master docs corpus sample](/images/scraper-api/real-world/docs-corpus/00-master-corpus.png) */} ## The plan (what you should expect) Fewer stages than a directory lead list or retail catalog, because **Crawl discovers and extracts in one job**: ```text 1. Crawl → seed docs → linked pages + markdown 2. Parse → title, section, excerpt (local) 3. Merge → docs_corpus.json ``` POST /v1/crawl"] --> B["Poll GET /v1/crawl/job_id"] B --> C["Page markdown files"] end subgraph output["Output"] direction LR D["2. Parse
local"] --> E["3. Merge"] E --> F["docs_corpus.json"] end crawl_job --> D classDef api fill:#e8f5f3,stroke:#054D4D,color:#033333 classDef local fill:#f5f5f5,stroke:#666,color:#222 classDef result fill:#fff8e6,stroke:#b8860b,color:#333 class A,B api class D,E local class C,F result `} /> | API | Why we use it here | | ------------------------------------------------------ | ----------------------------------------------------------- | | [Crawl](/docs/scraper-api/guides/crawl/01_first-crawl) | One seed, unknown linked tree, content + discovery together | | Local parsing | Turn Markdown into indexable fields | **Search is not used.** You already know the docs URL. **Map is not required.** Map would only return URLs; you would still Batch Extract. Crawl returns content. **Batch is not required.** You do not have a URL list until Crawl finishes, and by then pages are already extracted. Compare with other real-world guides: | Guide | Why Crawl was skipped | | ------------------------------------------------------------------------------------------------- | ---------------------------------------------------- | | [B2B agency lead list](/docs/scraper-api/real-world/build_b2b_agency_lead_list) | Directory + pagination gave profile URLs; then Batch | | [Retail category price catalog](/docs/scraper-api/real-world/build_retail_category_price_catalog) | Sitemap + PLP Extract gave product URLs; then Batch | *** ## Before you start You need: * A Geonode API key * Python 3.9 or later * The `requests` package ```bash export GEONODE_SCRAPER_API_KEY="YOUR_API_KEY" pip install requests ``` On Windows PowerShell: ```powershell $env:GEONODE_SCRAPER_API_KEY="YOUR_API_KEY" pip install requests ``` Never put your Geonode API key directly in your source code or commit it to your repository. Store it in an environment variable or `.env` file instead. *** ## Step 1: Crawl the docs section ### What we need Many docs pages with Markdown, without hand-picking every guide URL. ### What we will do Create one Crawl job from the Scraper API docs seed. Cap the walk with `limit` and `depth`. Keep `same_domain_only: true`. Docs are mostly static, so start with `render_js: false`. ### What you should expect A `202` response with `job_id`. Poll `GET /v1/crawl/{job_id}` until `status` is `completed`. Save each completed page’s Markdown. ```python import os import time import requests api_key = os.environ["GEONODE_SCRAPER_API_KEY"] base = "https://scraper.geonode.io" response = requests.post( f"{base}/v1/crawl", headers={ "X-Api-Key": api_key, "Content-Type": "application/json", }, json={ "url": "https://docs.geonode.com/docs/scraper-api", "formats": ["markdown"], "limit": 30, "depth": 3, "same_domain_only": True, "include_subdomains": False, "render_js": False, }, ) response.raise_for_status() job_id = response.json()["job_id"] while True: job = requests.get( f"{base}/v1/crawl/{job_id}", headers={"X-Api-Key": api_key}, ).json() print(job["status"], job.get("completed_pages"), "/", job.get("total_pages")) if job["status"] in {"completed", "failed", "cancelled"}: break time.sleep(5) # job["results"] → save data.markdown per page ``` Crawl job progress Crawling only `https://docs.geonode.com/` with shallow depth may stop at a handful of top-level hubs (Getting Started, Proxies, Scraper API). Seeding **`/docs/scraper-api`** walks nested Extraction, Batch, Crawl, Map, and Search guides, better for a RAG demo. ### Takeaway **Crawl answers “give me linked pages with content from this seed.”** Bound the job with `limit` and `depth` so demos stay cheap and predictable. *** ## Step 2: Parse pages locally ### What we need Structured records: URL, title, section, short excerpt. ### What we will do Parse saved Markdown locally (no Geonode call). Prefer the first `#` heading or front-matter `title:` as the page title. Derive `section` from the URL path (`/docs/scraper-api/...` → `scraper-api`). ### What you should expect A `pages.json` with one object per crawled URL, plus section counts. Parsed docs page JSON ### Takeaway **Scraper API gets you the pages. Local parse makes the corpus indexable.** Keep excerpts short for demos; store full Markdown files separately if you need RAG chunks later. *** ## Step 3: Merge the master corpus ### What we need The demo deliverable: one slim file a search or AI pipeline can load. ### What we will do Drop failed pages and internal file names. Keep `url`, `title`, `section`, `path`, `depth`, `excerpt`, `markdown_len`. Add `page_count` and section histogram. ### What you should expect `docs_corpus.json`, the knowledge-corpus equivalent of a lead list or product catalog master file. Master docs corpus file ### Takeaway Same pattern as other real-world guides: **Geonode fetches; your merge step is the business artifact.** *** ## Why this API mix worked ### Crawl was required You had **one seed** and needed **many unknown linked pages with content**. That is Crawl’s job. ### Map would have been incomplete | Tool | What you get | | ------- | ------------------------------------------------ | | Map | URLs only, still need Extract/Batch | | Extract | One page per call, you write the link walker | | Batch | Needs a URL list you already have | | Crawl | Discover + extract, async, capped by limit/depth | ### When not to use Crawl * You already have every URL → **Batch** * You only need a link inventory → **Map** * You need one page → **Extract** * You do not know which site → **Search** first, then Crawl that site *** ## Cost and request usage On [request-based pricing](/docs/scraper-api/additional-resources/request_based_pricing), successful page extractions in a crawl consume requests (roughly one per completed page). | Stage | Rough volume | Notes | | ------------- | ------------------: | ------------------ | | Crawl | up to `limit` pages | Demo uses limit 30 | | Parse / merge | 0 | Local only | Keep `limit` small while teaching. Raise it when you need a fuller corpus. *** ## Limitations and good practice * Prefer first-party or permitted docs sites for demos. * Docs trees change; re-run Crawl when the corpus must stay fresh. * Shallow depth on a marketing homepage may miss nested guides, seed the section you care about. * Excerpts are not chunking strategy; production RAG usually splits Markdown further. *** ## Recap 1. **Crawl** the docs section with `limit` / `depth`. 2. **Parse** Markdown into titles, sections, excerpts. 3. **Merge** into `docs_corpus.json` for RAG or search. That is the Crawl-first story: **one seed, unknown URLs, content included**, the gap left by the B2B lead list and retail catalog examples. # Build a Retail Category Price Catalog from a Public Sitemap (/docs/scraper-api/real-world/build_retail_category_price_catalog) This guide walks you through a real e-commerce workflow with the Geonode Scraper API: **as if we are building it together**. You want a **competitive catalog snapshot**: what a retailer sells in one category, at what price, on promo or not. You **already know the website**. You do **not** need Search to find it. By the end, you will understand **which API to reach for at each stage**, why Map on a homepage is not the same as the XML sitemap you open in a browser, and why Extract does not paginate for you. Companion example code lives in `geonode-scraper-examples/bedbathandbeyond-catalog/`. ## Use case Imagine you are on a category, marketplace, or pricing team. You need a spreadsheet-ready product list with: * Product title, brand, SKU / item number * Current price (USD) and sale flag * Category breadcrumb * Product URL and retailer product id Starting point: **a known retailer**, not a search query. Example: Bed Bath & Beyond. Sitemap and category used in this walkthrough: ```text https://www.bedbathandbeyond.com/sitemap.xml https://www.bedbathandbeyond.com/sitemap/ctaxonomy/ctaxonomy.xml https://www.bedbathandbeyond.com/c/towels/bath-towels?t=18652 ``` The same pattern works on other public retailers that publish a sitemap index and `/c/...` category pages (product listing pages, or **PLPs**). Each SKU lives on a product detail page (**PDP**), often ending in `/product.html`. ## What we are going to build A master JSON catalog file. Each product looks roughly like this (shortened): ```json { "product_id": "33411469", "sku": "37850619", "url": "https://www.bedbathandbeyond.com/Bedding-Bath/.../33411469/product.html", "title": "American Soft Linen 100% Cotton Turkish Bath Towels...", "brand": "American Soft Linen", "price": 55.49, "list_price": null, "on_sale": true, "promo": "Labor Day Sale", "currency": "USD", "category": ["Bedding & Bath", "Bath Linens", "Towels", "Bath Towels"], "availability": "unknown" } ``` In a demo run against **two PLP pages** of Bath Towels: | Metric | Approx. result | | --------------------------------- | -------------: | | Category PLPs in taxonomy sitemap | \~1430 | | Product URLs from 2 listing pages | \~36 | | PDP HTML files saved | \~34 | | With a parsed price | \~30 | | Price range (USD) | \~$33–$120 | Exact counts depend on pagination depth, Batch failures, and site availability. Master catalog sample ## The plan (what you should expect) We move through six stages. At each stage we pick one Geonode product on purpose: ```text 1. Map → sitemap.xml (index of sub-sitemaps) 2. Map → ctaxonomy.xml (all category PLP URLs) 3. Extract → one PLP + pagination → product URLs 4. Batch → all PDPs → HTML 5. Parse → title, price, SKU, sale (local) 6. Merge → products_master.json ``` /v1/map"] --> B["Sub-sitemap URLs"] B --> C["2. Map ctaxonomy.xml
/v1/map"] C --> D["Category PLPs /c/..."] end subgraph products["Collect products"] direction LR E["3. Extract PLP
/v1/extract + JS + page=N"] --> F["Product URLs"] F --> G["4. Batch PDPs
/v1/batch"] end subgraph output["Output"] direction LR H["5. Parse HTML
local"] --> I["6. Merge"] I --> J["products_master.json"] end discover --> products G --> H classDef api fill:#e8f5f3,stroke:#054D4D,color:#033333 classDef local fill:#f5f5f5,stroke:#666,color:#222 classDef result fill:#fff8e6,stroke:#b8860b,color:#333 class A,C,E,G api class H,I local class B,D,F,J result `} /> | API | Why we use it here | | -------------------------------------------------------------------------- | --------------------------------------------------------------------- | | [Map](/docs/scraper-api/guides/map/00_understanding_map) | Fast URL inventory from sitemap XML and taxonomy | | [Extract](/docs/scraper-api/guides/extraction/01_understanding_extraction) | Render a JS category grid and parse product links; one call = one URL | | [Batch](/docs/scraper-api/guides/batch/00_understanding_batch) | Many known PDP URLs in one job | | Local parsing | Turn HTML into price/SKU fields, then merge | **Search is not used.** You already know the retailer. **Crawl is not required for this demo.** After Extract you already have product URLs. Crawl is the right tool later if you want a BFS walk of one category with a page `limit` instead of looping `?page=`. *** ## Before you start You need: * A Geonode API key * Python 3.9 or later * The `requests` package ```bash export GEONODE_SCRAPER_API_KEY="YOUR_API_KEY" pip install requests ``` On Windows PowerShell: ```powershell $env:GEONODE_SCRAPER_API_KEY="YOUR_API_KEY" pip install requests ``` Never put your Geonode API key directly in your source code or commit it to your repository. Store it in an environment variable or `.env` file instead. *** ## Step 1: Map the sitemap index (not the homepage) ### What we need The retailer’s **sitemap index**: a list of sub-sitemap XML files (taxonomy, refinements, keyword pages, and so on). ### What we will do Open the sitemap in a browser so you know what “success” looks like: ```text https://www.bedbathandbeyond.com/sitemap.xml ``` You should see a `` with `` entries such as: ```text https://www.bedbathandbeyond.com/sitemap/ctaxonomy/ctaxonomy.xml https://www.bedbathandbeyond.com/sitemap/ctaxonomy/refinements.xml https://www.bedbathandbeyond.com/sitemap/keyword-search-pages/keyword-search-pages.xml ``` Sitemap index in the browser Then call Map on **that same URL** (`sitemap.xml`), not on `https://www.bedbathandbeyond.com`. ### Why homepage Map looks different Map means: **discover links from this seed**. It is not “pretty-print the XML file Chrome showed you.” | Seed you Map | Typical result | | ------------- | ------------------------------------------------------ | | Homepage `/` | Mix of site links, often **PDPs** (`.../product.html`) | | `sitemap.xml` | Inventory of URLs Map can see from that document | If Map on `sitemap.xml` does not return the `.xml` children you see in the browser, **Extract the sitemap with `render_js: false`** and parse `` tags locally. That is the same lesson as JS directories: Map first, Extract when the document content is what you need. ```python import os import requests api_key = os.environ["GEONODE_SCRAPER_API_KEY"] response = requests.post( "https://scraper.geonode.io/v1/map", headers={ "X-Api-Key": api_key, "Content-Type": "application/json", }, json={ "url": "https://www.bedbathandbeyond.com/sitemap.xml", "include_subdomains": True, }, ) response.raise_for_status() print(response.json()) ``` Map output vs browser sitemap ### Takeaway **Map answers “what URLs can we inventory from this seed?”** For a sitemap index, seed **`sitemap.xml`**. Save `ctaxonomy.xml` as the next seed. *** ## Step 2: Map the category taxonomy (all PLPs) ### What we need Every **category listing URL** (`/c/...`) so you can pick one demo category (Bath Towels). ### What we will do Map (or Extract-parse `` if Map times out) on: ```text https://www.bedbathandbeyond.com/sitemap/ctaxonomy/ctaxonomy.xml ``` In the browser this file is a `` of category pages, for example: ```text https://www.bedbathandbeyond.com/c/towels/bath-towels?t=18652 ``` Category taxonomy XML in the browser Taxonomy sitemaps can be huge. Map may return **HTTP 408** (Map request timed out during URL discovery). That is a documented Map error, not a bad API key. Retry, or Extract the XML and parse every `` that contains `/c/`. ```python response = requests.post( "https://scraper.geonode.io/v1/map", headers={ "X-Api-Key": api_key, "Content-Type": "application/json", }, json={ "url": "https://www.bedbathandbeyond.com/sitemap/ctaxonomy/ctaxonomy.xml", "include_subdomains": True, "ignore_query_parameters": False, }, ) response.raise_for_status() links = response.json().get("links") or [] # Keep URLs whose path contains /c/ ``` In this example run, taxonomy discovery produced **about 1430** category PLPs. We pick one: ```text https://www.bedbathandbeyond.com/c/towels/bath-towels?t=18652 ``` Mapped category PLP list ### Takeaway **Second Map is the category tree**, not product SKUs. PDPs usually live on listing pages, not in this taxonomy file. *** ## Step 3: Extract the category PLP (pagination is your loop) ### What we need **Product URLs** (`.../product.html`) from the Bath Towels grid. ### What we will do 1. Extract **page 1** with JavaScript rendering and `wait_config`. 2. Parse the highest `page=` in the Markdown (the UI can show the last page on page 1). 3. Call Extract again for `?page=2`, `?page=3`, … yourself. ### Extract does not auto-paginate One Extract request = **one URL**. There is no “follow all pages” flag. That is the same pattern as directory listings in the [B2B lead list guide](/docs/scraper-api/real-world/build_b2b_agency_lead_list). Bed Bath & Beyond uses a query parameter: ```text https://www.bedbathandbeyond.com/c/towels/bath-towels?t=18652&page=33 ``` PLP pagination in the browser ### wait\_config (when the grid is JS) From [Waiting for Dynamic Content](/docs/scraper-api/guides/making-requests/04_waiting_for_dynamic_content), wait order is: ```text wait_until → wait_for → wait_timeout → extract ``` For a product grid: ```python response = requests.post( "https://scraper.geonode.io/v1/extract", headers={ "X-Api-Key": api_key, "Content-Type": "application/json", }, json={ "url": "https://www.bedbathandbeyond.com/c/towels/bath-towels?t=18652", "formats": ["markdown"], "render_js": True, "processing_mode": "sync", "proxy": {"country": "US", "type": "residential"}, "wait_config": { "wait_until": "networkidle", "wait_for": 'a[href*="product.html"]', "wait_timeout": 5000, }, }, ) response.raise_for_status() markdown = response.json()["data"]["markdown"] # Parse product.html links; parse max page=N; repeat Extract for page=2... ``` `429` means **request throttled or work concurrency limit reached** ([Error Handling](/docs/scraper-api/troubleshooting/error-handling)). Slow down, honor `Retry-After` if present, and retry with backoff. Sync Extract with `render_js` uses workers. This demo extracted **2 PLP pages** and collected **\~36** product URLs. {/* ![Extract markdown containing product links](/images/scraper-api/real-world/retail-catalog/03-extract-product-links.png) */} ### Takeaway **Extract discovers PDPs on JS category pages.** You implement pagination. Crawl is optional if you prefer one async job with a `limit` instead of a page loop. *** ## Step 4: Batch extract all product pages ### What we need HTML for every PDP so we can read price, SKU, and title. ### What we will do Submit **one Batch job** with all known product URLs instead of calling Extract in a loop. ### What you should expect You get a `job_id`. Poll `GET /v1/batch/{job_id}` until the job completes. Save each result’s HTML (use the numeric product id in the filename so files are not all named `product.html`). Some URLs may fail; keep going. In this run: **36** submitted, **34** HTML files, **2** failed. ```python product_urls = [ "https://www.bedbathandbeyond.com/Bedding-Bath/.../33411469/product.html", # ... URLs from step 3 ] response = requests.post( "https://scraper.geonode.io/v1/batch", headers={ "X-Api-Key": api_key, "Content-Type": "application/json", }, json={ "urls": product_urls, "ignore_invalid_urls": True, "formats": ["html"], "render_js": True, "proxy": {"country": "US", "type": "residential"}, "wait_config": { "wait_until": "networkidle", "wait_timeout": 3000, }, }, ) response.raise_for_status() job_id = response.json()["job_id"] # Poll GET /v1/batch/{job_id} until completed, then save each result HTML ``` Batch job for product pages ### Takeaway **Many known URLs → Batch.** Extract was for PLP discovery. Batch is for volume on PDPs you already have. *** ## Step 5: Parse product HTML locally ### What we need Structured product rows: title, price, SKU, sale flag, category. ### What we will do Parse saved HTML locally (no Geonode call). Retailer PDPs often embed a compact analytics object (in this example, `ensighten.items` with `price`, `sku`, `productId`, `productName`) plus breadcrumbs. ### What you should expect JSON objects with fields such as: * `product_id`, `sku`, `url`, `title`, `brand` * `price`, `list_price`, `on_sale`, `currency` * `category` (breadcrumb) You may **not** get a reliable in-stock flag from HTML alone (add-to-cart vs “out of stock” copy is noisy). Treat `availability` as best-effort. {/* ![Parsed product JSON object](/images/scraper-api/real-world/retail-catalog/05-parsed-product-json.png) */} ### Takeaway **Scraper API gets you the pages. Local parse makes the pricing spreadsheet.** Prefer structured blobs in the HTML over scraping the entire 3000-line document. *** ## Step 6: Merge the master catalog ### What we need The final deliverable: one slim file a category or pricing person can use. ### What we will do Drop parse internals (`source_file`). Keep business fields. Add stats (`with_price`, `on_sale_count`, min/max/avg USD). Optionally list Batch-failed URLs so you can retry later. ### What you should expect `products_master.json` with a `products` array and summary metrics. Master catalog file ### Takeaway **Same idea as a B2B master lead file:** Geonode fetches pages; your merge step is the business artifact. *** ## Why this API mix worked ### You already knew the site (no Search) Search is for **market intent** (“Berlin agencies”). Here the seed is a **known domain + public sitemap**. ### Map was the right first tool (unlike the JS directory) | Situation | Better tool | | ------------------------------------- | ----------------------------------------------------------- | | Sitemap index / taxonomy XML | **Map** (Extract `` if Map 408 or misses XML children) | | JS category grid / pagination | **Extract** with `render_js` + `wait_config` | | Many known PDP URLs | **Batch** | | Deep walk of one site with a page cap | **Crawl** (optional; not used in this demo) | ### Extract + Batch covered listing → SKUs → HTML 1. Map builds the category list from taxonomy. 2. Extract discovers product URLs from rendered PLPs (+ your page loop). 3. Batch pulls all PDP HTML. 4. Local code turns HTML into a catalog JSON. That is usually enough for **sitemap → category → products → prices**. ### When to add Crawl later Use [Crawl](/docs/scraper-api/guides/crawl/01_first-crawl) from one PLP if you want async BFS with `depth` and `limit` instead of writing a `?page=` loop. Keep `limit` small for demos. Do not Crawl the entire retailer from the homepage for this use case. *** ## Cost and request usage On [request-based pricing](/docs/scraper-api/additional-resources/request_based_pricing), successful **page extractions** consume requests. Map and job-status polling do not count as page extractions the same way content extraction does. Approximate request shape for a run like this example (2 PLP pages): | Stage | Rough volume | Notes | | -------------------------- | -----------: | ---------------------------------------------- | | Map sitemap + taxonomy | 2 Map calls | Plus Extract on XML only if Map misses `` | | PLP Extract | 2 pages | page 1–2 with JS rendering (`--max-pages 2`) | | PDP Batch | \~36 URLs | 1 request per successfully extracted product | | **Total page extractions** | **\~38+** | Order-of-magnitude for this demo depth | If you Extract **all** Bath Towels pages (for example 30+), listing Extract grows linearly. Batch then grows with unique PDPs. ### Cost contrast: Crawl the whole site A homepage Crawl with a high `limit` can fetch hundreds of mixed URLs (guides, search pages, PDPs). For one category, **Extract that PLP + Batch those PDPs** stays bounded. Pricing and plan details can change: check: * [Request-Based Pricing](/docs/scraper-api/additional-resources/request_based_pricing) * [Unlimited Scraper API Pricing](/docs/scraper-api/additional-resources/unlimited_scraper_api_pricing) * [Choosing a Scraper API Plan](/docs/scraper-api/additional-resources/choosing_scraper_api_plan) Unlimited plans bill by **concurrency (threads)**, not a monthly request balance: useful when you run large Batches often. *** ## Limitations and good practice * Use publicly available pages and respect site terms / robots rules for your jurisdiction and use case. * Prices and promo flags change; this is a **point-in-time snapshot**, not a live feed unless you re-run Batch. * `availability` parsed from HTML can be wrong; confirm in the UI if stock is a business-critical field. * Pagination windows in the UI may not show the true last page on page 1; still parse `page=` links and cap `--max-pages` for demos. * This guide produces a **catalog file**, not a commercial scraping product. Apply your own policies. *** ## Recap 1. **Map** `sitemap.xml` to find sub-sitemaps (Extract XML if Map misses ``). 2. **Map** `ctaxonomy.xml` to list category PLPs (handle Map **408** on huge files). 3. **Extract** one PLP with JS + `wait_config`; loop `?page=` yourself. 4. **Batch** all PDPs to HTML. 5. Local parse + merge produces the master price catalog. Companion: B2B directories use **Search → Map-often-fails → Extract listings**. Retailers with a sitemap use **Map twice → Extract PLP → Batch PDPs**. # Scraper API Request Parameters (/docs/scraper-api/snippets/requests-parameters) import Link from "next/link";
  • Output Formats
  • JavaScript Rendering
  • Waiting for Dynamic Content
  • Processing Modes
  • Proxy and Geo-Targeting
  • Using Custom Headers
# Code Examples (/docs/scraper-api/troubleshooting/code-examples) These examples call `POST /v1/extract` in synchronous mode and print the extracted Markdown. Before running them, set your Scraper API base URL and API key. ```bash export SCRAPER_API_BASE_URL="https://scraper.geonode.io" export GEONODE_SCRAPER_API_KEY="YOUR_API_KEY" ``` ## cURL Use this example when you want to test the API from a terminal before writing application code. ```bash curl -X POST "$SCRAPER_API_BASE_URL/v1/extract" \ -H "X-Api-Key: $GEONODE_SCRAPER_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "url": "https://example.com", "formats": ["markdown"], "render_js": false, "processing_mode": "sync" }' ``` ## Python This example uses Python's standard library so you do not have to install any extra packages. ```python import json import os from urllib import request, error base_url = os.environ["SCRAPER_API_BASE_URL"] api_key = os.environ["GEONODE_SCRAPER_API_KEY"] payload = { "url": "https://example.com", "formats": ["markdown"], "render_js": False, "processing_mode": "sync", } req = request.Request( f"{base_url}/v1/extract", data=json.dumps(payload).encode("utf-8"), headers={ "X-Api-Key": api_key, "Content-Type": "application/json", }, method="POST", ) try: with request.urlopen(req, timeout=60) as response: result = json.loads(response.read().decode("utf-8")) print(result["data"]["markdown"]) except error.HTTPError as exc: body = exc.read().decode("utf-8") print(f"Request failed with HTTP {exc.code}: {body}") ``` ## Node.js This example uses the built-in `fetch` API available in current Node.js versions. ```javascript const baseUrl = process.env.SCRAPER_API_BASE_URL; const apiKey = process.env.GEONODE_SCRAPER_API_KEY; const response = await fetch(`${baseUrl}/v1/extract`, { method: "POST", headers: { "X-Api-Key": apiKey, "Content-Type": "application/json", }, body: JSON.stringify({ url: "https://example.com", formats: ["markdown"], render_js: false, processing_mode: "sync", }), }); const result = await response.json(); if (!response.ok) { throw new Error(`Request failed with HTTP ${response.status}: ${JSON.stringify(result)}`); } console.log(result.data.markdown); ``` # Error Handling (/docs/scraper-api/troubleshooting/error-handling) The Scraper API returns standard HTTP status codes and JSON error bodies. Handle both validation errors and extraction errors in your client, because a request can fail before extraction starts or after the API tries to process the target page. ## HTTP Status Codes | HTTP status | Meaning | Returned by | | ----------- | --------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------- | | `200` | Synchronous extraction, map request, job lookup, statistics request, webhook lookup, or health check succeeded. | Extract (sync), map, get job, get batch/crawl status, list jobs, statistics, webhook get/list, health | | `201` | Webhook subscription was created. | Create webhook | | `202` | Async extraction, batch, or crawl job was accepted, or a running job cancellation was accepted. | Extract (async), create batch, create crawl, cancel batch, cancel crawl | | `204` | Webhook subscription was deleted. | Delete webhook | | `400` | Invalid request. | Extract, map | | `401` | API key is missing or invalid. | All authenticated endpoints | | `402` | Payment required or insufficient request balance. | Extract, map, crawl | | `404` | Job, webhook, or other requested resource was not found. | Get job, get batch/crawl status, cancel batch/crawl, webhook get/update/delete/deliveries | | `408` | Map request timed out during URL discovery. | Map | | `409` | Batch or crawl job cannot be cancelled in its current state. | Cancel batch, cancel crawl | | `422` | Validation error or extraction failed. | Extract (body validation and extraction errors), create batch, create crawl, create/update webhook | | `429` | Request throttled or work concurrency limit reached. | Extract, create batch, create crawl, map | | `500` | Internal server error. | Extract, map, webhook create/get/update/delete/list/deliveries | | `502` | Billing service returned an upstream error. Retryable depends on the billing error. | Extract, map, create crawl | | `503` | Service or billing service temporarily unavailable. | Extract, map, create batch, create crawl | | `504` | Synchronous extraction timed out waiting for the browser worker. Retryable. | Extract (sync) | ## Validation Error Validation errors can return a `detail` array. This usually means the request body does not match the expected schema, such as an invalid URL or unsupported value. ```json { "detail": [ { "type": "value_error", "loc": ["body", "url"], "msg": "Value error, URL must contain a valid hostname.", "input": "not-a-url" } ] } ``` ## Extraction Error Extraction failures return an `error` object and `tokens_charged`. ```json { "error": { "code": "UNPROCESSABLE_CONTENT", "message": "The target page could not be extracted.", "retryable": false, "details": null }, "tokens_charged": 0 } ``` The response field is currently named `tokens_charged` in the API schema. In Scraper API docs and billing language, treat this value as the number of requests charged. ## Extraction Error Codes Extraction error codes include: * `RATE_LIMITED` * `WORK_CONCURRENCY_LIMITED` * `WORK_CONCURRENCY_LEASE_EXPIRED` * `TEMPORARY_BLOCK` * `NETWORK_ERROR` * `CAPTCHA_CHALLENGE` * `PERMANENT_BLOCK` * `INVALID_URL` * `UNPROCESSABLE_CONTENT` * `AUTH_REQUIRED` * `TIMEOUT` * `PAYMENT_REQUIRED` * `INTERNAL_ERROR` * `PROXY_ERROR` Use the `retryable` field to decide whether a retry may help. For retryable failures, use exponential backoff and avoid retrying in a tight loop. The OpenAPI schema includes `429` responses, but no public numeric rate limit is specified in the current contract. If you receive `429`, slow down the client and retry after a delay. # API Overview (/docs/scraper-api/v1) The Geonode Scraper API helps you extract, discover, and process web content without managing browsers, proxies, or scraping infrastructure. Send a URL and receive structured content as Markdown or HTML. The API also supports JavaScript rendering, geo-targeting, batch processing, website crawling, and webhook notifications. ## What You Can Do With the Scraper API, you can: * Extract content from webpages * Process multiple URLs in batch jobs * Crawl websites and discover pages * Receive webhook notifications when jobs complete * Use geo-targeted proxy routing * Extract content from JavaScript-powered websites * Retrieve links found on webpages ## Available APIs Choose the API that best matches your use case. | API | Use When | | ---------- | -------------------------------------------------------- | | Extraction | You already know the URL and want to extract its content | | Batch | You have multiple URLs that need to be processed | | Crawl | You want to discover and extract pages across a website | | Webhooks | You want to receive notifications when jobs complete | ## When to Use the Scraper API Use the Scraper API when you need: * Clean Markdown or HTML output * JavaScript rendering * Geo-targeted extraction * Managed proxy infrastructure * Batch processing * Website crawling * Asynchronous processing * Webhook notifications If you need complete control over browser automation, request handling, or custom scraping logic, consider using the Geonode Proxy API instead. ## Getting Started If you're new to the Scraper API, start with the Quick Start Guide. The Quick Start Guide covers: 1. Creating an API key 2. Sending your first request 3. Understanding responses 4. Exploring available APIs ## Next Steps Continue to the **Quick Start Guide** to make your first request and explore the Scraper API. # Reference (/docs/scraper-api/v1/reference) Use the API Reference when you already know what you want to build and need the exact endpoint for a specific operation. If you're new to the Scraper API, start with the Quick Start Guide. ## Base URL All API requests use the following base URL: ```text https://scraper.geonode.io ``` You can store it as an environment variable: ```bash title="request.sh" export SCRAPER_API_BASE_URL="https://scraper.geonode.io" ``` ## Authentication All Scraper API requests require an API key. Send your API key using the `X-Api-Key` request header. ```http X-Api-Key: YOUR_API_KEY ``` You can store the API key as an environment variable: ```bash title="request.sh" export GEONODE_SCRAPER_API_KEY="YOUR_API_KEY" ``` Keep your API key private and never expose it in frontend applications, public repositories, screenshots, or logs. ## API Categories The Scraper API is organized into the following categories. | Category | Purpose | | ---------- | ------------------------------------------- | | Extraction | Extract content from webpages | | Batch | Process multiple URLs in a single job | | Crawl | Discover and extract pages across a website | | Map | Discover URLs from a website | | Statistics | Retrieve usage statistics | | Webhooks | Receive notifications when jobs complete | | System | Service health and status | ## Endpoints ### System | Method | Endpoint | Description | | ------ | --------- | ----------------------- | | `GET` | `/health` | Check API health status | ### Extraction | Method | Endpoint | Description | | ------ | ---------------------- | ------------------------------ | | `POST` | `/v1/extract` | Extract content from a webpage | | `GET` | `/v1/extract/{job_id}` | Retrieve an extraction job | | `GET` | `/v1/extract/jobs` | List extraction jobs | ### Batch | Method | Endpoint | Description | | -------- | -------------------- | -------------------- | | `POST` | `/v1/batch` | Start a batch job | | `GET` | `/v1/batch/{job_id}` | Retrieve a batch job | | `DELETE` | `/v1/batch/{job_id}` | Cancel a batch job | ### Map | Method | Endpoint | Description | | ------ | --------- | ---------------------------- | | `POST` | `/v1/map` | Discover URLs from a website | ### Crawl | Method | Endpoint | Description | | -------- | -------------------- | -------------------- | | `POST` | `/v1/crawl` | Start a crawl job | | `GET` | `/v1/crawl/{job_id}` | Retrieve a crawl job | | `DELETE` | `/v1/crawl/{job_id}` | Cancel a crawl job | ### Statistics | Method | Endpoint | Description | | ------ | ---------------- | ------------------------- | | `GET` | `/v1/statistics` | Retrieve usage statistics | ### Webhooks | Method | Endpoint | Description | | -------- | ----------------------------------------- | ----------------------- | | `POST` | `/v1/webhooks` | Create a webhook | | `GET` | `/v1/webhooks` | List webhooks | | `GET` | `/v1/webhooks/{webhook_id}` | Retrieve a webhook | | `PATCH` | `/v1/webhooks/{webhook_id}` | Update a webhook | | `DELETE` | `/v1/webhooks/{webhook_id}` | Delete a webhook | | `POST` | `/v1/webhooks/{webhook_id}/rotate-secret` | Rotate a webhook secret | | `GET` | `/v1/webhooks/{webhook_id}/deliveries` | List webhook deliveries | ## Next Steps Choose the endpoint category that matches your use case and open its guides or endpoint reference pages. # Some Apps Are Not Using the Proxy (/docs/proxies/additional-resources/fixing-common-issues/apps-bypassing-proxy) Some applications may bypass proxy configurations, causing inconsistent behavior or incorrect IP routing.\ This guide explains why that happens and how to fix it. *** ## Common Symptoms | Issue | Description | | --------------------------- | -------------------------------------------------------------------------------- | | **App ignores proxy** | The app connects directly to the internet instead of routing through your proxy. | | **Inconsistent IP results** | IP changes work in the browser but not in the app. | | **Authentication errors** | The app doesn’t support proxy login credentials. | *** ## Steps to Fix ### 1. Use a Proxy-Compatible Browser Some built-in browsers (like Edge WebView or internal app browsers) ignore system proxy settings.\ ✅ Try using **Google Chrome**, **Mozilla Firefox**, or **Brave**, which fully support proxies. *** ### 2. Check App Proxy Settings Certain apps have their **own proxy configuration** separate from the system.\ 🛠 Open the app’s **Network** or **Connection** settings and enter your proxy details manually. *** ### 3. Confirm Proxy Authentication Type If your proxy requires **username/password authentication**, make sure the app supports it.\ Some mobile or desktop apps only support IP whitelisting instead. *** ### 4. Use a VPN as an Alternative If the app completely ignores proxy settings, use a **VPN** service instead.\ VPNs tunnel all traffic system-wide, ensuring every app is routed through the same network path. *** ## FAQs Some apps do not support proxy connections or require manual configuration. {" "} Only by using a VPN or a system-level proxy manager that intercepts all connections. Usually no — browsers like Chrome and Firefox respect system or manual proxy settings. *** ## Summary * Some apps ignore system proxy settings by design. * Always check if the app has its own proxy configuration field. * Use Chrome or Firefox for consistent proxy testing. * If nothing works, switch to a VPN or a proxy manager for full traffic control. # Frequent Authentication Pop-Ups (/docs/proxies/additional-resources/fixing-common-issues/authentication-popups) If you keep getting authentication pop-ups while using a proxy — like this: Authentication Pop-Ups — it usually means there’s an issue with authentication or IP whitelisting.\ Follow these steps to fix it. *** ## Troubleshooting Steps ### 1. Whitelist Your IP Address If your IP is not whitelisted, the proxy will repeatedly ask for credentials.\ 👉 [How to Whitelist My IP](/docs/proxies/getting-started/setup_and_configuration/advance-configuration/whitelist-ip) *** ### 2. Check Your Login Credentials Ensure that you’re using the correct **username** and **password** for your Geonode proxy.\ Incorrect or expired credentials will trigger the authentication window every time. *** ### 3. Use an Authentication-Supported App Some browsers or apps do not handle proxy authentication correctly.\ ✅ Try using **Google Chrome**, **Mozilla Firefox**, or another modern browser that supports proxy logins. *** ## FAQs Your IP may not be whitelisted, or you might be using an incompatible browser. No, authentication is required for secure proxy usage.\ To avoid pop-ups, whitelist your IP so the proxy doesn’t ask for credentials. *** ## Summary * Repeated pop-ups usually mean your IP isn’t whitelisted. * Double-check your proxy username and password. * Use a browser that supports proxy authentication. * Whitelisting your IP eliminates most repeated login prompts. # Proxy Connection Keeps Dropping (/docs/proxies/additional-resources/fixing-common-issues/connection-dropping) If your proxy connection keeps disconnecting or timing out, this guide will help you identify the cause and fix it quickly. *** ## Troubleshooting Steps ### 1. Ensure Your Internet Connection Is Stable Unstable or weak internet can cause frequent proxy disconnections.\ ✅ Try switching to a wired connection or move closer to your Wi-Fi router. *** ### 2. Reconfigure Proxy Settings Incorrect or outdated proxy settings may cause the connection to drop.\ 🛠 Go to your **Network or Wi-Fi Settings**, remove the existing proxy configuration, and re-enter your **Geonode proxy details**. *** ### 3. Restart Your Router If you’re on a home or shared Wi-Fi network, restart your router.\ This helps clear temporary network conflicts that might interrupt proxy communication. *** ### 4. Test With Another App or Browser Try connecting your proxy in a different browser or app.\ If the issue doesn’t repeat, it may be caused by the original app’s network configuration. *** ## FAQs The most common causes are unstable internet connections or incorrect proxy configuration. {" "} Sometimes — switching to a different port can bypass temporary routing issues. Not necessarily. Check if your connection is stable and test the same proxy on another device before assuming downtime. *** ## Summary * Check and stabilize your internet connection. * Re-enter your proxy credentials and configuration. * Restart your router to clear temporary issues. * Test on a different app or browser to isolate the problem. # Proxy Not Changing My IP (/docs/proxies/additional-resources/fixing-common-issues/ip-not-changing) If your IP address doesn’t change after setting up a proxy, it’s usually due to app configuration or proxy setup issues.\ Follow these steps to verify and fix the problem. *** ## Troubleshooting Steps ### 1. Check Your IP on a Verification Site Visit [ShowMyIP](https://www.showmyip.com/) or [ip-api.com](https://ip-api.com/) to confirm whether your IP has actually changed.\ Sometimes, caching or DNS delays can make it look like your IP is the same. *** ### 2. Restart Your Browser or App After updating proxy settings, restart your browser or app to apply the configuration.\ ⚙️ Some apps only load proxy settings at startup — a restart ensures they take effect. *** ### 3. Ensure Your App Supports Proxies Certain applications bypass system proxy settings entirely.\ ✅ Try testing your proxy using **Google Chrome**, **Firefox**, or another proxy-compatible browser. *** ### 4. Confirm Manual Configuration Double-check that your proxy details were entered correctly under **Manual Proxy Setup** (IP, Port, Username, and Password).\ A missing field or typo can prevent the proxy from activating. *** ### 5. Test Another Proxy Type If your IP still doesn’t change, switch between **HTTP** and **SOCKS5** proxies to see if the app or service supports one better than the other. *** ## FAQs Some apps bypass system proxy settings. Test using a proxy-compatible browser like Chrome or Firefox. {" "} Yes — some websites detect your previous session via cookies or cache, even after the IP changes. Yes. If a VPN is active, it overrides proxy routing. Disable VPNs before testing proxy connections. *** ## Summary * Verify your IP with an external site like ShowMyIP or ip-api.com. * Restart your app to apply new proxy settings. * Ensure the proxy is configured correctly and supported by your application. * Disable VPNs or conflicting network tools before testing. # No Save Option in Proxy Settings (/docs/proxies/additional-resources/fixing-common-issues/no-save-option) If your device doesn’t show a **“Save”** button when setting up a proxy, don’t worry — some systems handle proxy saving automatically.\ Follow the steps below to confirm your settings are applied correctly. *** ## Troubleshooting Steps ### 1. Exit Settings After Configuration On many devices, proxy changes are **auto-saved** as soon as you leave the settings page.\ Try exiting the menu normally — the configuration is likely already active. *** ### 2. Restart Your Device If changes don’t take effect immediately, restart your device.\ A quick reboot often forces the system to apply pending network configurations. *** ### 3. Try Another Network Some Wi-Fi networks or administrators **block proxy changes** for security reasons.\ Connect to a different network and try setting up the proxy again. *** ## FAQs Some devices automatically save proxy configurations when you exit the settings menu. {" "} Usually no — but restarting can help apply the settings if the proxy doesn’t activate immediately. Yes. Most operating systems save proxy settings silently once you leave the configuration screen. *** ## Summary * Many systems auto-save proxy settings when you close the menu. * Restarting your device applies the new configuration if it doesn’t activate right away. * If settings still don’t apply, try switching to a different Wi-Fi network. # Pages Not Loading After Proxy Setup (/docs/proxies/additional-resources/fixing-common-issues/pages-not-loading) If web pages fail to load or load slowly after setting up a proxy, this guide will help you identify and fix the issue. *** ## Troubleshooting Steps ### 1. Check Proxy Server Status Make sure your proxy is **active and running** on the [Geonode Dashboard](https://app.geonode.com/).\ If it’s inactive or expired, your connection requests won’t go through. *** ### 2. Disable and Re-enable the Proxy Sometimes, reapplying the configuration helps.\ Go to your **Wi-Fi or Network Settings**, disable the proxy, save changes, and then enable it again. *** ### 3. Clear Browser Cache Cached or outdated data can conflict with new proxy settings.\ 🧹 Clear your browser cache and cookies, then restart the browser or app. *** ### 4. Try a Different Network Some Wi-Fi networks — especially in offices or public places — block proxy connections via firewalls.\ ✅ Test your proxy setup using a different Wi-Fi or mobile hotspot. *** ### 5. Test the Proxy Visit [ShowMyIP](https://www.showmyip.com/) or [ip-api.com](https://ip-api.com/) to confirm whether your proxy is active and your IP address has changed. *** ## FAQs Some websites block certain proxy servers or IP ranges. Try switching to another proxy location or type (HTTP/SOCKS5). {" "} Visit [ShowMyIP](https://www.showmyip.com/) or [ip-api.com](https://ip-api.com/) to confirm your proxy IP and location. Yes. Cached DNS and cookies can store old network data that conflicts with new proxy settings. *** ## Summary * Verify your proxy is active and properly configured. * Reapply proxy settings if pages aren’t loading. * Clear your browser cache and restart the app. * Test with another network or proxy location if the issue persists. # Proxy Not Working (/docs/proxies/additional-resources/fixing-common-issues/proxy-not-working) If your proxy isn’t working or failing to connect, follow these steps to diagnose and fix the issue. *** ## Troubleshooting Steps ### 1. Check Proxy Details Make sure you’ve entered the **correct Proxy IP and Port** from your [Geonode Dashboard](https://app.geonode.com/).\ ⚠️ A small typo in the IP or port number can prevent the connection entirely. *** ### 2. Restart Your Device After configuring the proxy, **restart your device** to ensure the settings are properly applied. *** ### 3. Switch Networks Try switching between **Wi-Fi and mobile data**.\ Some networks — especially corporate or public ones — block proxy traffic for security reasons. *** ### 4. Verify Proxy Compatibility Not all applications support proxies equally.\ ✅ Test your proxy with a browser like **Google Chrome** or **Firefox** to confirm it works as expected.\ If it does, the issue might be app-specific. *** ### 5. Check Proxy Authentication If your proxy requires credentials, double-check your **username** and **password**.\ Incorrect authentication can prevent the proxy from connecting. *** ## FAQs Ensure that you entered the correct Proxy IP, Port, Username, and Password. Also, confirm that the proxy is active in your Geonode Dashboard. {" "} Yes. In most cases, you’ll need to manually enter proxy details in your device’s or browser’s network settings. Some apps bypass system proxy settings. Try using a proxy-compatible browser or configure the proxy directly within the app. *** ## Summary * Double-check Proxy IP, Port, and authentication credentials. * Restart your device to apply settings. * Try a different network or browser to test compatibility. * If the proxy works elsewhere, the issue is likely app-specific. # Troubleshooting Proxy Issues on Android (/docs/proxies/additional-resources/fixing-common-issues/troubleshooting-common-issue) Setting up a proxy on Android can surface the same issues covered in other troubleshooting guides (no connection, pages not loading, repeated auth prompts, etc.). To avoid duplication, this page highlights only Android-specific checks and points you to the detailed articles for each issue. *** ## Quick Android Checks * Ensure the proxy is set under **Wi‑Fi → Advanced → Proxy → Manual** for the network you’re using. * Some Android apps bypass system proxies; test with **Chrome** or **Firefox** first. * Toggle Wi‑Fi off/on or reboot the device after changing proxy settings. * If using IP whitelist, confirm your current IP is added in the dashboard. *** ## Issue Index (Android) * **Proxy not working** — follow the main guide and re-check Wi‑Fi manual proxy config on Android: [Proxy Not Working](/docs/proxies/additional-resources/fixing-common-issues/proxy-not-working) * **Pages not loading** — see the primary steps for slow/no load; on Android also re-apply the proxy on the current Wi‑Fi: [Pages Not Loading](/docs/proxies/additional-resources/fixing-common-issues/pages-not-loading) * **Frequent authentication pop-ups** — confirm IP whitelist or credentials, and use a browser that supports auth prompts: [Authentication Popups](/docs/proxies/additional-resources/fixing-common-issues/authentication-popups) * **IP not changing** — verify the active network has the proxy set and test with a browser: [IP Not Changing](/docs/proxies/additional-resources/fixing-common-issues/ip-not-changing) * **No Save option in proxy settings** — many Android builds auto-save when you back out of settings: [No Save Option](/docs/proxies/additional-resources/fixing-common-issues/no-save-option) * **Connection keeps dropping** — apply the general stability steps and re-enter proxy details on your Wi‑Fi: [Connection Dropping](/docs/proxies/additional-resources/fixing-common-issues/connection-dropping) * **Apps bypassing proxy** — some apps ignore system proxies; use proxy-aware browsers or app-level proxy fields: [Apps Bypassing Proxy](/docs/proxies/additional-resources/fixing-common-issues/apps-bypassing-proxy) *** ## Conclusion Most Android proxy issues are resolved by confirming Wi‑Fi manual proxy settings, using a proxy-aware app, and applying the detailed steps in the linked articles above. If problems persist, visit the [Geonode Documentation](/) or contact **Geonode Support**. # Overview (/docs/proxies/api-reference/geo-targeting/geo-targeting-options) Geo-targeting is one of the most powerful features of the Geonode Proxy API. It allows you to route your proxy requests through specific geographic locations, giving you precise control over where your traffic appears to originate from. This is essential for location-specific testing, content access, market research, and compliance with regional requirements. ## What is Geo-Targeting? Geo-targeting enables you to specify the geographic location of the IP address that will be used for your proxy requests. Instead of getting a random IP from anywhere in the world, you can target: * **Countries**: Route traffic through specific countries (e.g., United States, United Kingdom, Germany) * **States/Regions**: Narrow down to specific states or provinces within a country * **Cities**: Target specific cities for even more precise location control * **ISPs/ASNs**: Route through specific Internet Service Providers or Autonomous System Numbers ## Why Use Geo-Targeting? Geo-targeting allows you to route your proxy requests through specific geographic locations, giving you control over where your traffic appears to originate from. ## Targeting Levels Geonode supports multiple levels of geo-targeting, from broad to highly specific: ### Country-Level Targeting The broadest level of targeting. Simply append `-country-` to your username to route traffic through a specific country. This is ideal when you need traffic from a particular country but don't need more specific location control. **Example**: `username-country-US` routes traffic through the United States. ### State-Level Targeting For countries with states or provinces, you can target specific regions. This is useful when you need traffic from a particular state but don't need city-level precision. **Example**: `username-country-US-state-california` routes traffic through California. You cannot target both state and city at the same time. Choose either state-level or city-level targeting for your requests. ### City-Level Targeting The most precise geographic targeting option. Target specific cities within a country for maximum location accuracy. **Example**: `username-country-US-city-newyork` routes traffic through New York City. ### ISP/ASN-Level Targeting For advanced use cases, you can target specific Internet Service Providers or Autonomous System Numbers. This is useful when you need traffic from a particular ISP or network infrastructure. **Example**: `username-type-residential-country-US-asn-12345` routes traffic through a specific ASN in the United States. ## Available Endpoints This section provides endpoints for different geo-targeting options: * **[Perform Country Targeting](/docs/proxies/api-reference/geo-targeting/get-country)**: Route traffic through specific countries * **[Perform State Targeting](/docs/proxies/api-reference/geo-targeting/get-state)**: Target specific states or regions * **[Perform City Targeting](/docs/proxies/api-reference/geo-targeting/get-city)**: Target specific cities * **[Perform ISP/ASN Targeting](/docs/proxies/api-reference/geo-targeting/get-isp)**: Route through specific ISPs or ASNs ## Finding Available Locations Before you can target a location, you need to know what locations are available. Use the [Retrieve Available Geo-locations](/docs/proxies/api-reference/available-geo-locations) endpoint to get a comprehensive list of: * Available countries and their codes * Cities within each country * States/regions within each country * ISPs and ASNs available in each location ## Best Practices Before targeting a location, verify that it's available using the available geo-locations endpoint. Choose the appropriate level of targeting for your needs—use country-level if you don't need more specific location control. Use ISO 3166-1 alpha-2 country codes (e.g., `US`, `GB`, `DE`) for country targeting. City and state names should match the exact format provided in the available locations list. # Perform City Targeting (/docs/proxies/api-reference/geo-targeting/get-city) Route your proxy requests through a specific city by including both the country code and city name in your username string. You cannot target both state and city at the same time. Choose either state-level or city-level targeting. Append `-city-` after `-country-` to target a specific city. For example: `username-country-US-city-newyork` targets New York City in the United States. ## Request ```bash curl --request GET \ -x "http://proxy.geonode.io:" \ --user "-country--city-:" \ --url "http://ip-api.com/json" ``` ## Response ### 200 Success Returns detailed geolocation information about the IP address assigned to your proxy connection. #### Response Fields | Field | Type | Description | | --------------- | ------ | ------------------------------------------- | | `status` | string | The status of the request (e.g., "success") | | `continent` | string | The name of the continent | | `continentCode` | string | The two-letter continent code | | `country` | string | The full name of the country | | `countryCode` | string | The two-letter country code | | `region` | string | The region code | | `regionName` | string | The full name of the region | | `city` | string | The name of the city | | `district` | string | The district name, if available | | `zip` | string | The postal code associated with the IP | | `query` | string | The queried IP address | #### Example Response ```json { "status": "success", "continent": "North America", "continentCode": "NA", "country": "United States", "countryCode": "US", "region": "AL", "regionName": "Alabama", "city": "Decatur", "district": "", "zip": 35601, "query": "68.191.141.86" } ``` # Perform Country Targeting (/docs/proxies/api-reference/geo-targeting/get-country) Route your proxy requests through a specific country by including the country code in your username string. Append `-country-` after your `` to target a specific country. For example: `username-country-US` targets the United States. ## Request ```bash curl --request GET \ -x "http://proxy.geonode.io:" \ --user "-country-:" \ --url "http://ip-api.com/json" ``` ## Response ### 200 Success Returns detailed geolocation information about the IP address assigned to your proxy connection. #### Response Fields | Field | Type | Description | | --------------- | ------- | ------------------------------------------------ | | `status` | string | The status of the request (e.g., "success") | | `continent` | string | The name of the continent | | `continentCode` | string | The two-letter continent code | | `country` | string | The full name of the country | | `countryCode` | string | The two-letter country code (ISO 3166-1 alpha-2) | | `region` | string | The region code | | `regionName` | string | The full name of the region | | `city` | string | The name of the city | | `district` | string | The district name, if available | | `zip` | string | The postal code associated with the IP | | `lat` | number | The latitude coordinate | | `lon` | number | The longitude coordinate | | `timezone` | string | The timezone of the IP location | | `offset` | integer | The time offset in seconds from UTC | | `currency` | string | The currency code of the country | | `isp` | string | The name of the Internet Service Provider | | `org` | string | The organization that owns the IP address | | `as` | string | The Autonomous System (AS) number | | `query` | string | The queried IP address | #### Example Response ```json { "status": "success", "continent": "North America", "continentCode": "NA", "country": "Canada", "countryCode": "CA", "region": "QC", "regionName": "Quebec", "city": "Montreal", "district": "", "zip": "H2Y", "lat": 45.5088, "lon": -73.5878, "timezone": "America/Toronto", "offset": -18000, "currency": "CAD", "isp": "Bell Canada", "org": "Bell Canada", "as": "AS577 Bell Canada", "query": "207.134.47.124" } ``` # Perform ISP/ASN Targeting (/docs/proxies/api-reference/geo-targeting/get-isp) Route your proxy requests through a specific ISP or Autonomous System Number (ASN) by including the ASN number in your username string. The username format supports multiple targeting options: * **Country**: `-country-` - Specify the country * **ASN**: `-asn-` - Filter by ASN number * **IP Type**: `-type-` - Choose residential, datacenter, or mix Refer to our [Geo-Locations List](https://docs.geonode.com/docs/proxies/api-reference/available-geo-locations) to find available ASN numbers and locations. ## Request ```bash curl -x "http://proxy.geonode.io:" \ --user "-type--country--asn-:" \ --url "http://ip-api.com/json" ``` ## Response ### 200 Success Returns detailed geolocation information about the IP address assigned to your proxy connection. #### Response Fields | Field | Type | Description | | --------------- | ------- | ------------------------------------------------------ | | `status` | string | The status of the request (e.g., "success") | | `continent` | string | The continent where the IP is located | | `continentCode` | string | The continent code | | `country` | string | The country where the IP is registered | | `countryCode` | string | The country code in ISO 3166-1 alpha-2 format | | `region` | string | The regional subdivision (state/province) | | `regionName` | string | The full name of the region | | `city` | string | The city associated with the IP address | | `district` | string | The district or subdivision of the city | | `zip` | string | The postal or ZIP code of the location | | `lat` | number | Latitude coordinate of the location | | `lon` | number | Longitude coordinate of the location | | `timezone` | string | Time zone in which the IP is located | | `offset` | integer | Time offset from UTC in seconds | | `currency` | string | Local currency used in the country | | `isp` | string | The name of the Internet Service Provider (ISP) | | `org` | string | The name of the organization associated with the IP | | `as` | string | The Autonomous System (AS) number and name | | `asname` | string | The full Autonomous System (AS) name | | `mobile` | boolean | Indicates whether the IP is from a mobile network | | `proxy` | boolean | Indicates whether the IP is being used as a proxy | | `hosting` | boolean | Indicates whether the IP belongs to a hosting provider | | `query` | string | The IP address queried in the request | #### Example Response ```json { "status": "success", "continent": "Europe", "continentCode": "EU", "country": "Russia", "countryCode": "RU", "region": "VGG", "regionName": "Volgograd Oblast", "city": "Volgograd", "district": "", "zip": "", "lat": 48.5044, "lon": 44.5838, "timezone": "Europe/Volgograd", "offset": 14400, "currency": "RUB", "isp": "JSC ER-Telecom Holding Volgograd branch", "org": "JSC Columbia-Telecom", "as": "AS50543 JSC ER-Telecom Holding", "asname": "SARATOV-AS", "mobile": false, "proxy": false, "hosting": false, "query": "83.167.79.185" } ``` # Perform OS Targeting (/docs/proxies/api-reference/geo-targeting/get-os-targeting) Route your requests through proxy IPs associated with devices running a specific operating system by adding the `-os-` parameter to your proxy username. Append `-os-` to your proxy username to target devices running a specific operating system. For example, `-os-ios-` routes traffic through iOS devices. ## Supported Operating Systems The following operating systems are supported: | Operating System | Username Value | | ---------------- | -------------- | | Windows | `windows` | | Android | `android` | | iOS | `ios` | | macOS | `mac` | ## Request ```bash curl --request GET \ -x "http://proxy.geonode.io:10009" \ --user "geonode_username-os-ios-lifetime-30-session-randomIOS:YOUR_PASSWORD" \ --url "http://ip-api.com/json" ``` ### Example ```bash curl -x proxy.geonode.io:10009 \ -U geonode_username-os-ios-lifetime-30-session-randomIOS:YOUR_PASSWORD \ http://ip-api.com/json ``` In this example, the operating system target is specified in the proxy username: ```text geonode_username-os-ios-lifetime-30-session-randomIOS ``` The important part is: ```text -os-ios- ``` You can replace `ios` with any of the supported operating systems: ```text -os-windows- -os-android- -os-ios- -os-mac- ``` ## Response ### 200 Success Returns information about the IP address assigned to your proxy connection. #### Response Fields | Field | Type | Description | | ------------- | ------ | -------------------------------------------- | | `status` | string | Status of the request. | | `country` | string | Country of the proxy IP address. | | `countryCode` | string | Two-letter country code. | | `region` | string | Region or state code. | | `regionName` | string | Full name of the region or state. | | `city` | string | City associated with the proxy IP. | | `zip` | string | ZIP or postal code. | | `lat` | number | Latitude of the proxy IP. | | `lon` | number | Longitude of the proxy IP. | | `timezone` | string | Time zone of the proxy IP. | | `isp` | string | Internet service provider. | | `org` | string | Organization associated with the IP address. | | `as` | string | Autonomous System (AS) information. | | `query` | string | Public IP address returned by the proxy. | #### Example Response ```json { "status": "success", "country": "United States", "countryCode": "US", "region": "PA", "regionName": "Pennsylvania", "city": "Philadelphia", "zip": "19133", "lat": 39.9934, "lon": -75.1425, "timezone": "America/New_York", "isp": "T-Mobile USA, Inc.", "org": "T-Mobile USA, Inc.", "as": "AS21928 T-Mobile USA, Inc.", "query": "172.56.216.58" } ``` # Perform State Targeting (/docs/proxies/api-reference/geo-targeting/get-state) Route your proxy requests through a specific state or region by including both the country code and state name in your username string. You cannot target both state and city at the same time. Choose either state-level or city-level targeting. Append `-state-` after `-country-` to target a specific state. For example: `username-country-US-state-california` targets California in the United States. ## Request ```bash curl --request GET \ -x "http://proxy.geonode.io:" \ --user "-country--state-:" \ --url "http://ip-api.com/json" ``` ## Response ### 200 Success Returns detailed geolocation information about the IP address assigned to your proxy connection. #### Response Fields | Field | Type | Description | | --------------- | ------ | ------------------------------------------- | | `status` | string | The status of the request (e.g., "success") | | `continent` | string | The name of the continent | | `continentCode` | string | The two-letter continent code | | `country` | string | The full name of the country | | `countryCode` | string | The two-letter country code | | `region` | string | The region code | | `regionName` | string | The full name of the region | | `city` | string | The name of the city | | `district` | string | The district name, if available | | `zip` | string | The postal code associated with the IP | | `query` | string | The queried IP address | #### Example Response ```json { "status": "success", "continent": "North America", "continentCode": "NA", "country": "Canada", "countryCode": "CA", "region": "QC", "regionName": "Quebec", "city": "Granby", "district": "", "zip": "J2H", "query": "192.168.1.1" } ``` # Perform strict matching (/docs/proxies/api-reference/geo-targeting/perform-strict-matching) ### Overview Previously, when you requested an IP address from a specific geolocation and none were available, the system would silently fall back to a similar location in the same country. We are changing this behavior to make it explicit and give you more control: 1. **Strict Matching (Default)** * By default, if your requested geolocation is unavailable, you will receive an error indicating that no IP address is available (`proxy-error`). * This means the system will **not** automatically provide an alternative IP from another location. 2. **Fallback with Flag** * If you **want** to allow fallback (i.e., you want an IP from another location in the same country when your exact requested location is unavailable), you must explicitly enable it by setting the flag `-strict-off`. ### URL Flag Usage You can now include one of the following flags in your username string: * `-strict-on` * Forces strict matching. If there are no available IPs for the specified location, the system returns a `proxy-error`. * This is the **default** if you do **not** specify any strict flag. * `-strict-off` * Allows fallback. If the exact location is unavailable, the system will return another IP from the same country (if available). #### Syntax ``` geonode_-type--country--asn--strict-on: ``` or ``` geonode_-type--country--asn--strict-off: ``` > **Note:** > > * You can apply `-strict-on` or `-strict-off` to any combination of geo-targeting flags (e.g., `country`, `city`, `state`, `asn`), but we recommend always including a `-country-` in your query. > * The `asn` flag must be used **with** the `country` flag; otherwise, you will receive an error. ### Example Usage #### 1. Fallback Enabled (`-strict-off`) If you want to **allow** fallback to a different location in the same country when your requested location is unavailable: ```bash curl -x : \ -U "geonode_username-type-residential-country-fr-asn-3215-strict-off:your_password" \ --url "http://ip-api.com" ``` * Here, you requested a French (`country-fr`) IP that belongs to ASN `3215`. * With `-strict-off`, if no IP is available exactly for ASN 3215 in France, the system will return a **different** IP from France. #### 2. Strict Matching (`-strict-on`) If you want **strict** matching for the exact country/ASN combination: ```bash curl -x : \ -U "geonode_username-type-residential-country-fr-asn-3215-strict-on:your_password" \ --url "http://ip-api.com" ``` * If no IP is available that exactly matches France + ASN 3215, the system will return: ```json { "proxy-error": "country-fr-asn-3215 target was not found" } ``` ### Error Cases 1. **No Exact Location Found (Strict Matching)** * **Error**: `{"proxy-error":"country-fr-asn-3215 target was not found"}` * Occurs when using `-strict-on` (or default strict matching) and the system cannot find an IP for that location. 2. **`asn` Without `country`** * **Error**: `{"proxy-error":"ASN should be used with country."}` * The `asn` flag must be paired with a specific country code. 3. **`-strict-on` or `-strict-off` Without Any Geo-Targeting Flag** * **Error**: If you use `-strict-on` or `-strict-off` **without** specifying any location or ASN. * You must include at least one geo-targeting parameter (e.g. `-country-xx`, `-asn-xxxxx`, etc.) for the request to make sense. ### Summary * **Default**: Strict matching is enforced (the system does **not** fall back to another location). * **Use `-strict-off`**: to allow fallback to a similar location within the same country. * **Always Pair `asn` With `country`**: `-asn-50543` must be accompanied by `-country-xx`. * **Error Handling**: You will receive JSON error responses if no IP is found under strict conditions or if flags are incorrectly combined. By explicitly controlling strict matching via `-strict-on` or `-strict-off`, you can now decide whether to **always** require a specific location or to **allow** fallback within the same country in case your requested location is unavailable. # Perform ASN Exclusion (/docs/proxies/api-reference/exclude-targeting/exclude-asn) Route your proxy requests while excluding specific ASNs (Autonomous System Numbers) from the routing. You cannot mix ASN exclusions with other exclusion types (country, city, or state) in the same request. Use only one exclusion type per request. Append `-not.asn-,` after your ``. * **Single ASN**: `-not.asn-31898` * **Multiple ASNs**: `-not.asn-31898,12345,67890` * **With country targeting**: `-country-us-not.asn-31898` ## Request ```bash curl --request GET \ -x "http://proxy.geonode.io:9000" \ --user "-country-us-not.asn-:" \ --url "http://ip-api.com/json" ``` ## Response ### 200 Success Returns detailed geolocation information about the IP address, excluding the specified ASNs. #### Response Fields | Field | Type | Description | | ------------- | ------ | ------------------------------------------- | | `status` | string | The status of the request (e.g., "success") | | `country` | string | The full name of the country | | `countryCode` | string | The two-letter country code | | `region` | string | The region code | | `regionName` | string | The full name of the region | | `city` | string | The name of the city | | `zip` | string | The postal code associated with the IP | | `lat` | number | The latitude coordinate | | `lon` | number | The longitude coordinate | | `timezone` | string | The timezone of the IP location | | `isp` | string | The name of the Internet Service Provider | | `org` | string | The organization that owns the IP address | | `as` | string | The Autonomous System (AS) number | | `query` | string | The queried IP address | #### Example Response ```json { "status": "success", "country": "United States", "countryCode": "US", "region": "MA", "regionName": "Massachusetts", "city": "Springfield", "zip": "01101", "lat": 42.0986, "lon": -72.5931, "timezone": "America/New_York", "isp": "RingSquared CC", "org": "", "as": "AS7849 RingSquared CC", "query": "161.77.215.31" } ``` # Perform City Exclusion (/docs/proxies/api-reference/exclude-targeting/exclude-city) Route your proxy requests while excluding specific cities from the routing. You can only exclude cities. You cannot combine city exclusions with country, state, or ASN exclusions in the same request. Append `-not.city-,` after your ``. * **Single city**: `-not.city-charlotte` * **Multiple cities**: `-not.city-charlotte,newyork,houston` * **With country**: `-country-US-not.city-charlotte,newyork` ## Request ```bash curl --request GET \ -x "http://proxy.geonode.io:" \ --user "-country--not.city-,:" \ --url "http://ip-api.com/json" ``` ## Response ### 200 Success Returns detailed geolocation information about the IP address, excluding the specified cities. #### Response Fields | Field | Type | Description | | ------------- | ------ | ------------------------------------------- | | `status` | string | The status of the request (e.g., "success") | | `country` | string | The full name of the country | | `countryCode` | string | The two-letter ISO country code | | `region` | string | The region code | | `regionName` | string | The full name of the region | | `city` | string | The name of the city | | `zip` | string | The postal code associated with the IP | | `lat` | number | The latitude coordinate | | `lon` | number | The longitude coordinate | | `timezone` | string | The timezone of the IP location | | `isp` | string | The name of the Internet Service Provider | | `org` | string | The organization that owns the IP address | | `as` | string | The Autonomous System (AS) number | | `query` | string | The queried IP address | #### Example Response ```json { "status": "success", "country": "United States", "countryCode": "US", "region": "NC", "regionName": "North Carolina", "city": "Charlotte", "zip": "28202", "lat": 35.2327, "lon": -80.8461, "timezone": "America/New_York", "isp": "FiberPower LLC", "org": "FiberPower LLC", "as": "AS214483 FiberPower LLC", "query": "38.13.166.129" } ``` # Perform Country Exclusion (/docs/proxies/api-reference/exclude-targeting/exclude-country) Route your proxy requests while excluding specific countries from the routing. You can only exclude countries. You cannot combine country exclusions with city, state, or ASN exclusions in the same request. Append `-not.country-,` after your ``. * **Single country**: `-not.country-US` * **Multiple countries**: `-not.country-US,CA,MX` ## Request ```bash curl --request GET \ -x "http://proxy.geonode.io:" \ --user "-not.country-,:" \ --url "http://ip-api.com/json" ``` ## Response ### 200 Success Returns detailed geolocation information about the IP address, excluding the specified countries. #### Response Fields | Field | Type | Description | | ------------- | ------ | ------------------------------------------- | | `status` | string | The status of the request (e.g., "success") | | `country` | string | The full name of the country | | `countryCode` | string | The two-letter ISO country code | | `region` | string | The region code | | `regionName` | string | The full name of the region | | `city` | string | The name of the city | | `zip` | string | The postal code associated with the IP | | `lat` | number | The latitude coordinate | | `lon` | number | The longitude coordinate | | `timezone` | string | The timezone of the IP location | | `isp` | string | The name of the Internet Service Provider | | `org` | string | The organization that owns the IP address | | `as` | string | The Autonomous System (AS) number | | `query` | string | The queried IP address | #### Example Response ```json { "status": "success", "country": "Mexico", "countryCode": "MX", "region": "AGU", "regionName": "Aguascalientes", "city": "Aguascalientes", "zip": "20326", "lat": 21.9419, "lon": -102.2756, "timezone": "America/Mexico_City", "isp": "Uninet S.A. de C.V.", "org": "UNINET", "as": "AS8151 UNINET", "query": "187.232.239.178" } ``` # Perform State Exclusion (/docs/proxies/api-reference/exclude-targeting/exclude-state) Route your proxy requests while excluding specific states or regions from the routing. You can only exclude states. You cannot combine state exclusions with city, country, or ASN exclusions in the same request. Append `-not.state-,` after your ``. * **Single state**: `-not.state-california` * **Multiple states**: `-not.state-california,newyork,texas` * **With country**: `-country-US-not.state-california,newyork` ## Request ```bash curl --request GET \ -x "http://proxy.geonode.io:" \ --user "-country--not.state-,:" \ --url "http://ip-api.com/json" ``` ## Response ### 200 Success Returns detailed geolocation information about the IP address, excluding the specified states. #### Response Fields | Field | Type | Description | | ------------- | ------ | ------------------------------------------- | | `status` | string | The status of the request (e.g., "success") | | `country` | string | The full name of the country | | `countryCode` | string | The two-letter ISO country code | | `region` | string | The region code | | `regionName` | string | The full name of the region | | `city` | string | The name of the city | | `zip` | string | The postal code associated with the IP | | `lat` | number | The latitude coordinate | | `lon` | number | The longitude coordinate | | `timezone` | string | The timezone of the IP location | | `isp` | string | The name of the Internet Service Provider | | `org` | string | The organization that owns the IP address | | `as` | string | The Autonomous System (AS) number | | `query` | string | The queried IP address | #### Example Response ```json { "status": "success", "country": "United States", "countryCode": "US", "region": "NY", "regionName": "New York", "city": "Queens", "zip": "11436", "lat": 40.6744, "lon": -73.8016, "timezone": "America/New_York", "isp": "Charter Communications", "org": "Spectrum", "as": "AS12271 Charter Communications Inc", "query": "72.227.174.55" } ``` # Remove Whitelisted IPs (/docs/proxies/api-reference/whitelisting-ip/delete) You can remove up to 10 IP addresses per request. Remove IP addresses from your whitelist. Once removed, these IPs will no longer have automatic access to your proxy services. ## Request ```bash curl -X DELETE "https://app-api.geonode.com/api/configuration/active/whitelisted-ips" \ -H "Authorization: Basic base64(username:password)" \ -H "Content-Type: application/json" \ -d '{"ips":[{"ip":"161.142.148.121"},{"ip":"161.142.148.147"}]}' ``` ### Request Body Parameters | Field | Type | Required | Description | | ---------- | ------ | -------- | --------------------------------------- | | `ips` | array | Yes | List of IP objects to remove | | `ips[].ip` | string | Yes | The IP address to remove from whitelist | ### Example Request Body ```json { "ips": [ { "ip": "161.142.148.121" }, { "ip": "161.142.148.147" } ] } ``` ## Response ### 200 Success The IP addresses have been successfully removed from your whitelist. #### Response Fields | Field | Type | Description | | -------------------- | ------ | ------------------------------------------------------ | | `data` | array | A list of whitelisted IPs that have been removed | | `data[].ip` | string | The IP address that was removed from the whitelist | | `data[].description` | string | The user-defined label for the removed IP | | `data[]._id` | string | A unique identifier assigned to the removed IP | | `message` | object | Provides status details about the removal operation | | `message.title` | string | A short message indicating the status of the operation | | `message.body` | string | A detailed message about the removal operation | | `message.variant` | string | The status variant indicating success or failure | #### Example Response ```json { "data": [ { "ip": "161.142.148.141", "description": "mac-m1", "_id": "67a875f693afe58d40f2e93d" } ], "message": { "title": "Updated", "body": "Whitelisted IPs removed.", "variant": "success" } } ``` ### Error Responses #### 400 Bad Request Invalid input data provided. | Field | Type | Description | | ------- | ------ | ------------- | | `error` | string | Error message | ```json { "error": "Invalid input data." } ``` #### 401 Unauthorized Invalid API credentials provided. | Field | Type | Description | | ------- | ------ | ------------- | | `error` | string | Error message | ```json { "error": "Unauthorized - Invalid credentials." } ``` # Retrieve Whitelisted IPs (/docs/proxies/api-reference/whitelisting-ip/get) Retrieve all IP addresses currently on your whitelist. ## Request ```bash curl -X GET "https://app-api.geonode.com/api/configuration/active/whitelisted-ips" -u "username:apiKey" \ -H "Authorization: Basic base64(username:password)" ``` ## Response ### 200 Success Returns a list of all active whitelisted IP addresses. #### Response Fields | Field | Type | Description | | -------------------- | ------ | ---------------------------------------------- | | `data` | array | List of whitelisted IP addresses | | `data[].ip` | string | The IP address of the whitelisted entity | | `data[].description` | string | A brief description associated with the IP | | `data[]._id` | string | The unique identifier assigned to the IP entry | #### Example Response ```json { "data": [ { "ip": "161.142.148.150", "description": "updated-description", "_id": "67a0a5bbb412ddaf6f139b3b" } ] } ``` ### Error Responses #### 400 Bad Request Invalid parameters provided. | Field | Type | Description | | ------- | ------ | ------------- | | `error` | string | Error message | ```json { "error": "Invalid parameters provided." } ``` #### 401 Unauthorized Invalid API credentials provided. | Field | Type | Description | | ------- | ------ | ------------- | | `error` | string | Error message | ```json { "error": "Unauthorized - Invalid API key or credentials." } ``` #### 500 Internal Server Error An internal server error occurred. # Add Whitelisted IPs (/docs/proxies/api-reference/whitelisting-ip/post) You can add up to 10 IP addresses per request. Add IP addresses to your whitelist to allow them to access your Geonode proxy services without authentication. ## Request ```bash curl -X POST "https://app-api.geonode.com/api/configuration/active/whitelisted-ips" \ -H "Authorization: Basic base64(username:password)" \ -H "Content-Type: application/json" \ -d '{"ips": [{"ip": "161.142.148.140", "description": "mac-m1"}]}' ``` ### Request Body Parameters | Field | Type | Required | Description | | ------------------- | ------ | -------- | --------------------------------------------- | | `ips` | array | Yes | List of IP objects to add to the whitelist | | `ips[].ip` | string | Yes | The IP address being added to the whitelist | | `ips[].description` | string | No | A user-defined label or identifier for the IP | ### Example Request Body ```json { "ips": [ { "ip": "161.142.148.140", "description": "mac-m1" } ] } ``` ## Response ### 200 Success The IP addresses have been successfully added to your whitelist. #### Response Fields | Field | Type | Description | | -------------------- | ------ | -------------------------------------------------- | | `data` | array | A list of successfully whitelisted IP addresses | | `data[].ip` | string | The IP address that was whitelisted | | `data[].description` | string | A user-defined label or identifier for the IP | | `data[]._id` | string | A unique identifier assigned to the whitelisted IP | | `message` | object | Details about the status of the whitelist update | | `message.title` | string | A short title summarizing the update status | | `message.body` | string | A message detailing the outcome of the operation | | `message.variant` | string | The status variant indicating success or failure | #### Example Response ```json { "data": [ { "ip": "161.142.148.150", "description": "updated-description", "_id": "67a0a5bbb412ddaf6f139b3b" } ], "message": { "title": "Updated", "body": "Whitelisted IPs saved.", "variant": "success" } } ``` ### Error Responses #### 400 Bad Request Invalid input data provided. | Field | Type | Description | | ------- | ------ | ------------- | | `error` | string | Error message | ```json { "error": "Invalid input data." } ``` #### 401 Unauthorized Invalid API credentials provided. | Field | Type | Description | | ------- | ------ | ------------- | | `error` | string | Error message | ```json { "error": "Unauthorized - Invalid API key or credentials." } ``` # Update Whitelisted IP Description (/docs/proxies/api-reference/whitelisting-ip/put) Update the description or label associated with a whitelisted IP address. ## Request ```bash curl -X PUT "https://app-api.geonode.com/api/configuration/active/whitelisted-ips" \ -H "Authorization: Basic base64(username:password)" \ -H "Content-Type: application/json" \ -d '{"ip":"161.142.148.150","description":"updated-description"}' ``` ### Request Body Parameters | Field | Type | Required | Description | | ------------- | ------ | -------- | ------------------------------------------ | | `ip` | string | Yes | The IP address to update | | `description` | string | Yes | The new description for the whitelisted IP | ### Example Request Body ```json { "ip": "161.142.148.150", "description": "updated-description" } ``` ## Response ### 200 Success The IP description has been successfully updated. #### Response Fields | Field | Type | Description | | -------------------- | ------ | --------------------------------------------------- | | `data` | array | A list of updated whitelisted IPs | | `data[].ip` | string | The IP address that was updated | | `data[].description` | string | The updated description for the whitelisted IP | | `data[]._id` | string | A unique identifier assigned to the whitelisted IP | | `message` | object | Provides status details about the update operation | | `message.title` | string | A short message indicating the status of the update | | `message.body` | string | A detailed message about the update operation | | `message.variant` | string | The status variant indicating success or failure | #### Example Response ```json { "data": [ { "ip": "161.142.148.150", "description": "just update", "_id": "67a0a5bbb412ddaf6f139b3b" } ], "message": { "title": "Updated", "body": "Whitelisted IPs updated.", "variant": "success" } } ``` ### Error Responses #### 400 Bad Request Invalid input data provided. | Field | Type | Description | | ------- | ------ | ------------- | | `error` | string | Error message | ```json { "error": "Invalid input data." } ``` #### 401 Unauthorized Invalid API credentials provided. | Field | Type | Description | | ------- | ------ | ------------- | | `error` | string | Error message | ```json { "error": "Unauthorized - Invalid API key or credentials." } ``` # Overview (/docs/proxies/api-reference/whitelisting-ip/whitelisting-ips) IP whitelisting is a security feature that allows you to control which IP addresses can access your Geonode proxy services without requiring authentication. This is particularly useful for securing your proxy infrastructure and ensuring that only authorized IPs can connect to your account. ## What is IP Whitelisting? When you whitelist an IP address, you're essentially creating a trusted list of IPs that can bypass the standard authentication process. This is ideal for scenarios where: * You have a fixed server or application that always connects from the same IP * You want to enhance security by restricting access to specific IP addresses * You need to simplify authentication for automated systems * You want to prevent unauthorized access from unknown locations ## Key Features You can manage up to 10 IP addresses per request, making it easy to bulk update your whitelist. * **Add IPs**: Add up to 10 IP addresses at once with optional descriptions for easy identification * **List IPs**: Retrieve all currently whitelisted IP addresses for your account * **Update Descriptions**: Modify the labels associated with whitelisted IPs for better organization * **Remove IPs**: Remove IP addresses from your whitelist when they're no longer needed ## Use Cases IP whitelisting is useful when you need to allow specific IP addresses to access your proxy services without authentication. Common scenarios include automated scripts, server applications, and controlled access environments. ## Getting Started To start using IP whitelisting, you'll need to: 1. **Add IPs to your whitelist** using the [Add Whitelisted IPs](/docs/proxies/api-reference/whitelisting-ip/post) endpoint 2. **View your current whitelist** with the [Retrieve Whitelisted IPs](/docs/proxies/api-reference/whitelisting-ip/get) endpoint 3. **Update IP descriptions** as needed using the [Update Whitelisted IP Description](/docs/proxies/api-reference/whitelisting-ip/put) endpoint 4. **Remove IPs** when they're no longer needed via the [Remove Whitelisted IPs](/docs/proxies/api-reference/whitelisting-ip/delete) endpoint Once an IP is whitelisted, it can access your proxy services without authentication. Make sure to only whitelist trusted IP addresses and regularly review your whitelist to remove any IPs that are no longer needed. # ASN/ISP Targeting (/docs/proxies/getting-started/knowledge-base/asn-isp-targeting) This guide explains what ASN/ISP targeting is and how to use it in Geonode for precise proxy control and optimized network performance. ## What Is ASN/ISP Targeting? **ASN/ISP targeting** allows you to filter proxy connections based on an Internet Service Provider (ISP) or **Autonomous System Number (ASN)**.\ This helps you choose proxies from specific ISPs to improve reliability, compliance, and performance for targeted use cases. ## How ASN/ISP Targeting Works * Every ISP is assigned a unique ASN (Autonomous System Number). * Geonode enables users to select proxies by **country** and **ASN**. * This is especially useful for: * Accessing geo-restricted content * Market research * Increasing connection consistency and anonymity ### Example: Targeting a Specific ASN To target a specific ISP, include both the country code and the ASN in your request: ```text country-us-asn-7018 ``` # Proxy Endpoint Formats (/docs/proxies/getting-started/knowledge-base/endpoint-formats) This guide explains what proxy endpoints are, the types available in Geonode, and how to choose the most suitable format for your specific use case. ## What Is a Proxy Endpoint? A **proxy endpoint** is the address used to connect to a proxy server.\ It defines how your requests are routed through Geonode’s network, ensuring secure, reliable, and efficient data transmission. ## DNS vs. IP-Based Endpoints Geonode allows users to connect either via a **DNS-based hostname** or a **direct IP address**. ### 1. DNS-Based Endpoints (Recommended) * Use a domain name such as `proxy.geonode.io`. * Easier to maintain — DNS automatically updates IPs when servers change. * Reduces connection issues caused by IP rotation. ### 2. IP-Based Endpoints * Use a direct IP address, e.g. `123.45.67.89:9000`. * Skips DNS resolution, which can be slightly faster in some setups. * Ideal for apps or devices that **don’t support DNS hostnames**. ## Available Endpoint Formats Geonode supports multiple endpoint formats to ensure compatibility with different systems and authentication methods. Choosing an Endpoint Format ### 1. `hostname:port` * Example: `proxy.geonode.io:9000` * Best for: Simple connections without authentication. ### 2. `hostname:port:username:password` * Example: `proxy.geonode.io:9000:user:pass` * Includes authentication credentials for secure access. * Best for: APIs and applications requiring basic authentication. ### 3. `hostname:port@username:password` * Example: `proxy.geonode.io:9000@user:pass` * Alternative authentication format. * Best for: Legacy or custom proxy clients. ### 4. `username:password@hostname:port` * Example: `user:pass@proxy.geonode.io:9000` * Credentials placed before the host for compatibility. * Best for: Apps that authenticate before establishing a connection. ### 5. `http://username:password@server:port` * Example: `http://user:pass@proxy.geonode.io:9000` * Explicit HTTP/HTTPS proxy format. * Best for: Secure browsing and authenticated API requests. ## Choosing the Right Proxy Endpoint Format | Format | Best For | | -------------------------------------- | --------------------------------------------- | | `hostname:port` | Standard connections without authentication | | `hostname:port:username:password` | Secure authentication for APIs & applications | | `hostname:port@username:password` | Custom or legacy proxy configurations | | `username:password@hostname:port` | Systems needing pre-authentication | | `http://username:password@server:port` | Secure browsing, authenticated API access | ## Tips and Best Practices * **Start simple:** Use `hostname:port` if no authentication is required. * **Prefer DNS:** It’s more reliable and updates automatically. * **Use IP endpoints** only for tools that can’t resolve hostnames. * **Always use HTTPS** or encrypted channels when handling sensitive credentials. * Test different formats — some apps accept only specific syntaxes. ## Summary Proxy endpoints define how your connection routes through Geonode’s network.\ Choosing the correct format ensures: * Smooth compatibility with your app or client, * Secure authentication where required, * Stable performance with minimal downtime. # Gateways (/docs/proxies/getting-started/knowledge-base/gateway) This guide explains what proxy gateways are and how Geonode uses them to optimize proxy routing, improve performance, and maintain anonymity. ## What Is a Proxy Gateway? A **proxy gateway** is a server that routes your internet traffic through a specific geographic location.\ By selecting a gateway, you define the **entry point** for your proxy requests — controlling how and where your traffic appears to originate. Geonode provides multiple gateway locations to: * Access region-restricted content * Improve browsing speed and stability * Maintain privacy and network anonymity ## Geonode’s Available Gateways Geonode currently offers the following gateway locations: Gateway * **France** * **United States** * **Singapore** Each gateway routes your traffic through a regional hub, providing faster speeds and greater accessibility for local or restricted online services. ## Why Use a Gateway? ### 1. Access Geo-Restricted Content * Browse the internet as if you’re in another country * Useful for streaming, e-commerce, and international research ### 2. Improve Connection Speed * Routes traffic through optimized, low-latency locations * Reduces delay and improves access to nearby content ### 3. Enhance Privacy and Security * Masks your real IP and location * Adds a layer of protection when using public or sensitive networks ## Choosing the Right Gateway | Gateway | Best For | | ----------------- | -------------------------------------------------------------- | | **France** | EU-based content, GDPR-compliant data collection | | **United States** | Streaming, U.S. market research, and e-commerce | | **Singapore** | Low-latency connections in Asia, accessing region-locked sites | ## Tips and Best Practices * **Use the closest gateway** to minimize latency and maximize performance. * **Select gateways strategically** based on your target region or content source. * **Switch gateways** if your connection feels slow — load balancing can vary by region. ## Summary Proxy gateways determine where your connection enters Geonode’s network.\ By selecting the right gateway, you can: * Improve connection speed, * Access region-specific content, * And enhance your overall privacy and anonymity online. # Geo-Targeting (/docs/proxies/getting-started/knowledge-base/geo-targeting) This guide explains what Geo-Targeting is and how Geonode allows you to choose proxies based on country, state, and city for precise connection control and regional flexibility. ## What Is Geo-Targeting? **Geo-Targeting** lets you filter and select proxy locations by **country**, **state**, and **city**.\ It enables businesses, developers, and marketers to access content and services as if they were browsing directly from a specific geographic area. ## How Geo-Targeting Works in Geonode Geonode supports **three levels of geo-targeting**, shown in the example below: Geo-Targeting 1. **Country Targeting** – Select a country for your proxy connection. 2. **State Targeting** – Choose a particular state or region within that country. 3. **City Targeting** – Narrow it down further by selecting a specific city, or set it to **Any** for broader access. ## Why Use Geo-Targeting? ### 1. Access Geo-Restricted Content * Bypass region-based restrictions on websites, apps, and streaming services. * View local search engine results, prices, and ads as a user from that area. ### 2. Improve Localized Testing and Marketing * Run **ad verification** and A/B tests by region. * Test **localized websites, apps, and payment systems** from multiple markets. ### 3. Ensure Compliance and Security * Simulate user behavior from different regions for compliance testing. * Avoid detection by using realistic, location-accurate IPs. ## Choosing the Right Geo-Targeting Option | Geo-Targeting Level | Best For | | --------------------- | --------------------------------------------------------------------- | | **Country Targeting** | General browsing, international research, region-based content access | | **State Targeting** | Regional services, ad verification, localized e-commerce | | **City Targeting** | Precision testing, local SEO, hyper-targeted ad campaigns | ## Notes and Recommendations * The availability of **state** and **city** targeting depends on the selected country. * Some cities may have fewer proxy IPs — use **“Any”** for better coverage. * For faster performance, select a location closer to your target audience or testing region. ## Best Practices * Use **Country Targeting** for broad, region-specific access. * Choose **State** or **City Targeting** for more precise, location-based control. * Start with **Country Targeting** and refine your selection as needed for campaigns or tests. ## Summary Geo-Targeting gives you granular control over your proxy location — from country down to city level.\ By choosing the right targeting depth, you can: * Access localized content, * Run accurate regional tests, and * Optimize performance for specific markets. # IP Address (/docs/proxies/getting-started/knowledge-base/ip-address) This guide explains what an IP address is, how it functions, and why it’s important when using proxies. ## What Is an IP Address? An **IP address (Internet Protocol Address)** is a unique identifier assigned to every device connected to the internet.\ Think of it as a mailing address for your device — without it, data wouldn’t know where to go. When you visit a website, your IP address tells the site where to send the information you requested.\ It can also reveal your **approximate location**, **internet provider**, and **network type**. ## How IP Addresses Work * Every device gets an IP address from its **Internet Service Provider (ISP)**. * When you make a request (like opening a website), your IP acts as a return address. * The server sends data back to that address — completing the connection. ## Types of IP Addresses There are several types of IPs depending on their use, assignment method, and format. ### 1. Public vs. Private IP Addresses | Type | Description | | ---------------------- | --------------------------------------------------------------------------- | | **Public IP Address** | Assigned by your ISP and used to communicate over the internet. | | **Private IP Address** | Used within local networks (e.g., home Wi-Fi) to identify internal devices. | ### 2. Static vs. Dynamic IP Addresses | Type | Description | | ---------------------- | -------------------------------------------------------------------- | | **Static IP Address** | Fixed and unchanging. Common for servers, businesses, and VPNs. | | **Dynamic IP Address** | Changes periodically. Used by most home connections for flexibility. | ### 3. IPv4 vs. IPv6 | Type | Description | | -------- | -------------------------------------------------------------------------------- | | **IPv4** | Classic format with four number sets, e.g. `192.168.1.1`. Still the most common. | | **IPv6** | Newer format supporting many more devices, e.g. `2001:db8::ff00:42:8329`. | ## Why IP Addresses Matter in Proxies When using proxies, your **real IP** is replaced with another one — masking your location and identity.\ This provides privacy, allows region-based access, and reduces the risk of detection or blocking. ### 1. Residential vs. Datacenter IPs | Type | Description | | ------------------- | ------------------------------------------------------------------------------- | | **Residential IPs** | Provided by ISPs to real devices. Highly trusted and less likely to be flagged. | | **Datacenter IPs** | Generated by data centers or hosting services. Faster, but easier to detect. | ### 2. Rotating vs. Sticky IPs | Type | Description | | ---------------- | ------------------------------------------------------------------------------------ | | **Rotating IPs** | Change with every request or at timed intervals — ideal for scraping and automation. | | **Sticky IPs** | Remain constant for a session — best for logins, account management, or testing. | ## How to Check Your IP Address You can easily find your public IP: * Search **“What is my IP”** on Google. * Visit [whatismyip.com](https://www.whatismyip.com). * Check network details in your router or device settings. ## Tips and Best Practices * Use proxies to protect your identity and location online. * **Residential IPs** → more privacy and authenticity. * **Datacenter IPs** → more speed and cost efficiency. * **IPv6** is expanding, but **IPv4** remains dominant for most users. * Choose a reliable provider like **Geonode** to balance **security, anonymity, and performance**. ## Summary An IP address is your device’s online identifier — essential for routing internet traffic.\ Using a proxy lets you control or hide that identity, giving you: * More **privacy**, * Better **regional access**, and * Stronger **protection** online. # IP Types (/docs/proxies/getting-started/knowledge-base/ip-type) This guide explains what IP types are and how Geonode provides them for different proxy use cases. ## What Are IP Types? **IP types** define how an IP address is sourced and used within proxy networks.\ Geonode offers three main categories: IP Types * **Residential IPs** – Real-user IPs assigned by Internet Service Providers (ISPs). * **Datacenter IPs** – IPs generated by third-party data centers. * **Mixed IPs** – A combination of Residential and Datacenter IPs for flexibility. Each serves a unique purpose — from bypassing geo-restrictions to powering high-speed automation. ## 1. Residential IPs ### What Are Residential IPs? Residential IPs come from **real user connections** provided by ISPs.\ Because they look like genuine home users, websites see them as legitimate traffic. ### Best For * Web scraping with minimal detection risk * Accessing geo-restricted content * Managing e-commerce or social media accounts * Secure streaming and gaming ### Pros ✔️ High trust level — less likely to be blocked\ ✔️ Works on strict, detection-sensitive sites ### Cons ❌ Slower than datacenter IPs\ ❌ More expensive due to limited availability ## 2. Datacenter IPs ### What Are Datacenter IPs? Datacenter IPs are **server-based IPs** generated by hosting providers, not ISPs.\ They offer higher speed and lower cost but are easier to detect. ### Best For * Bulk web scraping and data collection * SEO monitoring and automation * High-performance or bot-driven tasks ### Pros ✔️ Fast and reliable\ ✔️ Cost-effective for large-scale operations ### Cons ❌ Easier to detect and block on strict websites\ ❌ Limited success with region-locked or sensitive content ## 3. Mixed IPs ### What Are Mixed IPs? Mixed IPs combine both **Residential** and **Datacenter** sources — balancing realism and performance. ### Best For * Scalable web scraping and automation * Testing and development environments * Multi-purpose e-commerce or social media tasks ### Pros ✔️ Balanced speed, security, and price\ ✔️ Greater flexibility across use cases ### Cons ❌ Not as anonymous as pure Residential IPs\ ❌ Some tasks may require dedicated IP types ## Comparison: Residential vs Datacenter vs Mixed | Feature | Residential IPs | Datacenter IPs | Mixed IPs | | ------------- | ----------------------------------- | ------------------------------ | ------------------------------------------------ | | **Speed** | Moderate | Fast | Balanced | | **Anonymity** | High (real users) | Low (easily detected) | Moderate | | **Cost** | Expensive | Affordable | Mid-range | | **Best For** | Geo access, social media, streaming | Bulk scraping, SEO, automation | Versatile tasks needing both reliability & speed | ## How to Choose the Right IP Type * **Residential IPs** → for high anonymity and geo-specific access. * **Datacenter IPs** → for fast, large-scale automation. * **Mixed IPs** → for a balanced approach between performance and stealth. ## Tips and Best Practices * Start with **Mixed IPs** to test both speed and reliability. * Use **Residential IPs** for sensitive or geo-restricted sites. * Use **Datacenter IPs** for speed-critical automation or scraping. * Always verify which type performs best for your target websites. ## Summary Each IP type serves a distinct purpose: * **Residential** for trust and authenticity, * **Datacenter** for speed and scale, * **Mixed** for flexibility. Choosing the right type ensures stable, efficient, and secure proxy performance for your specific needs. # IP Whitelisting (/docs/proxies/getting-started/knowledge-base/ip-whitelist) **IP whitelisting** is a security feature that allows only approved IP addresses to connect to your network or service.\ It blocks unauthorized access and ensures safer connections. When using proxies, **whitelisted IPs** let you connect without a password — making access both **simple** and **secure**. ## Benefits of IP Whitelisting * **Better Security** – Only trusted IPs can connect. * **More Control** – You decide who can access your system. * **Easy Access** – No need for manual logins or credentials. ## Best Practices To maximize security and efficiency, follow these recommendations: * **Add Only Trusted IPs** – Whitelist secure, verified networks. * **Use Clear Labels** – Name IPs (e.g., “Home,” “Office”) for easy identification. * **Review Regularly** – Remove unused or outdated IPs. * **Combine with Other Tools** – Use firewalls and authentication for layered protection. ## Troubleshooting Common Issues Ensure your proxy IP is whitelisted — connections may fail if it isn’t. Add your new IP address from your current location to regain access. * Verify that you entered the correct **public IP address**. * Check if your IP has changed and update it in the whitelist. ## Summary **IP whitelisting** strengthens your proxy setup by restricting access to trusted sources only.\ It simplifies authentication and prevents unauthorized use — ideal for both personal and enterprise environments. For a detailed setup guide, see:\ ➡️ [How to Whitelist Your IP Address](/docs/proxies/getting-started/setup_and_configuration/advance-configuration/whitelist-ip) # Overview (/docs/proxies/getting-started/knowledge-base/overview) Welcome to the **Proxy Service Guide** — your complete reference for understanding and configuring Geonode proxies. This guide covers all the essential concepts for effective proxy usage, including: * **IP Addresses** – Learn how IPs work and why they matter in proxy networks. * **Proxy Types** – Understand the differences between residential, datacenter, and mixed IPs. * **Authentication** – Explore IP whitelisting, username/password access, and security best practices. * **Geo-Targeting** – Configure country, state, and city-level targeting. * **Ports & Sessions** – Manage connections, rotation, and session persistence. By mastering these topics, you’ll be able to: * Optimize connection performance, * Maintain high security and privacy, * Avoid detection across different platforms and use cases. Explore each section to gain a clear, practical understanding of how Geonode proxies work — and how to use them efficiently for your specific needs. # Ports (/docs/proxies/getting-started/knowledge-base/port-type) This guide explains what ports are, how they work, and why they play a key role in proxy connections. ## What Is a Port? Imagine your device as an apartment building (your **IP address**) with many rooms inside — these rooms are **ports**.\ Each room serves a different purpose, helping data reach the right application. * Every device connected to the internet has a unique IP address. * **Ports** act as doorways that direct data traffic to the correct app or service. * Each service (browsing, messaging, streaming, etc.) uses its own port number to avoid confusion. Without ports, your device wouldn’t know which application incoming data belongs to — everything would arrive at the same “door.” ## How Ports Work in Practice Let’s say you open several browser tabs at once: * One tab has **WhatsApp Web**, * Another shows **Facebook**, * And the third streams **YouTube**. When a WhatsApp message arrives, your computer needs to know which app it’s for.\ Ports make this possible: * WhatsApp Web might use **port 5222**, * Facebook might use **port 443**, * YouTube might use another. Your system reads the port number, matches it with the right app, and delivers the data correctly.\ When that session closes, the port becomes available again — these are called **ephemeral ports** (temporary ports for short-lived connections). ## Ports in Proxy Configuration In proxy setups, ports function like **dedicated lanes** for your traffic.\ Different ports connect to different types of proxy behavior — such as rotating or sticky sessions. ### Common Proxy Ports in Geonode | Proxy Type | Port Range | | ------------------- | ----------- | | **SOCKS5 Rotating** | 11000–11010 | | **SOCKS5 Sticky** | 12000–12010 | | **HTTP Rotating** | 9000–9010 | | **HTTP Sticky** | 10000–10900 | Each range corresponds to a specific connection type, giving you control over how often your proxy IP changes. ## How to Configure Proxy Ports When setting up a proxy, choose a port according to your goal: * **Rotating proxies** → Use ports from the *rotating* range to get a new IP for each request. * **Sticky proxies** → Use ports from the *sticky* range to keep the same IP for a set duration. ➡️ For detailed setup steps, see:\ [Proxy Port Configuration](/docs/proxies/getting-started/setup_and_configuration/advance-configuration/port-configuration) ## Best Practices * Treat each **port** as a separate communication channel. * Choose **rotating** or **sticky** ports depending on your use case. * Once a port is assigned to a specific country, **remove the assignment** before reusing it. * Understanding ports helps ensure a stable and efficient proxy connection. ## Summary **Ports** are the internal “routes” that direct internet traffic where it needs to go.\ In proxy configurations, they define how your connection behaves — rotating for dynamic IPs, sticky for consistent sessions.\ Selecting the right port ensures smooth, secure, and optimized proxy performance. # Protocol Type (/docs/proxies/getting-started/knowledge-base/protocol-type) Choosing the right **protocol type** is essential for optimizing proxy performance, ensuring security, and maintaining reliable connections. This guide explains what a protocol is in the context of proxies and helps you decide between **HTTP/HTTPS** and **SOCKS5** for your specific use case. ## What Is a Protocol in Proxy Usage? A **protocol** is a set of rules that defines how data is transmitted between devices across a network.\ When you use a proxy, the protocol determines **how requests are sent, processed, and returned**. ## Types of Proxy Protocols Geonode supports two main protocol types: * **HTTP / HTTPS** * **SOCKS5** Let’s explore how they differ — and when to use each. *** ## 1. HTTP / HTTPS Proxies **HTTP/HTTPS proxies** handle standard web traffic.\ They forward web requests between your browser (or app) and the destination website.\ The **HTTPS** variant encrypts the connection, providing extra privacy and protection. ### Ideal For * General web browsing * APIs and web services * Web scraping and automation * Managing multiple online accounts ### Key Features * Easy to configure — compatible with most browsers and apps. * HTTPS ensures encrypted data transmission. * Works seamlessly with tools like **Postman**, **cURL**, and browser extensions. ### Port Range * **Rotating Proxy:** `9000–9010` * **Sticky Proxy:** `10000–10900` ### Common Use Cases * **Web Scraping:** Automate data collection while reducing detection risk. * **Account Management:** Maintain consistent login sessions. * **API Requests:** Manage large volumes of secure, authenticated requests. *** ## 2. SOCKS5 Proxies **SOCKS5** is a more flexible and powerful proxy protocol.\ Unlike HTTP, it works at a lower network level and doesn’t alter the transmitted data — making it ideal for **non-web traffic** and privacy-focused applications. ### Ideal For * Privacy and anonymity * P2P (peer-to-peer) connections * Bypassing firewalls and geo-restrictions * Streaming, gaming, and VoIP ### Key Features * Supports **both TCP and UDP**, enabling real-time communication. * Offers **higher anonymity** — does not inject identifying headers. * Handles all types of traffic: HTTP, FTP, VoIP, gaming, etc. ### Port Range * **Rotating Proxy:** `11000–11010` * **Sticky Proxy:** `12000–12010` ### Common Use Cases * **Bypassing Restrictions:** Access blocked content securely. * **Torrenting:** Stable and fast peer-to-peer transfers. * **VoIP & Streaming:** Lower latency and improved connection quality. *** ## How to Choose the Right Protocol | Criteria | **HTTP / HTTPS** | **SOCKS5** | | --------------- | ------------------------------ | ----------------------------------- | | **Security** | HTTPS encryption protects data | High anonymity, supports encryption | | **Speed** | Fast for web-based traffic | Faster for non-HTTP traffic | | **Flexibility** | Limited to web requests | Supports all internet protocols | | **Best For** | Websites, APIs, automation | P2P, VoIP, bypassing firewalls | *** ## Summary * Use **HTTP/HTTPS** for simplicity, compatibility, and secure web requests. * Choose **SOCKS5** for advanced use cases, full traffic support, and stronger anonymity. * If you’re just getting started, start with HTTP/HTTPS — you can always switch to SOCKS5 later as your needs evolve. # Protocols (/docs/proxies/getting-started/knowledge-base/protocols) Geonode’s proxy network supports several protocols, each designed for different use cases and levels of security.\ Choosing the right protocol ensures compatibility, performance, and privacy when connecting through Geonode. ## Supported Protocols Geonode currently supports the following proxy protocols: 1. **HTTP** — Hypertext Transfer Protocol 2. **HTTPS** — Hypertext Transfer Protocol Secure 3. **SOCKS5** — Socket Secure version 5 Each protocol serves a specific purpose — from standard web browsing to high-security, high-speed connections. *** ## Protocol Comparison | Protocol | Security Level | Best For | Availability | | ---------- | ------------------------- | ----------------------------------------- | ------------ | | **HTTP** | ❌ No encryption | General web scraping, browsing, APIs | ✅ Supported | | **HTTPS** | ✅ SSL/TLS encrypted | Secure websites, automation, transactions | ✅ Supported | | **SOCKS5** | ✅ High anonymity & secure | Streaming, gaming, bypassing restrictions | ✅ Supported | *** ## Unsupported Protocols The following protocols are **not supported** by Geonode proxies: * **UDP** (User Datagram Protocol) * **Email protocols** — such as SMTP, IMAP, and POP3 These protocols operate differently from HTTP/HTTPS and SOCKS5 and are not part of Geonode’s proxy infrastructure. *** ## Best Practices * Use **HTTP** for general data scraping and standard web requests. * Choose **HTTPS** for secure, encrypted communication and automated workflows. * Select **SOCKS5** for streaming, gaming, or tasks requiring maximum privacy and flexibility. Geonode focuses on stable, encrypted proxy connections — therefore, **UDP and email-related traffic are excluded** for security and performance reasons. *** ## Summary Geonode supports the three most widely used proxy protocols — **HTTP**, **HTTPS**, and **SOCKS5** — giving users the flexibility to choose between speed, compatibility, and security.\ Pick the protocol that fits your use case and connection needs to ensure optimal performance and reliability. # Proxy Authentication Methods (/docs/proxies/getting-started/knowledge-base/proxy-authentication) *** This guide will help you understand what the Proxy Authentication method is and the different ways geonode provides to authenticate it. ## **What is Proxy Authentication?** Proxy authentication is a way to verify your identity before accessing a proxy server. It helps keep your connection secure and ensures only authorized users can use the proxy. Geonode provides two ways to authenticate: 1. Using a username and password **(Basic Authentication)** 2. Whitelisting your IP address **(IP Authentication)** Each method has its advantages, depending on your use case. *** ## **1. Authenticating with Username and Password (Basic Authentication)** This method requires you to send your Geonode proxy username and password with every request. Many applications, APIs, and automation tools support this authentication method. ### **How Basic Authentication Works** 1. Format your credentials as: ``` username:password ``` 2. Convert this string to Base64 format (for security). 3. Add the encoded credentials to your request header like this: ``` Authorization: Basic BASE64_ENCODED_STRING ``` ➡️ For a full setup guide, check: [Set Up Proxy Authentication in API](/docs/proxies/getting-started/setup_and_configuration/advance-configuration/proxy-authentication-in-api) *** ## **2. Authenticating with IP Whitelisting** IP whitelisting allows you to skip entering your username and password by authorizing a specific IP address to access the proxy. ➡️[What is IP Whitelisting?](/docs/proxies/getting-started/setup_and_configuration/advance-configuration/whitelist-ip) ### **When to use IP Whitelisting** * You have a static IP and want seamless authentication. * You are automating tasks and don't want to include credentials in every request. * You want better security by restricting access to trusted IPs. ➡️ [Learn How to Whitelist Your IP Address](/docs/proxies/getting-started/setup_and_configuration/advance-configuration/whitelist-ip) *** ## **Choosing the Right Authentication Method** | Authentication Method | Best For | Requires Credentials? | | ----------------------- | -------------------------------------------------------- | --------------------- | | **Username & Password** | Most API calls, scripts, and browser extensions | ✅ Yes | | **IP Whitelisting** | Trusted networks, automation, and security-focused users | ❌ No | *** ## **Final Tips** * Use Username & Password for flexible authentication across different devices. * Use IP Whitelisting if you have a static IP and want to avoid entering credentials. * Always ensure your authentication details are kept secure and not exposed in scripts or public repositories. # Port Usage (/docs/proxies/getting-started/knowledge-base/proxy-usage) This guide explains how ports function when configuring **SOCKS5** and **HTTP** proxies in Geonode.\ Geonode provides **unlimited port ranges**, meaning you can generate as many proxies as needed — without performance limits. ## How Port Ranges Work Geonode assigns specific port ranges to organize different proxy types and session behaviors: | Proxy Type | Session Type | Port Range | | ---------- | ------------ | ------------- | | **HTTP** | Rotating | `9000–9010` | | **HTTP** | Sticky | `10000–10900` | | **SOCKS5** | Rotating | `11000–11010` | | **SOCKS5** | Sticky | `12000–12010` | These ranges act as **identifiers**, not limitations.\ You can reuse the same port as often as you like — each connection request automatically generates a **unique proxy IP**. *** ## Key Points to Remember ### Same Performance Across All Ports * Whether you use port `9000` or `9010`, performance and stability remain identical. * Reusing a port does **not** affect connection speed or session quality. ### Efficient Port Reuse * You don’t need to change ports for every connection. * Thousands of requests can run on the **same port** without issues or conflicts. * Port selection mainly helps organize your setup (for example, by type or location). *** ## Common Use Cases ### Using Multiple Proxies If you need multiple proxies for automation or scraping: * You can assign **all requests to the same port** (e.g., `9000`). * Each request will still produce a **unique IP address**, ensuring anonymity and diversity. * There’s no impact on speed or reliability. ### Rotating vs. Sticky Sessions * **Rotating Proxies** — Ports `9000–9010` (HTTP) and `11000–11010` (SOCKS5):\ Each request gets a **new IP address** automatically. * **Sticky Proxies** — Ports `10000–10900` (HTTP) and `12000–12010` (SOCKS5):\ The same IP remains active for a set duration (useful for account logins, testing, etc.). ### Country-Specific Port Assignments * When you assign a port to a specific **country**, it cannot be reused for another country until the previous assignment is **deleted**. * This ensures accurate geo-routing and prevents proxy conflicts. *** ## Common Misconceptions > “Fewer ports mean fewer proxies.” ❌ False.\ The **number of ports** has no effect on how many proxies you can use.\ Ports are simply **access points** — Geonode dynamically assigns IPs behind them. *** ## Summary * You can reuse the same port indefinitely — performance stays the same. * Rotating ports change IPs automatically; sticky ports keep the same IP for a set session. * Assign ports carefully if you’re working with geo-targeted proxies. * Geonode’s architecture allows **unlimited proxy generation** across all ports. By understanding how port usage works, you’ll be able to manage proxy sessions efficiently — without worrying about limits or performance degradation. # Proxy (/docs/proxies/getting-started/knowledge-base/proxy) This guide explains what a proxy is, how it works, and why it’s essential for online security, privacy, and automation. ## What Is a Proxy? A **proxy server** acts as a bridge between your device and the internet.\ Instead of connecting directly to a website, your request first passes through the proxy, which forwards it on your behalf. ### Analogy: A Proxy as a Messenger Imagine you’re ordering food but don’t want the restaurant to know your home address.\ You ask a **friend (proxy)** to pick it up and deliver it. The restaurant only sees your friend’s address — not yours. Likewise, when using a proxy: * Your real IP is hidden. * The proxy server communicates with websites for you. * Websites see the proxy’s IP instead of yours. *** ## How a Proxy Works When you connect through a proxy, your data follows four simple steps: 1. **Request Sent** — You request access to a website.\ The request goes to the proxy server first. 2. **IP Replaced** — The proxy swaps your IP with its own and sends the request onward. 3. **Response Received** — The target website sends data back to the proxy. 4. **Response Delivered** — The proxy forwards that data back to you, keeping your IP hidden. How Does a Proxy Work? *** ## Types of Proxies ### 1. Forward vs. Reverse Proxies | Type | Description | | ----------------- | -------------------------------------------------------------------------------------------------------- | | **Forward Proxy** | Protects the user by hiding their IP when browsing or scraping the internet. | | **Reverse Proxy** | Protects servers by managing inbound traffic, improving security and load balancing for hosted websites. | *** ### 2. Residential vs. Datacenter Proxies | Type | Description | | --------------------- | ----------------------------------------------------------------- | | **Residential Proxy** | Uses real IPs from ISPs — highly trusted and difficult to detect. | | **Datacenter Proxy** | Uses IPs from data centers — faster but easier to identify. | *** ### 3. Rotating vs. Sticky Proxies | Type | Description | | ------------------ | ------------------------------------------------------------------------------------------------- | | **Rotating Proxy** | Changes IP for every request or after a set time — great for scraping and large-scale automation. | | **Sticky Proxy** | Keeps the same IP for a session — best for logins, account management, and session-based tasks. | *** ## Why Use a Proxy? ### 1. Privacy & Anonymity * Hides your real IP and online identity. * Prevents tracking and profiling by websites. ### 2. Geo-Unblocking * Access region-restricted content (streaming, marketplaces, apps). * Simulate browsing from specific countries. ### 3. Security & Protection * Avoid IP bans and throttling. * Reduce exposure to threats by masking your origin. ### 4. Automation & Data Collection * Collect public data at scale without detection. * Power SEO, price-tracking, and analytics tools. *** ## How to Choose the Right Proxy | Use Case | Recommended Proxy Type | | ------------------------------- | ------------------------- | | **Browsing & Privacy** | Residential Proxy | | **Web Scraping & Automation** | Rotating Datacenter Proxy | | **Geo-Blocked Content** | Residential Proxy | | **Multiple Account Management** | Sticky Residential Proxy | *** ## Summary * Proxies act as intermediaries that protect your identity online. * Choose **rotating** proxies for scraping, **sticky** ones for stable sessions. * Residential proxies offer trust and geo-access, while datacenter proxies focus on speed. * Understanding proxy types helps you stay secure, anonymous, and efficient across all use cases. # Rotating Proxies (/docs/proxies/getting-started/knowledge-base/rotating-proxies) This guide explains what rotating proxies are, how they work, and how they help maintain anonymity and stability during high-volume requests. ## What Are Rotating Proxies? A **rotating proxy** automatically assigns a new IP address after each request — or after a defined time interval.\ This rotation helps prevent detection and blocking by websites that monitor for repeated activity from a single IP. *** ## How Rotating Proxies Work * Each outgoing request is assigned a **unique IP address** from a large proxy pool. * The IP rotates **automatically** after every request or based on a timer (e.g., every few minutes). * This makes requests appear as if they are coming from different users around the world. Rotating proxies are ideal for maintaining anonymity and avoiding IP bans in automation and scraping tasks. *** ## Key Benefits | Benefit | Description | | ------------------------ | ------------------------------------------------------------- | | **High Anonymity** | Constantly changing IPs prevent tracking and detection. | | **Bypass Rate Limits** | Allows multiple requests without triggering anti-bot systems. | | **Improved Scalability** | Enables large-scale web scraping and data collection. | | **Global Coverage** | Access data from various geographic locations automatically. | *** ## Common Use Cases | Use Case | Why It’s Useful | | --------------------------- | ---------------------------------------------------------- | | **Web Scraping** | Collects large datasets without detection or bans. | | **Ad Verification** | Checks ad placements from multiple IPs and regions. | | **Market Research** | Gathers pricing and trend data from competitor sites. | | **SEO Monitoring** | Tracks rankings without triggering search engine security. | | **Social Media Automation** | Manages multiple accounts without getting flagged. | *** ## Limitations While rotating proxies offer flexibility and anonymity, they’re not ideal for all scenarios: | Limitation | Impact | | ---------------------------------------- | --------------------------------------------------------- | | **Unstable Sessions** | Frequent IP changes break login-based activities. | | **Detection by Strict Websites** | Some systems still recognize automated behavior. | | **Potential for Inconsistent Responses** | Different IPs may yield varying localized or cached data. | *** ## Summary * **Rotating proxies** provide high anonymity and scalability for automation and data collection. * Use them for **web scraping**, **ad verification**, and **SEO monitoring**. * For login-based or session-sensitive tasks, switch to **sticky proxies** to maintain consistency. * Adjust rotation frequency to balance anonymity with connection stability. # Session Type (/docs/proxies/getting-started/knowledge-base/session-type) Choosing the right **session type** is key to optimizing proxy performance, maintaining stable connections, and avoiding detection.\ This guide explains what sessions are, how they work, and when to use **rotating** or **sticky** sessions. ## What Is a Session? In proxy configuration, a **session** refers to the period during which you use the same IP address.\ The session type determines whether your IP remains constant (sticky) or changes frequently (rotating). Geonode supports two main session types: * **Rotating Sessions** * **Sticky Sessions** *** ## Rotating Sessions A **rotating session** assigns a new IP address for every request or after a defined time interval.\ This setup is designed for maximum anonymity and large-scale operations where detection risk is high. ### Key Features | Feature | Description | | ------------------------- | ----------------------------------------------------------- | | **Automatic IP Rotation** | Changes IP after each request or at a set interval. | | **High Anonymity** | Prevents blocks by cycling through multiple IPs. | | **Scalable** | Ideal for bulk data collection and high-frequency requests. | ### Best For * Web scraping and data aggregation * Market or SEO research * Ad verification and automation tasks *** ## Sticky Sessions A **sticky session** maintains the same IP for a longer duration, creating a stable and continuous connection.\ This type is useful when session consistency or login persistence is required. ### Key Features | Feature | Description | | ---------------------------- | ------------------------------------------- | | **Consistent IP Address** | Keeps the same IP throughout the session. | | **Custom Duration** | Control how long the IP remains active. | | **Reduced Reauthentication** | Avoids frequent logouts and session resets. | ### Best For * Account management and login sessions * E-commerce or transactional activities * Application and website testing *** ## Comparison: Rotating vs. Sticky Sessions | Criteria | Rotating Sessions | Sticky Sessions | | --------------- | ----------------------------------------- | --------------------------------------- | | **IP Behavior** | Changes with each request or set interval | Remains the same for a defined duration | | **Anonymity** | High – ideal for stealth and scaling | Moderate – consistent IP for stability | | **Performance** | Suited for high-volume tasks | Suited for persistent sessions | | **Use Case** | Scraping, automation, data collection | Logins, transactions, testing | *** ## Summary * **Rotating sessions** provide higher anonymity and flexibility for automation, scraping, and large-scale operations. * **Sticky sessions** ensure reliability for logins, testing, and tasks that depend on stable IPs. * Combine both session types when needed — for instance, rotating proxies for data gathering and sticky proxies for account-based operations.\ Understanding how sessions work helps you balance **anonymity, performance, and connection stability** for your specific goals. # Sticky Session (/docs/proxies/getting-started/knowledge-base/sticky-session) A **sticky session** maintains the same IP address for a fixed period instead of changing it with every request.\ This allows for stable, consistent sessions — ideal for activities like account logins, e-commerce transactions, and automation tasks that rely on persistent identity. *** ## How Sticky Sessions Work * When you start a connection, a unique IP address is assigned to your session. * The IP remains active for a defined duration (e.g., 10–30 minutes, or longer). * Once the session expires, a new IP is automatically assigned on the next connection. * You can manually release a session early if you need to refresh your IP before expiration. Sticky sessions offer balance — they maintain stability without locking you to a single IP indefinitely. *** ## Key Benefits | Benefit | Description | | ----------------------------- | ---------------------------------------------------------------------- | | **Stable Identity** | Keeps the same IP during the session, ideal for login-based workflows. | | **Fewer CAPTCHA Challenges** | Reduces interruptions caused by frequent IP changes. | | **Persistent Access** | Prevents session resets on e-commerce and social media platforms. | | **Better Automation Control** | Ensures bots and tools operate smoothly across multi-step actions. | *** ## Common Use Cases | Use Case | Why It’s Useful | | ------------------------------ | ----------------------------------------------------- | | **Account Management** | Maintains session stability and prevents logouts. | | **E-Commerce & Checkout Bots** | Avoids disruptions during multi-step transactions. | | **Web Scraping (Login Sites)** | Enables consistent access for authenticated scraping. | | **Streaming Services** | Prevents interruptions or reauthentication requests. | | **SEO Monitoring** | Keeps IP identity consistent for ongoing checks. | *** ## Limitations | Limitation | Impact | | ----------------------------- | ---------------------------------------------------------------------------- | | **Overuse of One IP** | Extended use can make IPs easier to flag or block. | | **Limited Parallel Requests** | Using the same IP across many tasks can reduce efficiency. | | **Session Expiration** | Once the duration ends, the IP changes automatically, interrupting sessions. | *** ## Summary * **Sticky sessions** provide reliable, consistent connections — perfect for logins, form submissions, and stable workflows. * If an IP becomes blocked, you can **release the session** and instantly obtain a new one. * For **high anonymity and frequent IP changes**, use **rotating sessions** instead. * Choosing between sticky and rotating proxies depends on whether your task needs **stability or stealth**. # Threads in Proxy Usage (/docs/proxies/getting-started/knowledge-base/thread) Threads determine how many tasks can run at the same time when using proxies.\ Understanding how they work helps optimize performance, avoid bans, and make the most out of your proxy setup. *** ## What Are Threads? A **thread** is a unit of execution within a program.\ In proxy usage, threads allow you to run multiple actions—such as web requests or scrapers—**simultaneously**, improving efficiency and speed. ### Analogy: Threads Are Like Checkout Counters Imagine a supermarket with one cashier: everyone waits in line.\ Open 10 counters, and customers check out faster. * **Single-threaded** → one cashier, one queue * **Multi-threaded** → multiple cashiers, faster service That’s exactly how threads improve proxy-based automation and scraping. *** ## How Threads Work in Proxies Threads are used to send many requests at once, without waiting for each one to finish before starting the next. ### Without Threads (Single-Threaded) * Only one request runs at a time. * The next request waits for the previous one to finish. * Very slow when dealing with large datasets. ### With Threads (Multi-Threaded) * Many requests run at the same time. * No waiting between requests. * Data collection and automation happen much faster. *** ## Why Threads Matter in Proxy Usage | Benefit | Explanation | | ------------------------------- | ----------------------------------------------------------------- | | **Faster Data Collection** | Executes multiple requests in parallel. | | **Efficient Proxy Utilization** | Distributes requests across multiple proxies, reducing detection. | | **Large-Scale Capability** | Ideal for scraping thousands of pages or bulk API requests. | | **Lower Ban Risk** | Threads allow balanced load across IPs to minimize blocking. | *** ## Choosing the Right Number of Threads The ideal number of threads depends on your **proxy type**, **hardware resources**, and **task intensity**. | Scenario | Recommended Threads | | ------------------------------------- | -------------------------------- | | Small-scale scraping (few pages) | 5–10 threads | | Medium-scale scraping (moderate data) | 20–50 threads | | Large-scale scraping (massive data) | 100+ threads | | Using residential proxies | Fewer threads (avoid bans) | | Using datacenter proxies | More threads (faster processing) | ⚠️ **Too many threads** → may cause IP bans or overload your proxies.\ 🐢 **Too few threads** → slows down operations.\ 🎯 **Balance is key** — test and adjust based on performance. *** ## Threads vs. Concurrent Connections Threads and concurrent connections are related but not identical. | Feature | Threads | Concurrent Connections | | -------------- | --------------------------------------- | ------------------------------------------------- | | **Definition** | Execution units within a program | Multiple network connections happening at once | | **Purpose** | Controls how many tasks run in parallel | Defines how many requests are sent simultaneously | | **Example** | Running multiple scrapers in parallel | Opening 50 browser tabs at once | *** ## Summary * Threads enable multitasking and speed up proxy workflows. * Start with fewer threads and scale gradually. * Match thread count to your proxy pool — more proxies can safely support more threads. * Monitor CPU, RAM, and network limits to prevent system overload. Understanding threading is crucial for stable, efficient, and scalable proxy performance. # Access Credentials (/docs/proxies/getting-started/prerequisites/access-credentials) *** To use Geonode's proxy services or API, you need authentication credentials. These credentials include a **username** and **password**, which allow you to securely connect to the service. This guide will help you:\ ✔️ Access your credentials from the Geonode dashboard.\ ✔️ Secure them properly to prevent unauthorized access.\ ✔️ Use them for authentication in API requests. ### Access Your Dashboard * **Login to Geonode Dashboard:** Visit [Geonode Dashboard](https://app.geonode.com/) and log in using your credentials. * **Navigate to Proxy Section:** Go to the **Proxies** section in the dashboard. * **Get Your User Credentials:** * **Username:** Copy your unique API username. * **Password:** Copy your API password. Get Your User Credentials * **Secure Your Credentials:** Keep your API credentials safe and never share them publicly Once you've obtained your API credentials, you're ready to make your first API request or configure your proxy. # Proxy Server Information (/docs/proxies/getting-started/prerequisites/proxy-server-information) *** To connect to Geonode's proxy network, you need the **Proxy IP** and **Port**. These details allow you to configure your applications, scripts, or browser to route traffic through Geonode's servers. ## Steps: Get Proxy Credentials from Geonode 1. Login to the Geonode Dashboard Visit [Geonode Dashboard](https://app.geonode.com/) and sign in. 2. Go to the Proxy Section\ Navigate to the Proxies section to access your assigned proxy details. 3. Select the proxy format as: `hostname:port:username:password` Proxy Format 4. Copy the proxy details from the Geonode dashboard. * Proxy IP: This is the server address you will use. * Port: The port number required to establish a connection. Proxy Detail OR Copy Your Proxy dns and Port Geonode Proxy IP 5. **Save These Details Securely**\ Do not share your proxy details publicly to avoid unauthorized usage. ## Next Steps Once you have your proxy IP and port, you can:\ ✔️ Use them in API requests.\ ✔️ Configure them in your browser, terminal, or scripts. # Verify Proxy Connection (/docs/proxies/getting-started/setup_and_configuration/verify-proxy-connection) import ProxyOs from "../../../../snippets/proxy-os.mdx"; import BrowsersFaqs from "../../../../snippets/browsers-faqs.mdx"; import SupportParagraph from "../../../../snippets/support-paragraph.mdx"; You can verify your proxy connection in two ways: * Using an online verification tool * Using the command line (cURL) *** ## Method 1: Using an Online Tool ### Step 1: Check Your Current IP Address Before enabling your proxy, check your current IP address. 1. Visit [IP API](https://ip-api.com/). 2. Note your IP address for reference. ### Step 2: Connect to Your Proxy Set up and configure your proxy based on your device or browser.\ Follow the setup guides for your platform below: ### Step 3: Verify Your New IP Address Once connected to the proxy: 1. Go back to [IP API](https://ip-api.com/). 2. Verify if the displayed IP address matches your proxy location. 3. If the IP has changed, your proxy connection is active. Response from Web *** ## Method 2: Using the Command Line (cURL) You can also verify the proxy connection through the command line using `curl`. ```bash curl -x proxy.geonode.io:9000 http://ip-api.com ``` # Overview (/docs/proxies/guides/datacenter-proxies/overview) Rotating Datacenter Proxies give you access to high-speed datacenter IPs for browsing, automation, and data collection. They are billed pay-as-you-go by traffic (per GB), and you configure targeting, protocol, and session type from the Geonode dashboard. The dashboard experience is similar to [Residential Proxies](/docs/proxies/guides/residential-proxies/overview). The main difference is that traffic uses **datacenter** IPs instead of residential IPs. ## How Rotating Datacenter Proxies Work Rotating Datacenter Proxies use a traffic-based plan model: * Usage is billed per GB. * Your remaining bandwidth is shown in GB in the dashboard. * Traffic is consumed as you send requests through the proxy. * Pricing starts from the rate shown on your Rotating Datacenter Proxies plan. * You can configure endpoints, targeting, and session type before connecting. * You can top up bandwidth manually or enable auto top-up from the dashboard. See [Pricing](/docs/proxies/guides/datacenter-proxies/pricing) for Pay As You Go and Subscription plans. ## Access Your Credentials 1. Log in to the [Geonode Dashboard](https://app.geonode.com/). 2. Go to **Proxies → Rotating Datacenter Proxies**. 3. Open **Proxy configuration**. 4. Find your proxy credentials. Your credentials include: * **Username** * **Password** * **Host** * **Port** ## Configure Your Proxy You can configure Rotating Datacenter Proxies from the Geonode dashboard under **Proxies → Rotating Datacenter Proxies → Proxy configuration**. From Proxy Configuration, you can set: * Gateway * Country, state, and city targeting * ASN/ISP targeting * OS targeting * Protocol * Session type * Endpoint count and format * API credentials for username and password authentication The dashboard also includes **Port configuration**, **Active sticky sessions**, **Statistics**, **Block list**, and **Reseller**. See [Block List](/docs/proxies/guides/residential-proxies/block-list) to block outlier domains, IPs, or wildcards so the proxy never fetches them. Rotating Datacenter Proxy configuration ### Configuration Options | Option | Description | | --------------------- | -------------------------------------------------------- | | **Gateway** | Select the gateway region for your proxy traffic | | **Country targeting** | **Any** by default, or a specific country | | **State targeting** | **Any** by default, or a specific state or region | | **City targeting** | **Any** by default, or a specific city | | **ASN/ISP targeting** | **Any** by default, or a specific ASN/ISP when available | | **OS targeting** | Filter by operating system when available | | **Protocol** | HTTP/HTTPS or SOCKS5 | | **Session type** | Rotating or Sticky | | **Endpoints** | Number of endpoints and output format | ### Bandwidth The top of the Rotating Datacenter Proxies page shows your available bandwidth in GB. From there you can: * Check remaining bandwidth * Use **Manual top-up** to add traffic * Enable **auto top-up** * Open **See pricing** to review plan rates For full Pay As You Go and Subscription pricing, see [Pricing](/docs/proxies/guides/datacenter-proxies/pricing). ### Endpoint Format Use this format for tools that accept a single proxy string: ```text hostname:port:username:password ``` Example: ```text proxy.geonode.io:9000:USERNAME:PASSWORD ``` Replace `USERNAME` and `PASSWORD` with your dashboard credentials. ### Port Ranges For rotating sessions, Geonode uses these common port ranges: | Protocol | Port range | | -------------- | ------------- | | **HTTP/HTTPS** | `9000–9010` | | **SOCKS5** | `11000–11010` | The dashboard shows the active port range for your selected protocol under the generated endpoint list. ### Basic Connection #### HTTP ```bash curl -x http://USERNAME:PASSWORD@proxy.geonode.io:9000 https://ipinfo.io ``` #### SOCKS5 ```bash curl --proxy socks5://USERNAME:PASSWORD@proxy.geonode.io:11000 https://ipinfo.io ``` ## Targeting Rotating Datacenter Proxies support the same style of geo and network targeting as Residential Proxies: * **Country targeting** — route traffic through a specific country * **State targeting** — narrow traffic to a state or region * **City targeting** — target a specific city * **ASN/ISP targeting** — route through a specific ISP or ASN when available * **OS targeting** — filter by operating system when available For detailed targeting workflows, see: * [Target Specific Location](/docs/proxies/guides/residential-proxies/target-specific-location) * [Exclude Specific Location](/docs/proxies/guides/residential-proxies/exclude-specific-location) * [Switch Between Different Proxy Locations](/docs/proxies/guides/residential-proxies/switch-between-different-proxy-locations) ## Sessions Rotating Datacenter Proxies support both rotating and sticky sessions: * **Rotating** — a new IP is used across requests, typically over rotating ports such as HTTP `9000–9010` * **Sticky** — keep the same IP for a session when you need continuity For session workflows, see: * [Rotating Proxies](/docs/proxies/guides/residential-proxies/rotating-proxies) * [New Sticky Session](/docs/proxies/guides/residential-proxies/new-sticky-session) * [How to Release Sticky Sessions](/docs/proxies/guides/residential-proxies/release-a-sticky-session) * [How to Set Proxy Session Lifetime](/docs/proxies/guides/residential-proxies/identifying-session-lifetime) ## Endpoints and Credentials From **Proxy configuration**, you can: 1. Copy your API username and password. 2. Choose an endpoint format such as `hostname:port:username:password`. 3. Generate one or more proxy endpoints. 4. Use the generated endpoints in your tools or scripts. For code samples generated from your configuration, see [API Code Generator](/docs/proxies/guides/residential-proxies/api-code-generator). ## Monitoring Use the dashboard **Statistics** view and usage guides to monitor traffic and performance. See [Usage Stats & Analytics](/docs/proxies/guides/residential-proxies/usage-stats-analytics) for monitoring proxy usage. See [Latency Testing](/docs/proxies/guides/residential-proxies/success-latency) when you need to test response performance. ## Getting Started To start using Rotating Datacenter Proxies: 1. Open the Geonode dashboard. 2. Go to **Rotating Datacenter Proxies**. 3. Open **Proxy configuration**. 4. Set your targeting, protocol, and session type. 5. Copy your credentials and generated endpoints. If you are new to Geonode proxies, start with the [Quick Start Guide](/docs/proxies/getting-started/quick-start). # Pricing (/docs/proxies/guides/datacenter-proxies/pricing) Rotating Datacenter Proxies are billed by traffic (per GB). In the dashboard, you can choose between **Pay As You Go** and **Subscription**. Subscription plans save about **10%** compared with Pay As You Go rates for the same bandwidth tier. ## Billing Options | Option | Description | | ----------------- | ------------------------------------------------- | | **Pay As You Go** | Buy a fixed amount of bandwidth when you need it | | **Subscription** | Monthly bandwidth plans with a lower price per GB | Across both options: * Bandwidth is measured in GB. * Unused bandwidth rolls over until cancellation. * Crypto payment is available where shown in the dashboard. * Enterprise pricing is available through sales. ## Pay As You Go Pay As You Go lets you purchase bandwidth without a monthly subscription commitment. | Plan | Price per GB | Total price | Bandwidth | | -------------- | -----------: | ----------: | --------: | | **10 GB Plan** | $0.53 | $5.30 | 10 GB | | **25 GB Plan** | $0.49 | $12.25 | 25 GB | | **50 GB Plan** | $0.47 | $23.50 | 50 GB | ### Flexible Plan The **Flexible Plan** lets you choose a custom amount. | Detail | Value | | ------------------- | ----------------------------- | | **Price per GB** | From $0.53 to $0.16 | | **Minimum order** | 10 GB | | **Monthly minimum** | None | | **Bandwidth** | Rolls over until cancellation | The more bandwidth you buy, the lower the price per GB. ## Subscription Subscription plans bill monthly and include a set amount of bandwidth each month. Unused bandwidth rolls over until cancellation. | Plan | Price per GB | Monthly price | Bandwidth included monthly | | --------------- | -----------: | ------------: | -------------------------: | | **10 GB Plan** | $0.477 | $4.77 | 10 GB | | **25 GB Plan** | $0.441 | $11.03 | 25 GB | | **50 GB Plan** | $0.423 | $21.15 | 50 GB | | **100 GB Plan** | $0.396 | $39.60 | 100 GB | | **250 GB Plan** | $0.360 | $90.00 | 250 GB | | **500 GB Plan** | $0.333 | $166.50 | 500 GB | | **1 TB Plan** | $0.306 | $306.00 | 1,000 GB | | **2 TB Plan** | $0.288 | $576.00 | 2,000 GB | | **3 TB Plan** | $0.270 | $810.00 | 3,000 GB | | **5 TB Plan** | $0.252 | $1,260.00 | 5,000 GB | | **7.5 TB Plan** | $0.234 | $1,755.00 | 7,500 GB | | **10 TB Plan** | $0.225 | $2,250.00 | 10,000 GB | | **15 TB Plan** | $0.216 | $3,240.00 | 15,000 GB | | **20 TB Plan** | $0.198 | $3,960.00 | 20,000 GB | | **25 TB Plan** | $0.198 | $4,950.00 | 25,000 GB | | **30 TB Plan** | $0.189 | $5,670.00 | 30,000 GB | | **40 TB Plan** | $0.180 | $7,200.00 | 40,000 GB | | **50 TB Plan** | $0.171 | $8,550.00 | 50,000 GB | | **60 TB Plan** | $0.162 | $9,720.00 | 60,000 GB | | **75 TB Plan** | $0.153 | $11,475.00 | 75,000 GB | | **100 TB Plan** | $0.144 | $14,400.00 | 100,000 GB | Higher-volume plans have a lower price per GB. ## Choosing a Plan | If you need | Consider | | ----------------------------- | ------------------------- | | Occasional or irregular usage | Pay As You Go | | A custom bandwidth amount | Flexible Plan | | Predictable monthly usage | Subscription | | Lower price per GB at scale | Larger Subscription tiers | ## How Bandwidth Works * Your remaining bandwidth is shown in GB on the Rotating Datacenter Proxies dashboard. * Traffic is consumed as you send requests through the proxy. * You can top up bandwidth manually or enable auto top-up. * Unused bandwidth rolls over until the plan is cancelled. ## Need Custom Terms? If you need enterprise-grade solutions or a unique offer, [contact sales](https://geonode.com/contact). # Overview (/docs/proxies/guides/isp-proxies/isp-proxies) ISP Proxies provide proxy IPs associated with Internet Service Providers (ISPs). The Geonode dashboard lets you manage your assigned ISP proxy IPs, view their network details, organize them with tags and notes, and monitor usage. If you are new to proxies, see the [Proxy Service Guide](/docs/proxies/getting-started/knowledge-base/overview) to learn the basic proxy concepts before continuing. ## ISP Proxies Dashboard Open **ISP Proxies** from the **Proxies** section of the Geonode dashboard. The dashboard provides two main areas: * **Proxy List** — View and manage your assigned ISP proxy IPs. * **Statistics** — Monitor proxy usage and performance. The top section also shows information about your current ISP Proxies subscription and the number of assigned IPs. ## Subscription Overview The subscription section shows what your current ISP Proxies package includes. Depending on your account, you can see: * Total number of assigned IPs * Number of **Shared** proxies * Number of **Dedicated** proxies * Current subscription status * Subscription end or renewal information * Options to view plans, buy more IPs, or manage the subscription The number of IPs in the **Proxy List** matches your package. For example, if your plan includes **4 IPs** with **2 Shared** and **2 Dedicated** proxies, the dashboard shows **4 IPs** at the top and lists those four proxies below. Your package can include Shared IPs only, Dedicated IPs only, or a mix of both. The Shared and Dedicated counts in the summary explain why each IP appears in your list. Based on your subscription, IPs may be removed when the plan ends on a specific date. If your subscription allows it, you can re-activate the plan to restore access instead of losing the IPs permanently. For plan types and pricing, see [ISP Proxy Pricing](/docs/proxies/guides/isp-proxies/pricing). ## Proxy List The **Proxy List** contains the ISP proxy IPs assigned to your account based on your current package. ISP proxy list The proxy list provides information about each IP, including: | Field | Description | | ---------------- | ------------------------------------------------------- | | **IP** | IP address of the proxy. | | **Port** | Port used to connect to the proxy. | | **Username** | Username associated with the proxy. | | **Password** | Password associated with the proxy. | | **Country** | Country associated with the proxy IP. | | **City** | City associated with the proxy IP. | | **State** | State or region associated with the proxy IP. | | **ISP** | Internet Service Provider associated with the proxy IP. | | **Last checked** | Most recent time the proxy information was checked. | | **Created at** | Date when the proxy was created. | | **Type** | Whether the proxy is Shared or Dedicated. | | **Status** | Current status of the proxy. | | **Tags** | Tags assigned to the proxy. | ## Search Your Proxies Use the **Search** field above the proxy list to find a specific proxy. Search can help you locate an IP in a larger proxy list without manually checking each row. ## View IP Details Select a proxy from the list to open its **IP details**. IP details The IP details view shows the full information for the selected proxy, including: * IP address * Port * Country * State * City * ISP * Status (for example, Active) * **Type** — Shared or Dedicated * Tags * Notes Use **Type** to confirm whether the selected IP is **Shared** or **Dedicated**. This matches the Shared and Dedicated counts shown in your subscription summary. For example, in a package with 2 Shared and 2 Dedicated proxies, opening any IP in the list shows whether that specific IP is Shared or Dedicated. You can also add notes to the proxy. ### Add Notes Use the notes field to store information about the selected IP. For example, you can use notes to keep track of how you use a particular proxy within your own workflow. After adding or updating a note, select **Save Changes**. ### Delete or Re-activate an IP The IP details window provides a **Delete IP** option when you want to remove a selected proxy from your account manually. Separately, your package controls how long IPs stay available: * When a subscription ends on a specific date, IPs from that plan may be removed automatically. * If re-activation is available for your plan, you can re-activate the subscription to restore access to your ISP proxies. Deleting an IP removes it from your proxy list. Make sure you no longer need the proxy before deleting it. Package end dates can also remove IPs unless you renew or re-activate the subscription. ## Organize Proxies with Tags You can assign tags to your ISP proxy IPs from the proxy list. Select **add tag** for a proxy to open the Tags dialog. Proxy tags You can: * Choose a tag color. * Enter a tag name. * Save the tag. Tags can help you organize proxies according to your own workflow. For example, you could create tags based on an internal project, application, or other grouping that is useful to you. ## Export Proxy List The proxy list includes an **Export CSV** option. Use **Export CSV** to export the proxy list as a CSV file for use outside the Geonode dashboard. This can be useful when you need to work with your proxy information in another application or maintain your own records. The proxy list contains sensitive connection information, including usernames and passwords. Store exported files securely and do not share them publicly. ## Statistics Open the **Statistics** tab to view your ISP Proxy usage. The Statistics section allows you to select a time period and review your proxy activity. Available time ranges include: * **Last 24 hours** * **7 days** * **30 days** * **90 days** You can also select a custom date range. ## Usage Metrics The statistics dashboard provides an overview of your proxy activity for the selected period. The dashboard displays metrics such as: * **Total** — Total number of requests. * **Success Rate** — Percentage of successful requests. * **Average Duration** — Average request duration. * **Requests Used** — Number of requests used during the selected period. The dashboard also provides charts for request activity and request usage. ## Using Your ISP Proxies Once you have your ISP proxy IP, port, username, and password, you can configure the proxy in your application or tool. For information about proxy protocols and connection methods, see: * [Protocols](/docs/proxies/getting-started/knowledge-base/protocols) * [Proxy Endpoint Formats](/docs/proxies/getting-started/knowledge-base/endpoint-formats) * [Port Usage](/docs/proxies/getting-started/knowledge-base/proxy-usage) If you need to target a specific ISP or ASN, see [ASN/ISP Targeting](/docs/proxies/getting-started/knowledge-base/asn-isp-targeting). ## What's Next? You now know how to manage ISP Proxies from the Geonode dashboard. You can: * Understand how many Shared and Dedicated IPs your package includes. * View your assigned IPs in the proxy list. * Open IP details to confirm whether an IP is Shared or Dedicated. * Add tags and notes. * Delete IPs or re-activate a subscription when needed. * Export your proxy list. * Monitor proxy usage and performance. For plan options, see [ISP Proxy Pricing](/docs/proxies/guides/isp-proxies/pricing). For more information about using ISP and ASN targeting, see [ASN/ISP Targeting](/docs/proxies/getting-started/knowledge-base/asn-isp-targeting). # ISP Proxy Pricing (/docs/proxies/guides/isp-proxies/pricing) ISP Proxies are available in **Shared** and **Dedicated** options. The price per IP depends on the selected proxy type and the number of IPs in your plan. ## Choose Your Proxy Type When creating or changing an ISP Proxy subscription, you can choose between two proxy types. | Proxy Type | Description | | ------------------------- | ----------------------------------- | | **Shared ISP Proxies** | Shared by up to 3 users. | | **Dedicated ISP Proxies** | Exclusive access, used only by you. | The selected proxy type determines which pricing tiers apply to your subscription. ## Shared ISP Proxies Shared ISP Proxies are shared by up to 3 users. The price per IP decreases as you increase the number of IPs in your plan. ### Shared ISP Pricing Tiers | Number of IPs | Price per IP | | ------------- | -----------: | | 3–4 | $1.75 | | 5–19 | $1.70 | | 20–29 | $1.65 | | 30–39 | $1.60 | | 40–49 | $1.55 | | 50–74 | $1.50 | | 75–99 | $1.45 | | 100–149 | $1.40 | | 150–199 | $1.35 | | 200–999 | $1.30 | | 1000–10000 | $1.25 | ### Example If you select **15 Shared ISP IPs**, the dashboard shows the **5–19** pricing tier. The price is: ```text 15 × $1.70 = $25.50/month ``` The dashboard may also show the price for your current subscription alongside the price of the new subscription when you change the number of IPs. ## Dedicated ISP Proxies Dedicated ISP Proxies provide exclusive access to you. Dedicated ISP Proxies use a separate pricing structure from Shared ISP Proxies. ### Dedicated ISP Pricing Tiers | Number of IPs | Price per IP | | ------------- | -----------: | | 3–4 | $3.50 | | 5–29 | $3.00 | | 30–74 | $2.85 | | 75–299 | $2.75 | | 300–999 | $2.50 | | 1000–10000 | $2.25 | ### Example If you select **15 Dedicated ISP IPs**, the dashboard shows the **5–29** pricing tier. The price is: ```text 15 × $3.00 = $45.00/month ``` ## How Pricing Is Calculated The monthly price is based on: 1. The selected proxy type. 2. The number of IPs. 3. The price per IP for the applicable pricing tier. For example, selecting 15 IPs results in: | Proxy Type | IPs | Price per IP | Monthly Price | | ---------- | --: | -----------: | ------------: | | Shared | 15 | $1.70 | $25.50 | | Dedicated | 15 | $3.00 | $45.00 | ## Configure Your Plan When configuring an ISP Proxy subscription: 1. Select **Shared ISP Proxies** or **Dedicated ISP Proxies**. 2. Set the **Number of IPs** using the number field or the `+` and `−` controls. 3. Review the price per IP shown by the dashboard. 4. Review the monthly subscription amount. 5. Select **Proceed to checkout** to continue. The dashboard shows the current subscription and the new subscription separately when you are changing an existing plan. ## Increasing Your IP Count Increasing the number of IPs can move your subscription into a different pricing tier. For example, Shared ISP pricing changes as follows: ```text 3–4 IPs → $1.75/IP 5–19 IPs → $1.70/IP 20–29 IPs → $1.65/IP 30–39 IPs → $1.60/IP ... 1000–10000 IPs → $1.25/IP ``` The same tier-based pricing applies to Dedicated ISP Proxies using their dedicated pricing table. ## Pricing Tiers The pricing tables show the price **per IP** for each IP range. This means adding more IPs can move the subscription into a lower per-IP pricing tier. For example: ```text Shared ISP 15 IPs ↓ 5–19 tier ↓ $1.70 per IP ↓ $25.50/month ``` And: ```text Dedicated ISP 15 IPs ↓ 5–29 tier ↓ $3.00 per IP ↓ $45.00/month ``` ## Current Subscription vs New Subscription When modifying an existing subscription, the configuration screen can show both: * **Current subscription** * **New subscription** This allows you to compare the existing IP allocation and price with the configuration you are about to purchase. For example, the dashboard can show an existing subscription with 2 IPs and a new configuration with 17 IPs after adding 15 more IPs. Review the **New subscription** section before continuing to checkout. ## Choosing Between Shared and Dedicated Choose **Shared ISP Proxies** when you need ISP proxy IPs that can be shared by up to 3 users. Choose **Dedicated ISP Proxies** when you need exclusive access to the assigned proxy IPs. The two options have different prices, so select the proxy type based on your access requirements. ## What's Next? After selecting your ISP Proxy type and configuring the number of IPs, continue to checkout to create or update your subscription. Once your proxies are available, you can manage them from the [ISP Proxies Dashboard](/docs/proxies/guides/isp-proxies/isp-proxies). # Overview (/docs/proxies/guides/unlimited-residential-proxies/00_unlimited_residential_proxies) Unlimited Residential Proxies are speed-based plans designed for high-usage customers. They are billed as a monthly subscription, not per GB, and provide unlimited traffic within the limits of your selected plan. ## How Unlimited Residential Proxies Work Unlimited Residential Proxies use a speed-based plan model: * Plans are based on a speed limit such as 200 Mbps, 400 Mbps, or 1000 Mbps. * Traffic is unlimited. * Pricing is charged monthly. * Usage is not billed per GB. * Performance depends on your setup and the speed cap of your plan. Unlimited Residential currently supports the **US gateway only**. There are no thread or concurrency limits enforced by the plan. Performance depends on your setup and the speed limit of your selected plan. ## Configure Your Proxy You can configure your Unlimited Residential Proxy from the Geonode dashboard. Available configuration options include: * Gateway * Country targeting * OS targeting * Protocol * Session type and session time * Endpoints See [Configure Proxy Settings](/docs/proxies/guides/unlimited-residential-proxies/03_configure_proxy_settings) for setup and connection details. ## Plans Unlimited Residential plans are available at different speed levels. Each plan provides unlimited traffic with a different monthly price and speed limit. See [Pricing](/docs/proxies/guides/unlimited-residential-proxies/01_unlimited_residential_proxy_pricing) for the available plans and pricing. ## Monitoring To monitor requests, success rate, and data usage, see [Usage Stats & Analytics](/docs/proxies/guides/residential-proxies/usage-stats-analytics). To stop the proxy from fetching outlier domains you define, see [Block List](/docs/proxies/guides/residential-proxies/block-list). ## Getting Started To start using Unlimited Residential Proxies: 1. Open the Geonode dashboard. 2. Go to **Unlimited Residential Proxies**. 3. Open **Proxy Configuration**. 4. Configure your proxy settings and access your credentials. See [Configure Proxy Settings](/docs/proxies/guides/unlimited-residential-proxies/03_configure_proxy_settings) for the setup steps. # Pricing (/docs/proxies/guides/unlimited-residential-proxies/01_unlimited_residential_proxy_pricing) ## Available Plans Unlimited Residential Proxies use a speed-based monthly pricing model. Traffic is unlimited, and the monthly price depends on the selected speed. | Speed | Monthly Price | Traffic | | --------: | ------------: | --------- | | 200 Mbps | $1,800 | Unlimited | | 400 Mbps | $2,400 | Unlimited | | 600 Mbps | $3,000 | Unlimited | | 800 Mbps | $3,400 | Unlimited | | 1000 Mbps | $3,800 | Unlimited | Higher-speed plans provide a higher speed limit, while all listed plans provide unlimited traffic. Unlimited Residential plans are not billed per GB. Your monthly price is based on the selected speed tier. ## Choosing a Plan Choose a plan based on the throughput your workload requires. | If you need | Consider | | --------------- | -------------- | | Up to 200 Mbps | 200 Mbps plan | | Up to 400 Mbps | 400 Mbps plan | | Up to 600 Mbps | 600 Mbps plan | | Up to 800 Mbps | 800 Mbps plan | | Up to 1000 Mbps | 1000 Mbps plan | Unlimited refers to traffic volume. Your selected plan still has a defined speed limit. ## Monthly Subscription Unlimited Residential plans are billed as a monthly subscription. The plans are not charged based on the amount of data used. Instead, the monthly price is determined by the selected speed tier. ## Need Custom Terms? If you need a different duration or custom terms, [contact sales](https://geonode.com/contact) for an enterprise-grade solution. # Configure Proxy Settings (/docs/proxies/guides/unlimited-residential-proxies/03_configure_proxy_settings) To get started with Unlimited Residential Proxies, access your proxy credentials from the Geonode dashboard. ## Access Your Credentials 1. Log in to the [Geonode Dashboard](https://app.geonode.com/). 2. Go to **Residential Proxies**. 3. Open **Proxy Configuration**. 4. Find your proxy credentials. Your credentials include: * **Username** * **Password** * **Host** * **Port** ## Proxy Configuration You can configure Unlimited Residential Proxies from the Geonode dashboard under **Proxies → Unlimited Residential Proxies → Proxy configuration**. From Proxy Configuration, you can set: * Gateway * Country targeting * OS targeting * Protocol * Session type and session time * Endpoint count and format Unlimited Residential Proxy configuration ### Configuration Options | Option | Description | | --------------------- | ---------------------------------------------------------- | | **Gateway** | Currently **United States** only for Unlimited Residential | | **Country targeting** | **Any** by default, or a specific country | | **OS targeting** | Filter by operating system when available | | **Protocol** | HTTP/HTTPS or SOCKS5 | | **Session type** | Rotating or Sticky | | **Session time** | How long a sticky session keeps the same IP | | **Endpoints** | Number of endpoints and output format | Unlimited Residential currently supports the **US gateway only**. ### Endpoint Format Use this format for tools that accept a single proxy string: ```text hostname:port:username:password ``` Example: ```text residential-unlimited-us-01-proxy.geonode.io:13003:USERNAME:PASSWORD ``` Replace `USERNAME` and `PASSWORD` with your dashboard credentials. ### Basic Connection #### HTTP ```bash curl -x http://USERNAME:PASSWORD@residential-unlimited-us-01-proxy.geonode.io:13003 https://ipinfo.io ``` #### SOCKS5 ```bash curl --proxy socks5://USERNAME:PASSWORD@residential-unlimited-us-01-proxy.geonode.io:13003 https://ipinfo.io ``` ## Related Guides For detailed workflows, use the Residential Proxies guides: * [Target Specific Location](/docs/proxies/guides/residential-proxies/target-specific-location) — country and geo targeting * [Rotating Proxies](/docs/proxies/guides/residential-proxies/rotating-proxies) — rotating session setup * [New Sticky Session](/docs/proxies/guides/residential-proxies/new-sticky-session) — create sticky sessions * [How to Set Proxy Session Lifetime](/docs/proxies/guides/residential-proxies/identifying-session-lifetime) — session time * [How to Release Sticky Sessions](/docs/proxies/guides/residential-proxies/release-a-sticky-session) — release sticky sessions * [Usage Stats & Analytics](/docs/proxies/guides/residential-proxies/usage-stats-analytics) — monitor usage and statistics * [Block List](/docs/proxies/guides/residential-proxies/block-list) — block outlier domains, IPs, or wildcards # Troubleshooting Unlimited Residential Proxies (/docs/proxies/guides/unlimited-residential-proxies/06_troubleshooting) In this guide, you’ll learn how to troubleshoot common issues with Unlimited Residential Proxies and what to try when your results are slow, blocked, or failing. ## Slow Results If your results are slow, try: * Reduce local concurrency. * Release sessions to refresh IPs. * Check your local ISP or server limits. ## Target Blocks You If the target website blocks you, try: * Release or rotate the session. * Adjust country targeting. * Use sticky sessions for login or checkout flows. ## High Failure Rate If you experience a high failure rate, try: * Release sessions. * Try different country targeting. * Use fewer concurrent connections. Review your proxy configuration and statistics to identify whether the issue is related to sessions, country targeting, concurrency, or your local network setup. # API Code Generator (/docs/proxies/guides/residential-proxies/api-code-generator) The API Code Generator allows you to instantly create code snippets for different programming languages based on your configured proxy parameters.\ This helps you test, integrate, and automate API calls quickly and efficiently. *** ## Step 1: Configure Your Proxy Before generating code, set up your proxy endpoint. Follow the guide:\ ➡️ [How to Use the Endpoint Generator to Configure a Proxy](/docs/proxies/getting-started/setup_and_configuration/advance-configuration/proxy-configuration) This ensures you have the correct proxy details ready for code generation. *** ## Step 2: Access the API Code Generator Once your endpoint is created: 1. Scroll down to the **API Code Generator** section in your dashboard. 2. You will see code automatically generated based on your proxy configuration. 3. The code is available in multiple programming languages, including: * Python * Node.js * Go * And others Access the API Code Generator *** ## Step 3: Copy and Paste the Code 1. Choose your preferred language. 2. Copy the generated code snippet. 3. Paste it into your development environment or editor (e.g., VS Code, PyCharm, GoLand). *** ## Step 4: Run the Code Example: Running the Python code in **VS Code**. If your snippet uses Python’s `requests` package, install it first: ```bash pip install requests ``` Then run your file: `python app.py` *** ## Step 5: Verify the Output After running the script, check the console or terminal. The output should display the expected data from your API call, confirming that your proxy and code are properly configured. Verify the Output You’re now ready to generate, customize, and run API code seamlessly with Geonode. # Block List (/docs/proxies/guides/residential-proxies/block-list) The **Block list** lets you prevent your Geonode proxies from accessing specific domains, IPs, or wildcards. ## When to use the Block list Use the Block list when you want to prevent your proxy from accessing specific destinations, such as unwanted or background requests. For example, if your browser sends requests to `mozilla.dev` that you do not need, you can add it to the Block list. ## How it works Once a target is added to the Block list, requests to that target return **468 Target Forbidden**. Other destinations continue to work normally. ## Block list vs. Location Exclude These features control different things: | Feature | Controls | Example | | -------------------- | --------------------------------------- | ------------------- | | **Block list** | Which destinations the proxy can access | Block `mozilla.dev` | | **Location exclude** | Which proxy locations can be used | Exclude Germany | Use the **Block list** to control **where your proxy can connect** and [Location Exclude](/docs/proxies/guides/residential-proxies/exclude-specific-location) to control **which proxy locations you receive**. Image paths below are placeholders. Drop the dashboard captures into `/images/functionalities/how-to/block-list/` using the filenames in each step. *** ## Step 1: Open the Block list 1. Log in to the [Geonode Dashboard](https://app.geonode.com/). 2. Go to **Proxies → Residential Proxies**. 3. Open the **Block list** tab. Direct link: [app.geonode.com/proxies?tab=block-list](https://app.geonode.com/proxies?tab=block-list) The same tab is also available from **Rotating Datacenter Proxies** (and other products that use this proxy dashboard). If the list is empty, you will see **No data yet**. Block list tab *** ## Step 2: What you can add The input accepts: | Type | Example | Use when | | -------- | --------------- | ------------------------------- | | Domain | `mozilla.dev` | Block that host | | IP | `192.168.0.1` | Block a specific address | | Wildcard | `*.domain2.com` | Block a host and its subdomains | The placeholder in the dashboard shows the same formats: `domain1.com`, `192.168.0.1`, `*.domain2.com`. Do **not** add the site you actually need (for example your scrape target or `ip-api.com` if you use it to test). Only add outlier hosts you want the proxy to drop. *** ## Step 3: Add a single entry 1. Type a domain, IP, or wildcard in the field (for example `mozilla.dev`). 2. Click **Add entries**. 3. Confirm the row appears in the table. Wait a few seconds, then send a request (see [Verify the block](#verify-the-block)). {/* ![Add a Block list entry](/images/functionalities/how-to/block-list/add-entries.png) */} Populated Block list *** ## Step 4: Bulk add Use **Bulk add** when you have many hosts at once (trackers, helper domains, a list from a HAR file). 1. Click **Bulk add**. 2. Paste the domains, IPs, or wildcards. 3. Confirm. The new rows appear in the table. Bulk add You can also type several values in the main field (comma-separated, matching the placeholder examples) and click **Add entries**. *** ## Step 5: Search, delete, and export CSV ### Search Use **Search entries** to filter a long list by domain, IP, or wildcard. Search Block list entries ### Delete * Select one or more checkboxes, then click **Delete selected**. * Or use the trash icon on a single row. After you delete an entry, wait a few seconds. That host is allowed again. Delete selected entries ### Export CSV Click **Export CSV** to download the current list. Use it as a backup, to review entries offline, or to share the list with your team. Export Block list CSV *** ## Verify the block After `mozilla.dev` is on the list, send two requests through the same proxy: one allowed host, one blocked host. Replace `USERNAME` and `PASSWORD` with the values from **Proxy configuration**. For rotating residential HTTP, use port `9000` and a username that includes `-type-residential`. ### cURL ```bash curl -x "http://USERNAME:PASSWORD@proxy.geonode.io:9000" -I "http://ip-api.com/json" curl -x "http://USERNAME:PASSWORD@proxy.geonode.io:9000" -I "http://mozilla.dev" ``` ### Python ```python import requests username = "geonode_demouser-type-residential" password = "demopass" proxy = { "http": f"http://{username}:{password}@proxy.geonode.io:9000", "https": f"http://{username}:{password}@proxy.geonode.io:9000", } allowed = requests.get("http://ip-api.com/json", proxies=proxy, timeout=30) print(allowed.status_code, allowed.text[:200]) blocked = requests.get("http://mozilla.dev", proxies=proxy, timeout=30) print(blocked.status_code, blocked.reason) ``` ### What you should see | Target | Result | | ------------------------ | ------------------------------------------ | | `http://ip-api.com/json` | **200** — still allowed | | `http://example.com` | **200** — still allowed | | `http://mozilla.dev` | **468 Target Forbidden** | | `https://mozilla.dev` | Tunnel fails with **468 Target Forbidden** | The proxy refuses the blocked host immediately. It does not fetch the page. Platform security policies can return **464**. A host on *your* Block list returns **468 Target Forbidden**. If `mozilla.dev` still loads, wait 10–30 seconds and retry, or confirm the row is still in the table. Delete `mozilla.dev` from the list, wait, and run the same request again. It should succeed like any other allowed host. *** ## Tips * Block only outliers, not the destination you are trying to reach. * Browser sessions benefit the most: one page load can trigger many extra hosts. * After add or delete, wait a moment before testing. * Export CSV before a large cleanup so you can restore the list. * Failed Block list hits can show up as failed requests in [Usage Stats & Analytics](/docs/proxies/guides/residential-proxies/usage-stats-analytics). For other proxy error codes, see [Error Handling](/docs/proxies/api-reference/error-handling). # Exclude Specific Location (/docs/proxies/guides/residential-proxies/exclude-specific-location) import ProxyInfoFromGeonode from "../../../../snippets/get-proxy-info-from-geonode.mdx"; import RegionSpecificFAQs from "../../../../snippets/region-specific-faqs.mdx"; import SupportParagraph from "../../../../snippets/support-paragraph.mdx"; *** This guide will help you understand how to exclude proxies from specific cities, countries, states, or ISPs using the Geonode API. *** ## Prerequisites Before you begin, make sure you: * Have active Geonode proxy credentials. * Understand how to make API calls using tools like cURL or Python. *** ## What is Location Exclusion? Geonode allows you to exclude certain locations while routing traffic through proxies. This is helpful when you want to avoid specific regions due to content restrictions, compliance, or testing needs. ➡️ Want to route traffic through a location instead? [See Targeting Locations](/docs/proxies/guides/residential-proxies/target-specific-location) *** ## What Can You Exclude? You can exclude proxies based on: * Country * State * City * ASN (Autonomous System Number) You can't combine different exclusion types in a single request. For example, you can't exclude cities and ASNs together. *** ## Format for Exclusion To exclude a location, modify your proxy username like this: ```bash curl --request GET \ -x "http://proxy.geonode.io:" \ --user "-not.country-:" \ --url "http://ip-api.com/json" ``` Supported location types: * `not.country` * `not.city` * `not.state` * `not.asn` You can pass: * A single exclusion: `-not.city-tokyo` * Multiple exclusions: `-not.city-tokyo,kyoto,osaka` *** ## Get Proxy with Exclusions from Dashboard You can easily get the correct country codes, city names, state names, and ASN numbers directly from the Geonode Dashboard. Just go to the right-hand filters for Country, City, State, or ASN targeting. Once selected, they will appear in the proxy string for reference. Locations *** ## Calling API with Location Exclusions ### 1. Exclude by Country Use `-not.country-xx` or multiple like `-not.country-xx,yy,zz`. ```bash curl --request GET \ -x "http://proxy.geonode.io:" \ --user "-not.country-,:" \ --url "http://ip-api.com/json" ``` 📄 [See full API doc for country exclusion](/docs/proxies/api-reference/exclude-targeting/exclude-country) *** ### 2. Exclude by State Use `-not.state-`. ```bash curl --request GET \ -x "http://proxy.geonode.io:" \ --user "-country--not.state-:" \ --url "http://ip-api.com/json" ``` 📄 [See full API doc for state exclusion](/docs/proxies/api-reference/exclude-targeting/exclude-state) *** ### 3. Exclude by City Use `-not.city-`. ```bash curl --request GET \ -x "http://proxy.geonode.io:" \ --user "-country--not.city-:" \ --url "http://ip-api.com/json" ``` 📄 [See full API doc for city exclusion](/docs/proxies/api-reference/exclude-targeting/exclude-city) *** ### 4. Exclude by ASN (ISP) Use `-not.asn-`. You can also pass multiple values like `-not.asn-31898,12271`. ```bash curl --request GET \ -x "http://proxy.geonode.io:" \ --user "-country--not.asn-:" \ --url "http://ip-api.com/json" ``` 📄 [See full API doc for ASN exclusion](/docs/proxies/api-reference/exclude-targeting/exclude-asn) *** ## Example Response Here's an example API response when city or ASN is excluded: ```json { "status": "success", "country": "United States", "countryCode": "US", "region": "NC", "regionName": "North Carolina", "city": "Charlotte", "zip": "28202", "lat": 35.2327, "lon": -80.8461, "timezone": "America/New_York", "isp": "FiberPower LLC", "org": "FiberPower LLC", "as": "AS214483 FiberPower LLC", "query": "38.13.166.129" } ``` This shows the API successfully excluded the targeted city, and routed through an allowed location instead. *** ## Troubleshooting Tips * Make sure exclusions use correct names (e.g., "newyork" not "New York"). * City/state names should not have spaces. * Double-check that you are not mixing exclusion types in one request. * If you get `407` errors, check username/password. *** # How to Set Proxy Session Lifetime (/docs/proxies/guides/residential-proxies/identifying-session-lifetime) This guide explains how to configure session lifetime in Geonode to control how long each proxy session stays active. *** ## What Is Session Lifetime Session lifetime defines how long a proxy connection remains active before resetting.\ Setting the right duration helps to: * Maintain session consistency for account management. * Optimize proxy usage for scraping, automation, and security. * Prevent detection by avoiding overly frequent IP changes. *** ## Ways to Configure Session Lifetime You can set session lifetime in two ways: 1. Through the Geonode Dashboard (covered in this guide). 2. Via an API request → [Set Session Lifetime via API](/docs/proxies/api-reference/sticky-session/get-sticky-session/get-create). *** ## Step 1: Select a Service or Product Make sure you have an active **Residential Proxies** plan in your Geonode account.\ From the sidebar, open **Proxy Configuration** to start setting up your proxies. Service Selection *** ## Step 2: Configure Your Proxy 1. In **Proxy Configuration**, select the **Sticky** proxy type. 2. Choose a proxy protocol: **HTTP/HTTPS** or **SOCKS5**. * Changing session lifetime automatically updates the generated proxy list. * The new duration will apply to all selected sessions. Proxy Configuration *** ## Step 3: Set the Session Lifetime By default, session lifetime is **10 minutes**, but you can adjust it as needed. * **Minimum:** 3 minutes * **Maximum:** 24 hours (1440 minutes) You can enter the value in either minutes or hours, depending on your preference. Set Session Lifetime You can change this setting directly in the Dashboard or through the API. *** * Use shorter sessions for frequent IP changes (for example, scraping). - Choose longer sessions for stability in login or account-based workflows. - Adjust settings as needed to balance performance and anonymity. # New Sticky Session (/docs/proxies/guides/residential-proxies/new-sticky-session) import VerifyProxyConnectionComponent from "../../../../snippets/verify-proxy-connection-component.mdx"; import SupportParagraph from "../../../../snippets/support-paragraph.mdx"; import MonitorProxyUsage from "../../../../snippets/monitor-proxy-usage-component.mdx"; This guide explains how to create a new sticky session with Geonode. *** ## What Are Sticky Proxies Sticky proxies allow you to keep the same IP address for a set duration, making them ideal for activities that require stable connections — such as managing social media accounts, scraping sites with login sessions, or performing long-running data collection. ➡️ See the detailed explanation here: [What Are Sticky Session Proxies](/docs/proxies/guides/residential-proxies/new-sticky-session) *** ## Geonode's Sticky Proxy Ports Geonode assigns specific port ranges for sticky sessions to maintain consistent IPs over time: * **HTTP:** 10000–10900 * **SOCKS5:** 12000–12010 *** ## How to Create a Sticky Session via the Geonode API Before creating a session, configure your proxy settings in the Geonode dashboard. ### Step 1: Configure Your Proxy Settings 1. Log in to the **Geonode Dashboard**. 2. Open the **Proxies** section from the sidebar. 3. Select **Sticky** as your session type. 4. Set the session duration: * Minimum: 3 minutes * Maximum: 24 hours (1440 minutes) 5. Leave other fields as default unless you have specific preferences. Example configuration: * IP Type: *Residential* * Gateway: *France* * Port Range: *HTTP 10000–10900* * Session Type: *Sticky Session* Set Sticky Time Once configured, copy your proxy endpoint, which looks like this: ``` 92.204.164.15:10000:geonode_demouser-type-residential-lifetime-3-RAOnwR ``` Where: * IP → `92.204.164.15` * Port → `10000` * Username → `geonode_demouser-type-residential-lifetime-3-RAOnwR` * Password → `demopass` You can use this configuration for API access or in third-party tools such as Chrome extensions. *** ### Step 2: Create a Sticky Session via API To start a sticky session through the API, make a request to the `:` endpoint using your configured username and password. Example using **cURL**: ```bash curl -x 92.204.164.15:10000 \ --user "geonode_demouser-type-residential-lifetime-3-RAOnwR:demopass" \ --url "http://ip-api.com/json" \ --header "Accept: application/json" ``` Follow this Sticky session API Guide to learn more about the API [Create a New Sticky Session](/docs/proxies/api-reference/sticky-session/get-sticky-session/get-create) *** *** *** *** ## FAQs {" "} {" "} Sticky sessions maintain the same IP for a defined period, while rotating proxies assign a new IP with every request.{" "} {" "} The session duration depends on your configuration — from 3 minutes to 24 hours. You set this when creating the session.{" "} {" "} No, session parameters like IP and type cannot be changed once created. You can, however, release the current session and create a new one with updated settings.{" "} {" "} # Overview (/docs/proxies/guides/residential-proxies/overview) Residential Proxies give you access to real residential IPs for browsing, automation, and data collection. They are billed pay-as-you-go by traffic (per GB), and you configure targeting, protocol, and session type from the Geonode dashboard. ## How Residential Proxies Work Residential Proxies use a traffic-based plan model: * Usage is billed per GB. * Your remaining balance is shown in GB in the dashboard. * Traffic is consumed as you send requests through the proxy. * Pricing starts from the rate shown on your Residential Proxies plan. * You can configure endpoints, targeting, and session type before connecting. ## Configure Your Proxy You can configure your Residential Proxy from the Geonode dashboard under **Proxies → Residential Proxies → Proxy configuration**. Available configuration options include: * IP type set to Residential * Gateway selection * Country, state, and city targeting * ASN/ISP targeting * OS targeting * Protocol selection such as HTTP/HTTPS * Session type such as Rotating or Sticky * Endpoint count and format generation * API credentials for username and password authentication The dashboard also includes **Port configuration**, **Active sticky sessions**, **Statistics**, **Block list**, and **Reseller**. See [Block List](/docs/proxies/guides/residential-proxies/block-list) to block outlier domains, IPs, or wildcards so the proxy never fetches them. ## Targeting Residential Proxies support granular geo and network targeting: * **Country targeting** — route traffic through a specific country * **State targeting** — narrow traffic to a state or region * **City targeting** — target a specific city * **ASN/ISP targeting** — route through a specific ISP or ASN * **OS targeting** — filter residential devices by operating system See [Target Specific Location](/docs/proxies/guides/residential-proxies/target-specific-location) for geo-targeting workflows. See [Exclude Specific Location](/docs/proxies/guides/residential-proxies/exclude-specific-location) when you need to avoid certain locations. ## Sessions Residential Proxies support both rotating and sticky sessions: * **Rotating** — a new IP is used across requests, typically over rotating ports such as HTTP `9000–9010` * **Sticky** — keep the same IP for a session when you need continuity See [Rotating Proxies](/docs/proxies/guides/residential-proxies/rotating-proxies) to use rotating endpoints. See [New Sticky Session](/docs/proxies/guides/residential-proxies/new-sticky-session) and [How to Release Sticky Sessions](/docs/proxies/guides/residential-proxies/release-a-sticky-session) for sticky session workflows. ## Endpoints and Credentials From **Proxy configuration**, you can: 1. Copy your API username and password. 2. Choose an endpoint format such as `hostname:port:username:password`. 3. Generate one or more proxy endpoints. 4. Use the generated endpoints in your tools or scripts. For code samples generated from your configuration, see [API Code Generator](/docs/proxies/guides/residential-proxies/api-code-generator). ## Monitoring Use the dashboard **Statistics** view and usage guides to monitor traffic and performance. See [Usage Stats & Analytics](/docs/proxies/guides/residential-proxies/usage-stats-analytics) for monitoring residential proxy usage. See [Latency Testing](/docs/proxies/guides/residential-proxies/success-latency) when you need to test response performance. ## Getting Started To start using Residential Proxies: 1. Open the Geonode dashboard. 2. Go to **Residential Proxies**. 3. Open **Proxy configuration**. 4. Set your targeting, protocol, and session type. 5. Copy your credentials and generated endpoints. If you are new to Geonode proxies, start with the [Quick Start Guide](/docs/proxies/getting-started/quick-start). # How to Release Sticky Sessions (/docs/proxies/guides/residential-proxies/release-a-sticky-session) This guide explains how to release sticky sessions in Geonode to refresh your proxy connection, switch to a new IP, or free up resources. *** ## Why Release a Sticky Session Releasing a sticky session ends your current connection and starts a new one.\ This helps when you need to: * Switch to a different proxy IP. * Free up resources such as ports or threads. * Improve security by refreshing your active session. *** ## How to Release Selected Sticky Sessions Follow these steps to release specific sticky sessions from your Geonode Dashboard. ### Step 1: Select the Sessions to Release 1. Go to the **Active Sticky Sessions** section in your Geonode Dashboard. 2. Select the sessions you want to release — you can select multiple at once. Select Sessions *** ### Step 2: Release Selected Sessions 1. After selecting sessions, click **Release Selected**. 2. Confirm by clicking **Release Sessions**. 3. A success notification will appear once the sessions are released. Release Sessions *** ## How to Release All Sticky Sessions To disconnect all active sessions at once: 1. Click **Release All Sessions** in the Dashboard. 2. Confirm by selecting **Release All**. 3. All active connections will be immediately terminated. Release All *** * If a sticky session shows slow performance or fails to connect, release it to get a new IP. - Use **Release Selected** for specific sessions and **Release All** for bulk actions. - Refresh sticky sessions periodically to maintain security and performance. # Rotating Proxies (/docs/proxies/guides/residential-proxies/rotating-proxies) import VerifyProxyConnectionComponent from "../../../../snippets/verify-proxy-connection-component.mdx"; import SupportParagraph from "../../../../snippets/support-paragraph.mdx"; import ExtensionFAQs from "../../../../snippets/extensions-faqs.mdx"; This guide walks you through how to use Rotating Proxies in Geonode. *** ## What Are Rotating Proxies Rotating proxies automatically change your IP address with every request.\ This feature helps maintain anonymity, prevent detection, and ensure smoother automation. ➡️ For a detailed explanation, see [What Are Rotating Proxies](/docs/proxies/guides/residential-proxies/rotating-proxies) *** ## Geonode's Rotating Proxies Range Geonode provides rotating proxies via HTTP and SOCKS5 protocols, with the following ports: * **HTTP:** 9000–9010 * **SOCKS5:** 11000–11010 Each request you send will use a new IP address from a different location, helping to avoid IP-based blocking. *** ## How to Use Rotating Proxies with the Geonode API ### Step 1: Configure Your Proxy Settings In your Geonode dashboard: 1. Open **Proxy Configuration**. 2. Select: * IP Type: *Residential* * Gateway: *France* (or your preferred location) * Protocol: *HTTP/HTTPS* or *SOCKS5* * Session Type: *Rotating* Example configuration: Endpoints *** ### Step 2: Generate Proxy Endpoints Once configured, generate your proxy endpoints.\ They will typically look like this: ``` 92.204.164.15:9000:geonode_demouser-type-residential:demopass ``` Each endpoint includes your IP, port, and authentication credentials (username and password). *** ### Step 3: Make API Calls Using Rotating Proxies Here’s a Python example using the `requests` library: ```python import requests username = "geonode_demouser-type-residential" password = "demopass" GEONODE_DNS = "92.204.164.15:9000" url = "http://ip-api.com" proxy = { "http": f"http://{username}:{password}@{GEONODE_DNS}" } response = requests.get(url, proxies=proxy) print("Response:\n", response.text) ``` Each request will use a different IP address from Geonode’s proxy pool. *** *** *** ## FAQs {" "} {" "} Yes. You can add multiple proxies and switch between them by selecting the desired one and clicking **Connect**.{" "} {" "} Rotating proxies: - Maintain anonymity by changing IPs per request. - Prevent blocks tied to a single IP. - Work perfectly for scraping, SEO, and geo-restricted access.{" "} {" "} * **HTTP:** Best for browsing, scraping, and API requests (HTTP/HTTPS traffic). * **SOCKS5:** More flexible, supporting FTP, VoIP, and P2P connections.{" "} {" "} Yes, you can configure rotating proxies to target countries, cities, or ISPs to get region-specific IPs.{" "} {" "} If one proxy IP is blocked, Geonode automatically rotates to a new IP from the pool on the next request.{" "} {" "} Absolutely. They’re ideal for scraping since each request uses a new IP, reducing the risk of detection.{" "} {" "} You can track requests, performance, and errors directly from your Geonode dashboard.{" "} {" "} Usually, yes. Rotating proxies provide better anonymity and resistance to bans, while static proxies are easier to track and block.{" "} {" "} # Latency Testing (/docs/proxies/guides/residential-proxies/success-latency) This guide explains how to test the latency and success rate of your proxies using Python.\ You’ll learn how to measure proxy performance, analyze results, and visualize latency data. *** ## Overview The provided script uses Python’s `requests` library to send concurrent requests and measure performance metrics like: * **Latency (response time)** * **Success rate (status code analysis)** * **Error types (timeouts, connection issues)** It supports both **SOCKS5** and **HTTPS** proxies and uses [ip-api.com](http://ip-api.com) as the default target URL.\ You can also test alternative endpoints like **Cloudflare trace** for comparison. *** ## What You’ll Learn * How to run latency tests on multiple proxies. * How to configure proxy protocols (SOCKS5 or HTTP). * How to analyze average latency, median response time, and success rates. * How to interpret graphical results to detect issues. *** ## Default Settings | Parameter | Default Value | Description | | -------------- | ------------------------------- | ------------------------------------------ | | Protocol | HTTPS | Set `use_socks5 = True` for SOCKS5 proxies | | Target | [ip-api.com](http://ip-api.com) | Default endpoint for IP tests | | Concurrency | 5–10 threads | Recommended thread range | | Testing Volume | 2000+ requests | Recommended for stable averages | *** ## Variable Settings The script allows you to adjust several key variables: * **Target Country** – optional (leave blank to use random pool) * **Protocol** – choose between HTTP/HTTPS or SOCKS5 * **Session Type** – rotating (default) * **Target URL** – endpoint to test (default: ip-api.com) * **Number of Requests** – define test size * **Threads** – set concurrent workers (recommended: 5–10) Results may vary depending on which gateway you use. Geonode currently provides three gateway locations. *** ## Port Ranges | Session Type | Protocol | Port Range | Description | | ------------ | ---------- | ----------- | ---------------------------------- | | Rotating | HTTP/HTTPS | 9000–9010 | Changes IP for every request | | Rotating | SOCKS5 | 11000–11010 | Changes IP for every request | | Sticky | HTTP/HTTPS | 10000–10900 | Keeps same IP for session duration | | Sticky | SOCKS5 | 12000–12010 | Keeps same IP for session duration | *** ## Steps to Test Proxy Latency ### Step 1: Install Required Libraries Install dependencies before running the script: ```bash pip install requests matplotlib numpy ``` ### Libraries used * `requests` — send HTTP requests * `matplotlib` — visualize latency results * `numpy` — calculate average and median latency * `collections` — count occurrences of errors *** ## Step 2: Configure Proxy Settings You can test **SOCKS5** or **HTTP(S)** proxies by adjusting the `use_socks5` flag. ### SOCKS5 Proxy Example ```python use_socks5 = True proxies = { 'http': 'socks5://username:password@proxy.geonode.io:11009', 'https': 'socks5://username:password@proxy.geonode.io:11009' } ``` ### HTTP Proxy Example ```python use_socks5 = False proxies = { 'http': 'http://username:password@proxy.geonode.io:9008', 'https': 'http://username:password@proxy.geonode.io:9008' } ``` Replace `username` and `password` with your Geonode credentials ## Step 3: Configure Test Parameters Set your test parameters in the script: ```python url = 'http://ip-api.com' # Default test target num_requests = 5000 # Number of requests num_workers = 5 # Concurrent threads ``` For accuracy: * Use **2000+ requests** * Run with **5–10 threads** *** ## Step 4: Send Concurrent Requests The script uses `ThreadPoolExecutor` to send multiple requests simultaneously: ```python from concurrent.futures import ThreadPoolExecutor import requests, time def fetch_url(i): try: start_time = time.time() response = requests.get(url, proxies=proxies, timeout=60) latency = time.time() - start_time return response.status_code, latency except requests.exceptions.Timeout: return 'Timeout', time.time() - start_time except requests.exceptions.RequestException: return 'Error', time.time() - start_time ``` Each request logs: * **Status code** * **Latency (seconds)** * **Timeouts or errors** *** ## Step 5: Analyze and Visualize Results After completing all requests, the script plots a latency graph. | Color | Meaning | | ----------- | -------------------------------- | | **Blue** | Successful requests (status 200) | | **Red** | Non-200 responses (404, 500) | | **Green** | Timeouts | | **Magenta** | Other errors | *** ## Step 6: View Test Statistics Once the test completes, the script outputs key metrics: ``` Average Latency: 1.20 seconds Median Latency: 0.89 seconds Standard Deviation of Latency: 1.30 seconds Status Code Percentages: 200: 99.54% Error: 0.12% 401: 0.04% 402: 0.06% 500: 0.20% 502: 0.04% Total Requests: 5000 Successful (Status 200): 4977 Timeouts: 0 Other Errors: 6 Error Messages: Error: 6 ``` These values help assess **reliability**, **consistency**, and **stability** of your proxy connections. *** ## Step 7: Interpret Results Use the output to evaluate proxy quality: | Metric | Meaning | | ---------------------- | --------------------------------------- | | **Success Rate** | Percentage of status 200 responses | | **Latency** | Average response time per request | | **Error Distribution** | Frequency of timeouts or other failures | Rotating ports provide more accurate, diversified benchmarks than testing a single static proxy. *** ## Source Code You can find the full source code for this script on GitHub: [Geonode Proxy Testing Toolkit](https://github.com/geonodecom/proxy-testing-toolkit/tree/performance-testing/success-latency) ``` ``` # Switch Between Different Proxy Locations (/docs/proxies/guides/residential-proxies/switch-between-different-proxy-locations) import BrowsersFaqs from "../../../../snippets/browsers-faqs.mdx"; import VerifyProxyConnectionComponent from "../../../../snippets/verify-proxy-connection-component.mdx"; import SupportParagraph from "../../../../snippets/support-paragraph.mdx"; This guide explains how to switch between different proxy locations — for example, using proxies from various countries to access region-specific websites or services. *** ## Configure Proxy Based on Location To switch proxies by country, state, or city, follow the detailed setup instructions here: ➡️ [How to Target Specific Location with Proxies](/docs/proxies/guides/residential-proxies/target-specific-location) That guide explains how to choose and configure a proxy for a specific region, helping you set the right routing for your use case. *** ## Geonode Proxy Manager (Chrome Extension) If you use Google Chrome, Geonode provides a dedicated browser extension to manage and switch proxies quickly. ➡️ [How to Use the Geonode Chrome Extension for Proxy Management](/docs/proxies/getting-started/setup_and_configuration/extensions/geonode-proxy-manager) The extension lets you: * Switch between multiple proxy locations easily. * Save frequently used proxies. * Test and verify active connections in one click. *** ## FoxyProxy (Alternative Browser Extension) If you prefer a third-party tool, **FoxyProxy** is a popular extension that allows switching proxies by location or rule-based logic. ➡️ [How to Set Up Proxy in FoxyProxy](/docs/proxies/getting-started/setup_and_configuration/extensions/foxyproxy) FoxyProxy supports: * Multiple proxy profiles * Location-based routing * One-click switching between proxies You can use any other proxy management extension as long as it supports location-based switching. *** ## Rotating Proxy If you don’t need a specific location and prefer automatic IP rotation, use Geonode’s **Rotating Proxy Service**.\ It automatically assigns a new IP address from a different location for each request — ideal for tasks like scraping or multi-region testing. ➡️ [How to Set Up Rotating Proxies](/docs/proxies/guides/residential-proxies/rotating-proxies) *** *** *** # Target Specific Location (/docs/proxies/guides/residential-proxies/target-specific-location) import ProxyInfoFromGeonode from "../../../../snippets/get-proxy-info-from-geonode.mdx"; import RegionSpecificFAQs from "../../../../snippets/region-specific-faqs.mdx"; import SupportParagraph from "../../../../snippets/support-paragraph.mdx"; *** This guide will help you understand how to geo-target proxies based on locations like cities, countries, states, and ISPs using the Geonode API.. *** ## Prerequisites Before you begin, ensure that you: * Have valid Geonode API credentials. * Have basic knowledge of making API requests (e.g., using cURL or Python). *** ## What is Geo-Targeting? Geonode's **geo-targeting** capabilities allow you to route requests through proxies located in specific countries, cities, states, or regions. This enables precise location-based proxy management for your needs. ➡️ **Learn more about Geo-Targeting**: [What is Geo-Targeting](/docs/proxies/getting-started/knowledge-base/geo-targeting) *** ## Ways to Target Specific Locations in Geonode Geonode allows you to target proxies based on: * **Country** * **City** * **State** * **ISP/ASN** You can't target both **state** and **city** at the same time. *** ## API Endpoint Structure for Geo-Targeting Geonode's geo-targeting endpoints are designed to accept location parameters such as city, state, country, or ISP/ASN. Simply append the specific parameter (e.g., `-country-`) after the Geonode username. ```bash curl --request GET \ -x "http://proxy.geonode.io:" \ --user "-country-:" \ --url "http://ip-api.com/json" ``` Replace the placeholder values with your Geonode username, password, and desired location. *** ## Generating User Credentials Based on Location Since the location needs to be included in the username, Geonode makes it easy to generate a location-based proxy username via the Geonode Dashboard: * Go to the Geonode Dashboard. * Scroll down to Proxy Configuration. * On the right, you'll find various location targeting options. * Once you configure the location, you'll see the updated username, which you can use in API calls for geo-targeted requests. Locations *** ## Calling Geo-Targetting API To target a specific city, state, or country using the Geonode API, use the following endpoint formats: ### Geo-Targeting by Country To target a proxy by **country**, append `-country-` to the username. **Example** ```bash curl --request GET \ -x "http://proxy.geonode.io:" \ --user "-country-:" \ --url "http://ip-api.com/json" ``` Check out the detailed API docs here ---> [Perform Country Targeting](/docs/proxies/api-reference/geo-targeting/get-country) *** ### Geo-Targeting by State To target a proxy by **state**, append `-state-` to the username. **Example** ```bash curl --request GET \ -x "http://proxy.geonode.io:" \ --user "-country--state-:" \ --url "http://ip-api.com/json" ``` Check out the detailed API docs here ---> [Perform State Targeting](/docs/proxies/api-reference/geo-targeting/get-state) *** ### Geo-Targeting by City To target a proxy by **city**, append `-city-` to the username. **Example** ```bash curl --request GET \ -x "http://proxy.geonode.io:" \ --user "-country--city-:" \ --url "http://ip-api.com/json" ``` Check out the detailed API docs here ---> [Perform City Targeting](/docs/proxies/api-reference/geo-targeting/get-city) *** ### Geo-Targeting by ISP/ASN Geonode also allows you to target proxies based on the **ASN (Autonomous System Number)** of an ISP. To do this, append `-asn-` after specifying the country. **Example** ```bash curl --request GET \ -x "http://proxy.geonode.io:" \ --user "-country--asn-:" \ --url "http://ip-api.com/json" ``` Check out the detailed API docs here ---> [Perform ISP/ASN Targeting](/docs/proxies/api-reference/geo-targeting/get-isp) *** ## API Response The API will respond with detailed information about the targeted location, including: * **Region**, **City**, **ISP**, **Latitude**, **Longitude**, and more. **Example Response:** ```json { "status": "success", "city": "New York", "region": "New York", "country": "United States", "latitude": "40.7128", "longitude": "-74.0060", "isp": "ISP Example", "timezone": "America/New_York", "postal_code": "10001" } ``` This shows that the request was routed through the proxy located in **New York**, with corresponding geographical details. *** ## Troubleshooting If you encounter issues when targeting a specific location: * Double-check that the city, state, or country name is spelled correctly. * Ensure that the proxy is available and accessible in the targeted location. * Verify the authentication credentials if you receive errors related to access or authorization. *** # UDP over SOCKS5 for Residential Proxies (/docs/proxies/guides/residential-proxies/udp_over_socks5) UDP support is available for **Geonode Residential Proxies** when using the **SOCKS5** protocol. This guide explains how to enable UDP support, configure your proxy, understand how UDP works over SOCKS5, and verify your setup using a simple DNS test. ## Availability UDP support is available for **Geonode Residential Proxies** when using the SOCKS5 protocol. To enable UDP support, the proxy username must include the following flag: ```text -requireUdp-true ``` Use your normal proxy password. Do not change the password format. ## Configure Your Proxy Use the following proxy settings: ```text Protocol: SOCKS5 Host: proxy.geonode.io Port: 12000 Username: -requireUdp-true Password: ``` Example: ```text socks5://-requireUdp-true:@proxy.geonode.io:12000 ``` ## How UDP Works Geonode supports **UDP over SOCKS5** using the standard SOCKS5 `UDP ASSOCIATE` command. This is not a direct raw UDP connection between the client and the destination server. Instead, the client first establishes a standard SOCKS5 TCP connection with the proxy. After authentication, the client requests the proxy to create a UDP relay. The communication flow is: ```text 1. Client connects to the Geonode SOCKS5 proxy over TCP. 2. Client authenticates using the proxy username and password. 3. Client sends a SOCKS5 UDP ASSOCIATE request. 4. The proxy returns a UDP relay IP address and port. 5. Client sends UDP packets to the relay. 6. The proxy forwards the UDP packets to the destination server. 7. UDP responses return through the relay back to the client. ``` > The TCP control connection must remain open while UDP traffic is active. If the TCP connection closes, the UDP association may also terminate. ## Test UDP with DNS DNS is one of the simplest ways to verify UDP support because standard DNS queries commonly use UDP port `53`. You can use any SOCKS5 client, library, or tool that supports the SOCKS5 `UDP ASSOCIATE` command. > Some SOCKS5 clients only support TCP proxying. Those tools can verify SOCKS5 authentication but cannot verify UDP support. For example, send a DNS query through the SOCKS5 UDP relay: ```text DNS Server: 1.1.1.1 Port: 53 Domain: example.com ``` If a DNS response is returned, UDP over SOCKS5 is working correctly. ## DNS Test Workflow A typical DNS test follows these steps: ```text 1. Open a SOCKS5 TCP connection to proxy.geonode.io:12000. 2. Authenticate using the proxy username and password. 3. Send a SOCKS5 UDP ASSOCIATE request. 4. Receive a UDP relay IP address and port. 5. Send a UDP DNS query through the relay to 1.1.1.1:53. 6. Confirm that a DNS response is received. ``` ## Expected Output A successful test should produce results similar to the following: ```text SOCKS5 TCP connection: successful SOCKS5 authentication: successful UDP ASSOCIATE: accepted UDP relay returned: : UDP DNS query sent to: 1.1.1.1:53 UDP response: received DNS result: answer received for example.com Result: UDP over SOCKS5 is working. ``` ## Understanding the Results ```text SOCKS5 authentication: successful ``` The proxy username and password are valid. ```text UDP ASSOCIATE: accepted ``` The proxy accepted the SOCKS5 `UDP ASSOCIATE` request and created a UDP relay. ```text UDP response: received ``` The UDP packet was successfully forwarded through the proxy relay, reached the DNS server, and the response was returned to the client. ```text DNS result: answer received ``` The DNS server successfully resolved the requested domain. ## Troubleshooting If authentication fails: ```text Authentication failed SOCKS5 username/password auth failed status=1 ``` Verify the username, password, service access, and account status. Also confirm that you are using valid **Residential Proxy** credentials. If the UDP ASSOCIATE request fails: ```text UDP ASSOCIATE failed Command not supported Connection not allowed by ruleset SOCKS5 reply code 7 SOCKS5 reply code 2 ``` Confirm that you are using: * Residential Proxies * SOCKS5 * Port `12000` * The `-requireUdp-true` username flag If the UDP request times out: ```text timed out UDP response timeout No UDP response received ``` The SOCKS5 connection may have succeeded, and the proxy may have returned a UDP relay, but the UDP packet or response did not complete. For Geonode Residential Proxies, confirm that your username includes: ```text -requireUdp-true ``` If the flag is missing, authentication may still succeed, but the UDP relay test can time out because the session was not routed through the UDP-enabled path. Timeouts may also occur because of firewall rules, local network restrictions, destination-side filtering, or an incorrect proxy configuration. ## Summary To use UDP with Geonode Residential Proxies, use the **SOCKS5** protocol on port **12000** and include the `-requireUdp-true` flag in your proxy username. Geonode implements UDP support through the SOCKS5 `UDP ASSOCIATE` command, allowing the proxy to establish a UDP relay after successful authentication. A DNS query sent through the SOCKS5 proxy is the simplest way to verify that UDP has been configured correctly. # Usage Stats & Analytics (/docs/proxies/guides/residential-proxies/usage-stats-analytics) This guide explains how to track, analyze, and optimize your proxy usage using the Geonode dashboard.\ Monitoring usage statistics helps you manage bandwidth efficiently, improve performance, and troubleshoot issues early. *** ## Step 1: Access the Geonode Dashboard 1. Log in to your **Geonode account**. 2. Open the **Dashboard** to see an overview of your key information: * Available bandwidth * Billing cycle and renewal dates Geonode Dashboard *** ## Step 2: View Usage Statistics In the **Statistics** section, you can explore detailed proxy performance metrics. ### Request Metrics A visual overview of your request activity, including: * **Successful Requests** – Proxies that connected successfully. * **Failed Requests** – Requests that encountered errors or blocks. * **Total Requests Sent** – The complete count of requests made through your proxies. Request Metrics *** ### Data Usage Monitoring Keep track of your bandwidth and usage trends: * **Daily Bandwidth Consumption** – See how much data is used each day. * **Hourly Usage Trends** – Identify peak usage hours to optimize resources and balance workloads. This information helps you spot inefficiencies and adjust configurations for better proxy performance. Hourly Data Usage *** ## Final Tips * Check usage regularly to ensure efficient proxy operation. * Review failed requests to identify network or configuration issues. * Use hourly and daily analytics to optimize your bandwidth strategy. By leveraging Geonode’s analytics tools, you can maintain stable connections, reduce errors, and optimize your proxy performance. # How to Use the Secure EU Gateway (France – Whitelisted IP Only) (/docs/proxies/guides/residential-proxies/use-the-secure-eu-gateway) To ensure maximum uptime and performance, Geonode provides a secure EU gateway available only to paying users.\ This gateway offers enhanced reliability and protection through IP whitelisting. *** ## Gateway Details * **IP Address:** 92.204.164.13 * **Hostname:** prod-proxy.geonode.io * **Location:** France * **Name:** France (Whitelisted IP Only) * **Access Type:** IP-whitelisted access for paid users only *** ## How to Use 1. Log in to your **Geonode Dashboard**. 2. Go to the **IP Whitelist** section. 3. Add your current IP address. 4. In your proxy configuration, select the gateway location: *France (Whitelisted IP Only)*. 5. Save your configuration and connect through the new gateway. *** Access will be denied for any requests coming from non-whitelisted IP addresses. This gateway is recommended for users who need higher reliability and tighter access control. # Crawl Overview (/docs/scraper-api/dashboard-guides/crawl/00_crawl_overview) The Crawl dashboard provides a workspace for crawling a website, reviewing recent crawl jobs, and monitoring crawl usage. Crawl dashboard ## Account Overview At the top of the Crawl dashboard, you can view your current request usage and plan. The account overview shows: * **Available Requests** — Number of requests currently available, or Unlimited depending on your plan. * **Threads in use** — Current thread usage for your plan. * **Used last 24 hours** — Number of requests used during the last 24 hours. * **Plan** — Your current subscription plan and upgrade options. You can also use **See plans** to review available pricing. ## Crawl Workspace The main Crawl workspace is where you enter the website URL you want to crawl. Enter the URL in the field and use: * **Settings** to configure the crawl. * **API & Integrations** to access API-related options. * **Start Crawling** to submit the crawl. For information about configuring and starting a crawl from the dashboard, see [Crawl a Site](/docs/scraper-api/dashboard-guides/crawl/01_crawl_a_site). ## Recent Crawls The **Recent Crawls** section displays your previous crawl jobs. Each crawl displays: * **Status** — Current status of the crawl job. * **URL** — Starting URL that was crawled. * **Pages** — Number of pages crawled, such as `5 / 5`. * **Created** — Date and time when the crawl was created. * **Execution Time** — Time taken to complete the crawl. * **Output** — Output format, such as Markdown, with a download action. You can also: * Search previous jobs by URL. * Filter jobs by status. * Change the number of rows displayed per page. * Move between result pages. * Download crawl output from completed jobs. ## Crawl Status Crawl jobs can appear with different statuses depending on their current state. Use the **Status** filter to narrow the list of recent crawls by status. ## Statistics The Crawl dashboard also includes a **Statistics** tab for viewing crawl usage. Open **Statistics** to view your crawl activity and usage information. For more information about managing crawl jobs and monitoring usage, see [Manage Crawl Jobs](/docs/scraper-api/dashboard-guides/crawl/02_manage_crawl_jobs). ## What's Next? You now know the main sections of the Crawl dashboard. Continue to [Crawl a Site](/docs/scraper-api/dashboard-guides/crawl/01_crawl_a_site) to learn how to configure a crawl and start crawling from the dashboard. # Crawl a Site (/docs/scraper-api/dashboard-guides/crawl/01_crawl_a_site) Use the Crawl dashboard to enter a starting URL, configure crawl options, and start crawling a website. ## Enter the Website URL Open **Crawl a site** from the Web Data section of the dashboard. Enter the starting URL in the crawl field. This is the seed URL where the crawl begins. ## Configure Your Crawl Open **Settings** to configure crawl options before starting. Crawl settings Available settings include: | Setting | Description | | ----------------------- | ---------------------------------------------------------------------------- | | **Page limit** | Maximum number of pages to crawl. Maximum value is `10000`. | | **Depth** | How many link levels to follow from the starting URL. Maximum value is `10`. | | **Stay on same domain** | Keep the crawl limited to the same domain as the starting URL. | | **Include subdomains** | Allow crawling pages on subdomains of the starting domain. | | **Output format** | Format used for crawl output, such as Markdown. | | **JS Rendering** | Enable JavaScript rendering when the target pages require it. | | **Proxy Type** | Proxy network used for the crawl, such as Residential. | | **Proxy Country** | Optional country targeting when the target has geo-restrictions. | For detailed information about crawl parameters and how they affect your request, see [Configuring Crawl Requests](/docs/scraper-api/guides/crawl/02_configuring-crawl-requests). ## Start Crawling After entering the URL and configuring the available settings, select **Start Crawling**. The crawl is submitted as a job and appears in the **Recent Crawls** section. ## Review the Crawl Job Once the crawl is running or completed, find it in **Recent Crawls**. Each completed crawl shows: * Status * Starting URL * Pages crawled * Created time * Execution time * Output format Use the download action in the **Output** column to download the crawl results. ## What's Next? You now know how to configure a crawl, start crawling, and download results from the dashboard. Continue to [Manage Crawl Jobs](/docs/scraper-api/dashboard-guides/crawl/02_manage_crawl_jobs) to learn how to search, filter, and manage your previous crawls. # Manage Crawl Jobs (/docs/scraper-api/dashboard-guides/crawl/02_manage_crawl_jobs) The **Crawl dashboard** lets you manage previous crawls and monitor your crawl activity from one place. ## Recent Crawls Open the **Recent Crawls** tab to view your previous crawl jobs. Recent crawls The Recent Crawls section provides: * **Search by URL** — Find a previous crawl by its starting URL. * **Status** — Filter crawls by their current status. * **URL** — View the starting URL used for each crawl. * **Pages** — See how many pages were crawled. * **Created** — See when the crawl was created. * **Execution Time** — See how long the crawl took to complete. * **Output** — View the output format and download the results. You can also use the pagination controls to move between pages of crawl jobs and change the number of rows displayed per page. ## Filter Crawl Jobs Use **Search by URL** to find a specific crawl. For example, you can enter: ```text docs.geonode.com ``` to find crawls for that domain. You can also use the **Status** filter to narrow the list of crawls by their current status. ## Download Crawl Output Each completed crawl job includes a download action in the **Output** column. Select the download icon to export the crawl results for that job. ## Statistics The **Statistics** tab provides an overview of your Crawl API usage. Open **Statistics** to review crawl activity for a selected time period. You can typically: * Select a custom date range. * Choose a predefined time range, such as the last 24 hours, 7 days, 30 days, or 90 days. * View the total number of crawls. * Monitor your success rate. * Review the average crawl duration. * See the number of requests used during the selected period. ## What's Next? You now know how to review previous crawls, filter crawl jobs, download output, and monitor Crawl API usage. For the API-level details of crawl jobs, see [Managing Crawl Jobs](/docs/scraper-api/guides/crawl/04_managing-crawl-jobs). # Scraper Overview (/docs/scraper-api/dashboard-guides/extraction/00_extractor_overview) The Scraper dashboard lets you extract content from web pages without writing code. It is organized into three main sections that help you monitor your account, create extraction requests, and review previous jobs. This guide introduces each section of the dashboard. The next guides walk you through creating extraction requests and managing your jobs. Scraper Dashboard Overview ## 1. Account Overview The top section provides information about your account and subscription. Here you can: * View the number of available requests. * Check your free request allowance. * See when your request quota renews. * Review the number of requests used during the last 24 hours. * Upgrade your plan if you need additional requests or higher usage limits. This section gives you a quick overview of your current usage before submitting new requests. *** ## 2. Extraction Workspace The center section is where you create new extraction requests. From this workspace, you can: * Enter the URL you want to extract. * Switch between single and multiple URL extraction. * Open the **Settings** panel to configure your request. * Access **API & Integrations**. * Start a new extraction. This guide only introduces the workspace. The extraction workflows are covered in the following guides: * Extract Content from a Single URL * Extract Content from Multiple URLs For a detailed explanation of each extraction option, refer to the Extraction Guides. *** ## 3. Recent Jobs The bottom section displays your previous extraction requests. From here, you can: * View recent extraction jobs. * Search previous jobs by URL. * Filter jobs by status and output format. * Review execution times. * Open completed extraction results. Detailed job management is covered in the **Manage Extraction Jobs** guide. *** ## Next Steps Now that you're familiar with the Scraper dashboard, continue with **Extract Content from a Single URL** to create your first extraction request. # Extract Content from a Single URL (/docs/scraper-api/dashboard-guides/extraction/01_extract_single_url) import ScraperAPIRequestParameters from "../../snippets/requests-parameters.mdx"; The Scraper dashboard lets you extract structured content from a single web page without writing code. Enter the URL, configure the extraction settings if needed, and start the request. Once the extraction is complete, you can preview or download the results directly from the dashboard. ## Step 1. Enter the URL Enter the URL you want to extract into the URL field. After entering a valid URL, you can either start the extraction immediately or configure additional settings before submitting the request. ## Step 2. Configure the extraction (Optional) Click **Settings** to configure the extraction before starting the request. Extraction Settings The Settings panel allows you to configure: If required, select a specific proxy country or leave it set to **Any** to automatically choose the best available location. The **Advanced settings** section provides additional request options for more advanced extraction scenarios. ## Step 3. Start the extraction After reviewing your configuration, click **Start Extraction**. The dashboard submits the request and begins processing the page. ## Step 4. View the results When the extraction completes, select the job from **Recent Jobs** to open the result. The result page includes three output tabs: * **Preview** – Displays the extracted content in an interactive viewer. * **Markdown** – Displays the extracted Markdown output and allows you to copy or download it. * **HTML** – Displays the extracted HTML output and allows you to copy or download it. Switch between the available output formats using the tabs at the top of the result page. ## Step 5. Download the output You can download or copy the extracted content directly from the result page. ### Preview The **Preview** tab lets you inspect the extracted content before downloading it. Extraction Result ### Markdown The **Markdown** tab allows you to: * Copy the Markdown output. * Download the Markdown file. Markdown Output ### HTML The **HTML** tab allows you to: * Copy the HTML output. * Download the HTML file. HTML Output ## Next Step Need to extract multiple pages in a single request? Continue with **Extract Content from Multiple URLs**. # Extract Content from Multiple URLs (/docs/scraper-api/dashboard-guides/extraction/02_extract_multiple_urls) import ScraperAPIRequestParameters from "../../snippets/requests-parameters.mdx"; The Scraper dashboard lets you extract content from multiple web pages in a single batch request. You can add URLs manually or import them from a CSV file, configure the extraction settings, and download the results once processing is complete. ## Step 1. Open Batch Extraction From the Scraper dashboard, click **Set Multiple URLs**. Open Batch Extraction The **Batch Extraction** dialog opens, allowing you to submit multiple URLs in a single request. *** ## Step 2. Add URLs You can add URLs in two ways. ### Add URLs manually Enter a URL into the input field and click **Add in table**. Repeat this step until all required URLs have been added. ### Import a CSV file If you already have a list of URLs, click **Import CSV** to upload them in bulk. You can also download the sample CSV template by selecting **Download CSV template**. Add URLs The **Batch size** section displays the maximum number of URLs that can be included in a single batch based on your current plan. *** ## Step 3. Configure the extraction (Optional) Select **Settings** to configure the extraction request before starting the batch. Available options include: *** ## Step 4. Start the batch extraction After adding your URLs and reviewing the configuration, click **Start Batch Extraction**. Start Batch Extraction The dashboard creates a batch job and begins processing each URL. *** ## Step 5. Monitor the batch job Return to the **Recent Jobs** section and switch to the **Batches** tab. While the batch is running, you can monitor its progress. The dashboard displays: * Current job status * Total URLs in the batch * Completed URLs * Failed URLs * Creation time * Selected output format Batch Processing The completed and failed counters update as each URL finishes processing. *** ## Step 6. Download the results When the batch status changes to **Completed**, click the **Download** icon. Completed Batch The dashboard downloads the extraction results for every successfully processed URL. Depending on the selected output format, each URL is saved as an individual file. Downloaded Files *** ## Next Step Learn how to search, monitor, and manage your extraction history in the **Manage Extraction Jobs** guide. # Manage Extraction Jobs (/docs/scraper-api/dashboard-guides/extraction/03_manage_extraction_jobs) The **Recent Jobs** section lets you monitor all extraction requests submitted through the Scraper dashboard. You can switch between single URL and batch jobs, search previous requests, filter jobs by status, download completed results, and review extraction statistics. ## Switch between job types Open the **Recent Jobs** section and choose the type of jobs you want to view. * **Single URL** displays individual extraction requests. * **Batches** displays batch extraction requests. Recent Jobs ## Search and filter jobs For **Single URL** jobs, you can quickly find previous requests using the available filters. The dashboard allows you to: * Search by URL. * Filter jobs by status. * Filter jobs by output format. Search and Filter ## Filter by status Select the **Status** filter to display only jobs with a specific status. Available filters include: * Queued * Processing * Completed * Failed * Cancelled Status Filter ## Monitor batch jobs When viewing **Batch** jobs, the dashboard displays additional information about each extraction request. For every batch, you can see: * Current status * Number of URLs in the batch * Completed URLs * Failed URLs * Creation date * Execution time * Output format While a batch is running, the **Completed** and **Failed** counters update as each URL finishes processing. Once processing is complete, select the **Download** icon to download the extraction results. ## View extraction statistics Open the **Statistics** tab to review your extraction activity. Statistics The Statistics page includes: * Date range selection * Total extractions * Success rate * Average extraction duration * Requests used * Request activity over time * Request usage charts Use these statistics to monitor extraction performance and understand how your requests are being used over the selected period. # Map Overview (/docs/scraper-api/dashboard-guides/map/00_map_overview) import Link from "next/link"; The Map dashboard helps you discover the URLs available on a website before extracting content. It provides a simple interface for creating mapping requests, reviewing previous jobs, and monitoring your account usage. Before creating your first mapping request, we recommend reviewing the following guides to understand how the Map API works and the best practices for efficient website discovery.
  • Understanding Map
  • Map Workflows
  • Map Best Practices
These guides explain how the Map API discovers URLs, how mapping requests are processed, and the recommended workflows for achieving the best results. Map Dashboard Overview ## 1. Account Overview The top section provides information about your account and subscription. Here you can: * View your available requests. * Check your free request allowance. * See when your requests renew. * Review the number of requests used during the last 24 hours. * Upgrade your plan if you need additional requests. This section helps you monitor your available usage before submitting new mapping requests. ## 2. Mapping Workspace The center section is where you create new mapping requests. From this workspace, you can: * Enter the website URL you want to map. * Open the **Settings** panel to configure the request. * Access **API & Integrations**. * Start a new mapping request. This guide introduces the dashboard only. The complete mapping workflow is covered in the **Create Your First Map** guide. ## 3. Recent Maps The bottom section displays your previous mapping requests. From here, you can: * View recent mapping jobs. * Search previous requests by URL. * Filter jobs by status. * Review execution times. * View the number of discovered links. * Open completed mapping results. Managing previous mapping jobs is covered in the **Manage Map Jobs** guide. ## Next Step Continue with **Create Your First Map** to submit your first mapping request using the dashboard. # Discover URLs (/docs/scraper-api/dashboard-guides/map/01_map_discover_urls) import Link from "next/link"; The Map dashboard helps you discover URLs from a website before extracting content. Enter the website URL, optionally configure the mapping settings, and start the mapping request. Once the mapping is complete, you can review, copy, or download the discovered URLs. ## Step 1. Enter the website URL Enter the website you want to map into the URL field. Enter a Website URL Optionally, you can configure the mapping request before starting it. Available options include: * Filter URLs * Include subdomains * Ignore query parameters For a detailed explanation of these options and the recommended mapping workflow, refer to the following guides:
  • Understanding Map
  • Map Workflows
  • Map Best Practices
After reviewing the settings, click **Start Mapping**. ## Step 2. View the mapping job Once the mapping request has been submitted, it appears in the **Recent Maps** section. Recent Maps Each completed mapping request displays: * Status * Website URL * Creation date * Execution time * Number of discovered links Click the **View** icon to open the mapping results. ## Step 3. Review the discovered URLs The Map Preview displays all discovered URLs for the completed mapping request. Map Preview From the preview, you can: * Review all discovered URLs. * Copy the complete list to your clipboard. * Download the discovered URLs as a file. ## Next Step Continue with **Manage Map Jobs** to learn how to search, filter, and review previous mapping requests. # Manage Map Jobs (/docs/scraper-api/dashboard-guides/map/02_map_manage_jobs) The **Recent Maps** section lets you review all mapping requests submitted through the Map dashboard. You can search previous requests, filter them by status, open completed results, and monitor your mapping activity using the built-in statistics. ## Search and filter map jobs Open the **Recent Maps** tab to view your mapping history. Recent Maps The dashboard allows you to: * Search previous mapping requests by URL. * Filter jobs by status. * Open completed mapping results. The available status filters include: * Queued * Processing * Completed * Failed * Cancelled Click the **View** icon to open the discovered URLs for a completed mapping request. ## View mapping statistics Select the **Statistics** tab to review your mapping activity. Map Statistics The Statistics page provides an overview of your mapping requests for the selected time period. You can: * Select a custom date range. * Choose a predefined time range, such as the last 24 hours, 7 days, 30 days, or 90 days. * View the total number of mapping requests. * Monitor your success rate. * Review the average mapping duration. * See the number of requests used during the selected period. The page also includes charts showing request activity and usage trends over time, helping you monitor mapping performance and request consumption. ## Next Step You're now familiar with the complete Map dashboard workflow, including creating mapping requests, reviewing discovered URLs, and managing previous jobs. # Search Overview (/docs/scraper-api/dashboard-guides/search/00_search_overview) The Search dashboard provides a workspace for submitting searches, reviewing recent search jobs, and monitoring search usage. Search API dashboard ## Account Overview At the top of the Search dashboard, you can view your current request usage. The account overview shows: * **Available Requests** — Number of requests currently available. * **Free requests** — Number of free requests included with your account. * **Renewal date** — When your free requests renew. * **Used last 24 hours** — Number of requests used during the last 24 hours. You can also use the **See pricing** option to view available plans. ## Search Workspace The main Search workspace is where you enter your search query. Enter your query in the search field and use: * **Settings** to configure the search. * **API & Integrations** to access API-related options. * **Search** to submit the search. For information about making a search from the dashboard, see [Search and View Results](/docs/scraper-api/dashboard-guides/search/01_search_and_view_results). ## Recent Searches The **Recent Searches** section displays your previous search jobs. Each search displays: * **Status** — Current status of the search job. * **Query** — Search query that was submitted. * **Created** — Date and time when the search was created. * **Execution Time** — Time taken to complete the search. You can also: * Search previous jobs by query. * Filter jobs by status. * Change the number of rows displayed per page. * Move between result pages. * Open an individual search to view its results. ## Search Status Search jobs can appear with different statuses depending on their current state. Use the **Status** filter to narrow the list of recent searches by status. ## Statistics The Search dashboard also includes a **Statistics** tab for viewing search usage. Open **Statistics** to view your search activity and usage information. For more information about managing search jobs and monitoring usage, see [Manage Search Jobs](/docs/scraper-api/dashboard-guides/search/02_manage_search_jobs). ## What's Next? You now know the main sections of the Search dashboard. Continue to [Search and View Results](/docs/scraper-api/dashboard-guides/search/01_search_and_view_results) to learn how to submit a search and review the returned results. # Search and View Results (/docs/scraper-api/dashboard-guides/search/01_search_and_view_results) Use the Search dashboard to enter a search query, configure search options, and review the results directly in the dashboard. ## Configure Your Search Enter your search query in the Search field. You can open **Settings** to configure options such as: * **Locale** * **Time range** * **Safe search** Search settings For detailed information about the available search parameters and how they affect your request, see [Search Parameters](/docs/scraper-api/guides/search/03_search_parameters). ### Locale The **Locale** setting controls the language or region used for the search results. The default option is **Auto**. ### Time Range The **Time range** setting allows you to apply a time range to the search. The default option is **Off**. ### Safe Search The **Safe search** setting controls the level of filtering applied to the search results. Available options are: * **Off** * **Moderate** * **Strict** For the complete list of supported search parameters and their values, see [Search Parameters](/docs/scraper-api/guides/search/03_search_parameters). ## Submit the Search After entering your query and configuring the available settings, select **Search** to start the search. The search is then submitted and the results are displayed in the dashboard. ## Review Search Results Once the search is completed, the results section displays the returned search results. Search results The results view includes: * **Result count** — Number of results returned. * **Query** — Search query used for the request. * **Locale** — Locale used for the search. * **Safe search** — Safe search setting used for the search. * **Status** — Current status of the search. Each result can include: * Result position * Title * URL * Snippet * Displayed URL ## View Results as JSON The results section provides two views: * **Results** — Displays the search results in a readable format. * **JSON** — Displays the returned results as JSON. Select **JSON** when you need to inspect the returned search data in its JSON format. ## What's Next? You now know how to configure a search, submit a query, and review the returned results. Continue to [Manage Search Jobs](/docs/scraper-api/dashboard-guides/search/02_manage_search_jobs) to learn how to search, filter, and manage your previous searches. # Manage Search Jobs (/docs/scraper-api/dashboard-guides/search/02_manage_search_jobs) The **Search dashboard** lets you manage previous searches and monitor your search activity from one place. ## Recent Searches Open the **Recent Searches** tab to view your previous search jobs. Recent searches The Recent Searches section provides: * **Search by query** — Find a previous search by its query. * **Status** — Filter searches by their current status. * **Query** — View the query used for each search. * **Created** — See when the search was created. * **Execution Time** — See how long the search took to complete. You can also use the pagination controls to move between pages of search jobs and change the number of rows displayed per page. ## Filter Search Jobs Use **Search by query** to find a specific search. For example, you can enter: ```text web scraping best practice ``` to find searches containing that query. You can also use the **Status** filter to narrow the list of searches by their current status. ## View a Previous Search Each search job has an action button on the right side of the table. Select the action button to open the search and review its results. View search results The search results open in a separate view where you can switch between: * **Results** — View the returned search results. * **JSON** — View the results in JSON format. The search view also shows the search query and the status of the search. ## Statistics The **Statistics** tab provides an overview of your Search API usage. Search statistics ### Date Range You can select a custom date range or use one of the available time periods: * **Last 24 hours** * **7 days** * **30 days** * **90 days** ### Usage Summary The statistics view provides the following information for the selected period: * **Total** — Total number of searches. * **Success Rate** — Percentage of successful searches. * **Average Duration** — Average search duration. * **Requests Used** — Number of requests used during the selected period. ### Requests The **Requests** chart shows search requests over the selected period. The chart distinguishes between: * Successful requests * Failed requests ### Requests Usage The **Requests usage** chart shows request usage for the selected period. ## What's Next? You now know how to review previous searches, filter search jobs, open their results, and monitor Search API usage. For the API-level details of search jobs, see [Search Jobs](/docs/scraper-api/guides/search/06_search_jobs). # Understanding Batch (/docs/scraper-api/guides/batch/00_understanding_batch) Batch extraction lets you process multiple URLs in a single asynchronous request. Instead of sending a separate request for every URL, you submit all URLs together as one batch job. Geonode processes each URL in the background and lets you retrieve the results after the job has completed. The Batch endpoint immediately returns a job ID instead of the extracted content. Use the job ID to monitor progress and retrieve the results when processing finishes. ## When to Use Batch Batch extraction is useful when you need to process many webpages at once. Common use cases include: * Extracting product pages from an e-commerce website * Processing blog articles or news posts * Scraping documentation pages * Monitoring multiple websites * Running scheduled extraction jobs If you only need to extract content from a single webpage, use the Extraction endpoint instead. *** ## Choosing the Right Endpoint Geonode provides different endpoints depending on the task you want to perform. | If you want to... | Use | | ------------------------------------------ | ---------- | | Extract content from a single URL | Extraction | | Extract content from multiple URLs | Batch | | Crawl an entire website by following links | Crawl | *** ## Why Use Batch? Without batch extraction, every URL requires its own API request. E1["Extract"] U2["URL 2"] --> E2["Extract"] U3["URL 3"] --> E3["Extract"] U4["URL 4"] --> E4["Extract"] end subgraph B["With Batch"] B1["URL 1"] B2["URL 2"] B3["URL 3"] B4["URL 4"] B1 --> API["Batch API"] B2 --> API B3 --> API B4 --> API API --> JOB["Batch Job"] end `} /> Batch extraction groups multiple URLs into a single request, making it easier to process large collections of webpages. *** ## How Batch Works Batch extraction follows a simple asynchronous workflow. B["Submit Batch Request"] --> C["Batch Job Created"] --> D["Job ID Returned"] --> E["Processing"] --> F["Retrieve Results"] `} /> Unlike the Extraction endpoint, Batch does not return the extracted content immediately. Instead, it creates a job and returns a job ID that you can use to monitor progress and retrieve the results once processing has finished. *** ## Batch Workflow A typical batch request follows these steps: 1. Submit one or more URLs to the Batch endpoint. 2. Geonode creates a new batch job. 3. The API immediately returns a unique job ID. 4. Each URL is processed independently. 5. Retrieve the completed results using the job ID. *** ## Batch vs Extraction Although both endpoints extract webpage content, they are designed for different workloads. | Feature | Extraction | Batch | | ---------------- | --------------------------- | -------------------------------- | | Number of URLs | One | Multiple | | Processing | Synchronous or asynchronous | Asynchronous | | Best for | Individual webpages | Large collections of webpages | | Initial response | Extracted content or Job ID | Job ID | | Final result | Single extraction | Collection of extraction results | *** ## Key Concepts Before using the Batch endpoint, keep the following in mind: * Each URL is processed independently. * Some URLs may succeed while others fail within the same batch. * Batch requests return a job ID instead of the extracted content. * Results can be retrieved after processing has completed. *** ## Next Steps Now that you understand how batch extraction works, continue to **Your First Batch** to submit your first batch extraction request. # Batch Workflows (/docs/scraper-api/guides/batch/01_batch_workflows) A batch extraction is an asynchronous workflow. Instead of waiting for every URL to finish processing, the API immediately creates a job and returns a unique job ID. That job ID becomes the center of the workflow. You use it to monitor progress, retrieve results, or cancel the job if necessary. ## Complete Batch Workflow The following diagram shows how the Batch API endpoints work together. Every batch job follows this lifecycle, from creation to completion. *** ## Typical Workflow Most applications follow these steps when working with batch extraction. ### Step 1 — Create a Batch Start by submitting one or more URLs. ```http POST /v1/batch ``` The API immediately returns: * `job_id` * `status` * `status_url` * `accepted_urls` At this point, the extraction has been queued and continues in the background. *** ### Step 2 — Store the Job ID Save the returned `job_id`. You'll need this value for every subsequent operation, including: * Monitoring progress * Viewing results * Cancelling the batch *** ### Step 3 — Monitor Progress Use the Job Status endpoint until the batch finishes. ```http GET /v1/batch/{job_id} ``` A batch can move through the following states: ```text queued processing completed failed cancelled ``` Most applications poll this endpoint every few seconds until processing finishes. *** ### Step 4 — Process the Results Once the job reaches the `completed` state, the response contains: * Successfully processed URLs * Failed URLs * Output for each completed extraction Your application can now download, store, or process the extracted content. *** ## Finding Previous Jobs Sometimes an application loses the job ID due to a restart or network interruption. Instead of creating another batch, retrieve the existing one. ```http GET /v1/batch/jobs ``` This prevents duplicate processing and allows your application to continue from where it left off. *** ## Cancelling a Batch You can cancel a running batch if it is no longer needed. Common reasons include: * Incorrect URLs * Wrong proxy configuration * Incorrect request settings * User cancelled the operation ```http DELETE /v1/batch/{job_id} ``` After cancellation: * No new URLs are scheduled for processing. * URLs that are already running may finish. * Remaining URLs are skipped. *** ## End-to-End Example The complete workflow usually looks like this. | Step | Endpoint | Purpose | | ---- | --------------------------- | ------------------------------------------------- | | 1 | `POST /v1/batch` | Create a new batch job | | 2 | Save `job_id` | Store the identifier for future requests | | 3 | `GET /v1/batch/{job_id}` | Monitor progress until completion | | 4 | `GET /v1/batch/jobs` | Find previous jobs when the job ID is unavailable | | 5 | `DELETE /v1/batch/{job_id}` | Cancel the batch if it is no longer needed | *** ## Common Workflow Patterns ### Standard Batch Processing This is the most common workflow for asynchronous processing. *** ### Recovering After an Application Restart This allows applications to recover without creating duplicate jobs. *** ### Cancelling an Active Batch *** ## Best Practices * Store the returned `job_id` immediately after creating a batch. * Poll the Job Status endpoint instead of submitting duplicate batch requests. * Use the List Jobs endpoint to recover lost job IDs. * Cancel jobs that are no longer required instead of letting them continue processing. * Wait until the batch reaches the `completed` state before processing the final results. * Handle every possible job status (`queued`, `processing`, `completed`, `failed`, and `cancelled`) in your application. *** ## Next Steps You now understand the complete lifecycle of a batch job. Continue to the **Batch API Reference** for detailed request parameters, response fields, and endpoint-specific examples. # Your First Batch (/docs/scraper-api/guides/batch/02_your_first_batch) Batch jobs allow you to submit multiple URLs in a single API request. Instead of sending one extraction request per page, you create a single batch job. Geonode queues the job, processes each URL asynchronously, and lets you retrieve the results later. In this guide, you'll: * Create your first batch job * Understand the response * Learn how invalid URLs are handled * Retrieve the batch status *** ## Create Your First Batch Send a `POST` request to the batch endpoint. ```http POST /v1/batch ``` ### Request ```bash title="request.sh" curl -X POST "https://scraper.geonode.io/v1/batch" \ -H "X-Api-Key: YOUR_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "urls": [ "https://geonode.com", "https://docs.geonode.com", "https://example.com" ] }' ``` ### Response ```json title="response.json" { "job_id": "d16e56a0-dbe9-4586-a40d-028cf3c439a9", "status": "queued", "status_url": "/v1/batch/d16e56a0-dbe9-4586-a40d-028cf3c439a9", "accepted_urls": 3, "invalid_urls": [] } ``` The request creates a new batch job immediately. Geonode validates the request, queues the job, and returns a unique job ID that you can use to monitor progress. *** ## Understanding the Response The initial response confirms that the batch job has been accepted. | Field | Description | | --------------- | -------------------------------------------------------- | | `job_id` | Unique identifier for the batch job. | | `status` | Current processing status. A new job starts as `queued`. | | `status_url` | Endpoint used to retrieve the latest status and results. | | `accepted_urls` | Number of valid URLs accepted for processing. | | `invalid_urls` | URLs that were rejected during validation. | The actual extraction results are **not** returned immediately because batch jobs run asynchronously. *** ## What Happens Next? Once the request is accepted, Geonode processes each URL in the background. You can use the returned `job_id` or `status_url` to check the progress of the job at any time. *** ## Handling Invalid URLs By default, every URL in the batch request must be valid. If one or more URLs are invalid and `ignore_invalid_urls` is set to `false`, the entire request is rejected. ### Request ```json { "urls": [ "https://geonode.com", "https://docs.geonode.com", "not-a-valid-url" ], "ignore_invalid_urls": false } ``` ### Response ```json { "code": "VALIDATION_ERROR", "message": "Batch contains invalid URLs and ignore_invalid_urls is false", "correlation_id": "1fd0958b-6583-4c32-9e9b-5ad7960ac5ad", "retryable": false } ``` No batch job is created until the request contains only valid URLs. *** ## Skip Invalid URLs If you want Geonode to continue processing valid URLs, enable `ignore_invalid_urls`. ### Request ```json { "urls": [ "https://geonode.com", "https://docs.geonode.com", "https://example.com", "not-a-valid-url" ], "ignore_invalid_urls": true } ``` ### Response ```json { "job_id": "d16e56a0-dbe9-4586-a40d-028cf3c439a9", "status": "queued", "status_url": "/v1/batch/d16e56a0-dbe9-4586-a40d-028cf3c439a9", "accepted_urls": 3, "invalid_urls": [ "not-a-valid-url" ] } ``` Only the valid URLs are queued for processing. Invalid URLs are skipped and returned in the `invalid_urls` array. Use `ignore_invalid_urls` when importing URLs from user input, CSV files, or external systems where a few invalid entries shouldn't stop the entire batch. *** ## Check the Batch Status Batch jobs are processed asynchronously. After creating a job, use the returned `job_id` to check its status. ```http GET /v1/batch/{job_id} ``` You'll learn how to monitor a running job and retrieve its results in the next guide. *** ## Next Steps Now that you've created your first batch job, continue to **Working with Batch Inputs** to learn how to customize batch requests with output formats, JavaScript rendering, proxies, custom headers, and other request options. # Configuring Batch Requests (/docs/scraper-api/guides/batch/03_configuring_batch_requests) Every batch request starts with a list of URLs. You can further customize how those URLs are processed by configuring output formats, JavaScript rendering, proxy settings, custom headers, and other optional request parameters. All examples in this guide create a new batch job. Regardless of which request options you use, the Batch endpoint immediately returns a job ID and begins processing the accepted URLs in the background. ## Request Body Overview The following request fields are available when creating a batch job. | Field | Required | Description | | --------------------- | -------- | -------------------------------------------------- | | `urls` | ✅ | List of URLs to process. | | `ignore_invalid_urls` | No | Continue processing even if some URLs are invalid. | | `formats` | No | Choose the output format for extracted content. | | `render_js` | No | Render JavaScript before extraction. | | `processing_mode` | No | Configure how URLs are processed. | | `proxy` | No | Route requests through a proxy. | | `headers` | No | Include custom HTTP headers. | | `wait_config` | No | Wait for dynamic content before extraction. | All of these options are optional except `urls`. *** ## URLs Every batch request must include one or more valid URLs. ```json { "urls": [ "https://geonode.com", "https://docs.geonode.com" ] } ``` Each URL is processed independently as part of the same batch job. *** ## Ignoring Invalid URLs By default, every URL must be valid. If you want Geonode to continue processing valid URLs while skipping invalid ones, enable `ignore_invalid_urls`. ```json { "urls": [ "https://geonode.com", "https://docs.geonode.com", "not-a-valid-url" ], "ignore_invalid_urls": true } ``` For a complete example of how invalid URLs are handled, see **Your First Batch**. *** ## Output Formats Choose one or more output formats for the extracted content. ```json { "urls": [ "https://geonode.com" ], "formats": [ "markdown", "html" ] } ``` Learn more in **Output Formats**. *** ## JavaScript Rendering Enable JavaScript rendering for websites that load content dynamically. ```json { "urls": [ "https://geonode.com" ], "render_js": true } ``` Learn more in **JavaScript Rendering**. *** ## Processing Mode Control how the batch request is processed. ```json { "urls": [ "https://geonode.com" ], "processing_mode": "parallel" } ``` Learn more in **Processing Modes**. *** ## Proxy Configuration Use a proxy when accessing geo-restricted or protected websites. ```json { "urls": [ "https://geonode.com" ], "proxy": { "...": "..." } } ``` Learn more in **Proxy & Geo-Targeting**. *** ## Custom Headers Include custom HTTP headers with every request. ```json { "urls": [ "https://geonode.com" ], "headers": { "Authorization": "Bearer YOUR_TOKEN" } } ``` Learn more in **Using Custom Headers**. *** ## Waiting for Dynamic Content Delay extraction until specific content becomes available. ```json { "urls": [ "https://geonode.com" ], "wait_config": { "...": "..." } } ``` Learn more in **Waiting for Dynamic Content**. *** ## Complete Example The following example combines several request options. ```json { "urls": [ "https://geonode.com", "https://docs.geonode.com" ], "formats": [ "markdown" ], "render_js": true, "ignore_invalid_urls": true } ``` The response immediately returns a queued batch job. ```json { "job_id": "e9379def-8986-4624-86f1-f49c35afe711", "status": "queued", "status_url": "/v1/batch/e9379def-8986-4624-86f1-f49c35afe711", "accepted_urls": 2, "invalid_urls": [] } ``` *** ## Next Steps Continue to **Batch Results** to learn how to monitor batch jobs, retrieve completed results, and understand the different job states. # Monitoring Batch Jobs (/docs/scraper-api/guides/batch/04_monitoring_batch_jobs) Batch jobs are processed asynchronously. After creating a batch job, use the returned `job_id` to monitor its progress and retrieve the extraction results. *** ## Retrieve a Batch Job Use the following endpoint to retrieve the latest status and results for a batch job. ```http GET /v1/batch/{job_id} ``` ### Request ```bash title="request.sh" curl -X GET "https://scraper.geonode.io/v1/batch/YOUR_JOB_ID" \ -H "X-Api-Key: YOUR_API_KEY" ``` Replace `YOUR_JOB_ID` with the job ID returned when the batch was created. *** ## Example Response ```json title="response.json" { "job_id": "d54da3de-7afc-4127-a5ee-5f7e0378a3e6", "status": "completed", "created_at": "2026-07-05T16:14:22.721Z", "completed_at": "2026-07-05T16:14:25.318Z", "total_urls": 3, "completed_urls": 2, "failed_urls": 1, "pending_urls": 0, "cancelled_urls": 0, "token_summary": { "tokens_charged_total": 2, "tokens_reserved": 0 }, "results": [ { "input_index": 0, "url": "https://geonode.com", "status": "completed" }, { "input_index": 1, "url": "https://docs.geonode.com", "status": "completed" }, { "input_index": 2, "url": "https://this-domain-does-not-exist-123456789.com", "status": "failed" } ] } ``` *** ## Job Information The top-level response describes the overall batch job. | Field | Description | | -------------- | ------------------------------------ | | `job_id` | Unique identifier for the batch job. | | `status` | Current status of the batch job. | | `created_at` | Time the batch job was created. | | `completed_at` | Time processing finished. | *** ## Processing Statistics These fields summarize the progress of the batch. | Field | Description | | ---------------- | -------------------------------------------- | | `total_urls` | Total number of submitted URLs. | | `completed_urls` | URLs processed successfully. | | `failed_urls` | URLs that failed during processing. | | `pending_urls` | URLs that are still waiting to be processed. | | `cancelled_urls` | URLs cancelled before processing completed. | *** ## Token Usage The response also includes a summary of token usage. | Field | Description | | ---------------------- | ----------------------------------- | | `tokens_charged_total` | Total tokens consumed by the batch. | | `tokens_reserved` | Tokens reserved for processing. | *** ## Individual Results Each submitted URL appears in the `results` array. Each result contains information about a single URL. | Field | Description | | --------------- | -------------------------------------------- | | `input_index` | Position of the URL in the original request. | | `url` | The processed URL. | | `status` | Processing status for that URL. | | `error_code` | Error code when processing fails. | | `error_message` | Human-readable error description. | | `data` | Extracted content for successful requests. | | `metadata` | Processing metadata for the request. | *** ## Job Status Each URL has its own processing status. | Status | Description | | ------------ | -------------------------- | | `queued` | Waiting to be processed. | | `processing` | Currently being processed. | | `completed` | Successfully extracted. | | `failed` | Processing failed. | | `cancelled` | Processing was cancelled. | *** ## Partial Failures A batch job can complete successfully even if some URLs fail. For example, a batch containing three URLs may produce: * 2 completed URLs * 1 failed URL * Overall batch status: `completed` Failed URLs include additional error information. ```json { "url": "https://this-domain-does-not-exist-123456789.com", "status": "failed", "error_code": "INTERNAL_ERROR", "error_message": "An unexpected error occurred on the server" } ``` This allows you to retry only the failed URLs instead of rerunning the entire batch. Failed URLs do not prevent other URLs in the same batch from completing successfully. *** ## Metadata Each completed result includes processing metadata. | Field | Description | | ---------------- | ------------------------------------------------ | | `http_status` | HTTP status code returned by the target website. | | `duration_ms` | Time taken to process the URL. | | `tokens_charged` | Tokens consumed for that extraction. | This information can help monitor performance and troubleshoot extraction issues. *** ## Best Practices When working with batch jobs: * Save the returned `job_id` so you can retrieve the results later. * Wait until the job status is `completed` before using the extracted data. * Review the processing statistics to quickly identify failed URLs. * Retry only failed URLs instead of rerunning the entire batch. * Monitor token usage when processing large batches. *** ## Next Steps Now that you know how to monitor batch jobs and retrieve their results, continue to **Batch Workflows** to learn common patterns for processing large collections of URLs. # Listing Batch Jobs (/docs/scraper-api/guides/batch/05_listing_batch_jobs) Listing batch jobs allows you to retrieve previously created batch requests. This is useful when you want to review completed jobs, monitor running batches, or locate a job ID for checking its status. ## Batch Jobs Endpoint Use the following endpoint to retrieve previously created batch jobs. ```http GET /v1/batch/jobs ``` The response returns a paginated list of batch jobs. ## List Batch Jobs Use the following request to retrieve your recent batch jobs. ### Request ```bash title="request.sh" curl -X GET "https://scraper.geonode.io/v1/batch/jobs" \ -H "X-Api-Key: YOUR_API_KEY" ``` ### Response ```json title="response.json" { "jobs": [ { "job_id": "9e60e068-b6ac-408d-aa6e-f120f41faf0d", "status": "completed", "accepted_urls": 3, "completed_urls": 3, "failed_urls": 0, "config": { "render_js": false, "formats": [ "markdown" ], "proxy": { "country": null, "type": "residential" }, "headers": null, "wait_config": null }, "created_at": "2026-07-05T16:29:56.344150Z", "completed_at": "2026-07-05T16:30:13.963398Z" } ], "page": 1, "page_size": 10, "page_count": 1 } ``` ## Understanding the Response Each batch job contains summary information about the processing request. | Field | Description | | ---------------- | --------------------------------------------------------- | | `job_id` | Unique identifier for the batch job. | | `status` | Current status of the batch job. | | `accepted_urls` | Number of valid URLs accepted when the batch was created. | | `completed_urls` | Number of URLs processed successfully. | | `failed_urls` | Number of URLs that failed during processing. | | `config` | Configuration used when the batch job was submitted. | | `created_at` | Time the batch job was created. | | `completed_at` | Time the batch job finished processing. | The response also includes pagination information. | Field | Description | | ------------ | --------------------------------- | | `page` | Current page number. | | `page_size` | Number of jobs returned per page. | | `page_count` | Total number of available pages. | ## Filtering Batch Jobs You can filter the returned jobs using query parameters. | Query Parameter | Description | | --------------- | ---------------------------------------------------- | | `status` | Return only jobs with a specific status. | | `start_date` | Return jobs created on or after the specified date. | | `end_date` | Return jobs created on or before the specified date. | | `page` | Page number to retrieve. | | `page_size` | Number of jobs to return per page. | ### Filter by Status Retrieve only completed batch jobs. ```bash title="request.sh" curl -X GET "https://scraper.geonode.io/v1/batch/jobs?status=completed" \ -H "X-Api-Key: YOUR_API_KEY" ``` Supported values: ```text queued processing completed failed cancelled ``` ### Filter by Date Range Retrieve jobs created within a specific time period. ```bash title="request.sh" curl -X GET "https://scraper.geonode.io/v1/batch/jobs?start_date=2026-07-01&end_date=2026-07-31" \ -H "X-Api-Key: YOUR_API_KEY" ``` ### Pagination Retrieve a specific page of results. ```bash title="request.sh" curl -X GET "https://scraper.geonode.io/v1/batch/jobs?page=2&page_size=10" \ -H "X-Api-Key: YOUR_API_KEY" ``` ## Understanding Job Progress The job summary makes it easy to see how a batch performed without retrieving the full job details. | Scenario | Accepted URLs | Completed URLs | Failed URLs | | ------------------------------- | ------------: | -------------: | ----------: | | All URLs processed successfully | 3 | 3 | 0 | | One URL failed | 3 | 2 | 1 | | Two URLs failed | 5 | 3 | 2 | ## When to Use This Endpoint Use this endpoint to: * View recent batch activity. * Find a previous batch job. * Monitor completed or failed batches. * Search jobs created within a specific date range. * Retrieve a job ID before checking detailed status. * Build dashboards or reporting tools. ## Next Steps Now that you know how to retrieve and filter batch jobs, continue to **Cancelling Batch Jobs** to learn how to stop a running batch job. # Batch Limits (/docs/scraper-api/guides/batch/07_batch_limits) Before creating a batch job, it's important to understand the limits and validation rules enforced by the API. Following these guidelines helps prevent validation errors and ensures your requests are accepted successfully. ## Batch Request Limits The following limits apply when creating a batch job. | Constraint | Details | | -------------- | ----------------------------------------------------------------------------------------------------------------------- | | Required field | The `urls` field is required. Omitting it returns a `422 Unprocessable Entity` error. | | Minimum URLs | A batch request must contain at least **1 URL**. An empty array (`[]`) is not allowed. | | Maximum URLs | A single batch request can contain up to **1000 URLs**. | | Duplicate URLs | Duplicate URLs are accepted and processed as separate requests. | | Processing | Batch requests are processed asynchronously. A successful request returns a `job_id` instead of the extraction results. | ## Required URLs Field The `urls` field must always be included in the request body. ### Invalid Request ```json {} ``` ### Response ```json { "detail": [ { "type": "missing", "loc": [ "body", "urls" ], "msg": "Field required" } ] } ``` ## Empty URL List Providing an empty array is also considered an invalid request. ### Invalid Request ```json { "urls": [] } ``` ### Response ```json { "detail": [ { "type": "too_short", "loc": [ "body", "urls" ], "msg": "List should have at least 1 item after validation, not 0" } ] } ``` ## Maximum Batch Size A single batch request can contain up to **1000 URLs**. If you need to process more than 1000 URLs, split them into multiple batch requests. ## Duplicate URLs Duplicate URLs are allowed. Each URL in the request is treated as a separate extraction request, even if the same URL appears multiple times. For example: ```json { "urls": [ "https://example.com", "https://example.com", "https://example.com" ] } ``` All three URLs are accepted and processed independently. ## Asynchronous Processing Creating a batch job does not immediately return the extracted content. Instead, the API returns a `job_id` that can be used to monitor the job until it completes. ```json { "job_id": "9e60e068-b6ac-408d-aa6e-f120f41faf0d", "status": "queued", "status_url": "/v1/batch/9e60e068-b6ac-408d-aa6e-f120f41faf0d", "accepted_urls": 3, "invalid_urls": [] } ``` Batch processing is asynchronous. After creating a batch job, use the returned `job_id` to check the job status and retrieve the results. The complete workflow is covered in the next guide. ## Next Steps Continue to **Batch Processing Workflows** to learn how to create a batch job, monitor its progress, and retrieve the final results. # Understanding Crawl (/docs/scraper-api/guides/crawl/00_overview) The Crawl API collects content from multiple pages across a website, starting from a single URL. It automatically discovers linked pages, extracts their content, and processes them as an asynchronous crawl job. *** ## What is Crawl? Crawl is a website traversal and content extraction service. Starting from a single **seed URL**, it visits the page, discovers additional links, and extracts content from every eligible page it crawls. Unlike extracting a single webpage, Crawl automatically continues exploring connected pages until the crawl is complete. **Good to know** Crawl combines **link discovery** and **content extraction** into a single automated workflow. *** ## At a Glance * 🌐 Start from a single seed URL. * 🔍 Discover linked pages automatically. * 📄 Extract content from each page. * ⚙️ Process pages as an asynchronous crawl job. * 📦 Retrieve crawl results when processing completes. *** ## How Crawl Works Every crawl starts from a single **seed URL**. As each page is processed, the crawler extracts content, discovers new links, and continues exploring eligible pages until the crawl is complete. The crawl follows a simple cycle: 1. Start from a seed URL. 2. Visit and process the page. 3. Extract the page content. 4. Discover additional links. 5. Continue crawling eligible pages. 6. Finish when the crawl completes. *** ## Crawl Workflow From your application's perspective, every crawl follows the same lifecycle. Once a crawl job is created, pages are processed in the background. You can then monitor the job and retrieve its results when processing finishes. *** ## Crawl vs. Map vs. Extraction Each Scraper API endpoint is designed for a different task. | Feature | Extraction | Map | Crawl | | -------------------------- | ---------- | --- | ----- | | Extract page content | ✅ | ❌ | ✅ | | Discover URLs | ❌ | ✅ | ✅ | | Follow links automatically | ❌ | ❌ | ✅ | | Process multiple pages | ❌ | ❌ | ✅ | ### When should you use each endpoint? | Use Case | Recommended Endpoint | | ------------------------------------- | -------------------- | | Extract content from one page | **Extraction** | | Discover URLs on a website | **Map** | | Collect content across multiple pages | **Crawl** | *** ## Common Use Cases Crawl is commonly used for: * Documentation websites * Knowledge bases * Blogs and news sites * Product catalogs * Company websites * AI and RAG data collection *** ## Best Practices To build efficient crawls: * Start with a small crawl before scaling up. * Crawl only the sections you need. * Review crawl results before downstream processing. * Configure crawl requests based on your use case. *** ## Next Steps Now that you understand how Crawl works, you're ready to create your first crawl job. Continue to **Your First Crawl** to learn how to submit your first crawl request and understand the initial API response. # Your First Crawl (/docs/scraper-api/guides/crawl/01_first-crawl) In this guide, you'll create your first crawl job using the Crawl API. The Crawl API accepts a starting URL and returns a crawl job that runs asynchronously. Once the job is created, you can monitor its progress and retrieve the results using the job ID. *** ## Before You Begin Before creating a crawl job, make sure you have: * A GeoNode API key. * The Scraper API base URL. If you haven't completed the initial setup, see **Before You Start**. *** ## Create a Crawl Job Create a crawl by sending a `POST` request to the following endpoint. ```http POST /v1/crawl ``` For your first crawl, you only need a starting URL and can optionally specify the output format and page limit. ### Example Request ```bash curl --request POST \ --url https://api.geonode.com/v1/crawl \ --header "Authorization: Bearer " \ --header "Content-Type: application/json" \ --data '{ "url": "https://docs.geonode.com/", "formats": [ "markdown" ], "limit": 5 }' ``` You can also send the request body as JSON. ```json { "url": "https://docs.geonode.com/", "formats": [ "markdown" ], "limit": 5 } ``` *** ## Example Response If the request is accepted, the API returns a `202 Accepted` response similar to the following. ```json { "job_id": "8c8f8ad4-a0a0-46f8-92d5-253025c6e19f", "url": "https://docs.geonode.com/", "status": "queued", "status_url": "/v1/crawl/8c8f8ad4-a0a0-46f8-92d5-253025c6e19f", "estimated_pages": 5 } ``` *** ## Understanding the Response | Field | Description | | ----------------- | ------------------------------------------- | | `job_id` | Unique identifier for the crawl job. | | `url` | The seed URL used to start the crawl. | | `status` | Current status of the crawl job. | | `status_url` | Endpoint for checking the crawl job status. | | `estimated_pages` | Estimated number of pages to be processed. | The Crawl API processes requests asynchronously. A successful request creates a crawl job rather than returning the crawl results immediately. *** ## What Happens Next? After your crawl job is created, you can use the returned job ID to monitor its progress. The next guide explains how to configure crawl requests with additional options, while **Managing Crawl Jobs** covers how to monitor job progress and retrieve results. *** ## Next Steps Continue to **Configuring Crawl Requests** to learn about all available request options, including crawl depth, page limits, output formats, JavaScript rendering, proxy settings, and browser wait configuration. # Configuring Crawl Requests (/docs/scraper-api/guides/crawl/02_configuring-crawl-requests) The Crawl API provides several request parameters that let you control how websites are crawled. This guide explains each parameter and shows how to configure common crawl scenarios. *** ## Request Body Every crawl request requires a starting URL. All other parameters are optional. ```json { "url": "https://docs.geonode.com/" } ``` *** ## Required Parameter ### `url` The starting URL for the crawl. | Property | Value | | -------- | ------ | | Type | String | | Required | ✅ Yes | Example: ```json { "url": "https://docs.geonode.com/" } ``` *** ## Crawl Depth The `depth` parameter controls how many levels deep the crawler follows links from the starting URL. | Property | Value | | -------- | ------- | | Type | Integer | | Required | No | | Default | `2` | Example: ```json { "url": "https://docs.geonode.com/", "depth": 2 } ``` The maximum crawl depth available depends on your subscription plan. If the requested depth exceeds your plan limit, the API returns a validation error. ```json { "code": "VALIDATION_ERROR", "message": "Crawl depth limit exceeded: requested 3, plan allows at most 2", "correlation_id": "5d2012da-9bf9-4224-978b-b4337ab7026f", "retryable": false } ``` *** ## Page Limit The `limit` parameter specifies the maximum number of pages to crawl. | Property | Value | | -------- | ------- | | Type | Integer | | Required | No | Example: ```json { "url": "https://docs.geonode.com/", "limit": 5 } ``` *** ## Domain Scope ### `same_domain_only` Limits crawling to pages within the same domain as the starting URL. | Property | Value | | -------- | ------- | | Type | Boolean | | Required | No | Example: ```json { "url": "https://docs.geonode.com/", "same_domain_only": true } ``` *** ### `include_subdomains` Controls whether subdomains are included during the crawl. | Property | Value | | -------- | ------- | | Type | Boolean | | Required | No | Example: ```json { "url": "https://docs.geonode.com/", "include_subdomains": false } ``` *** ## Additional Request Options The following parameters can be combined to further customize crawl requests. | Parameter | Description | | ------------- | --------------------------------------------------- | | `formats` | Specifies the output format for crawl results. | | `render_js` | Enables JavaScript rendering for dynamic websites. | | `proxy` | Configures the proxy used during crawling. | | `wait_config` | Controls browser wait behavior before page capture. | Example: ```json { "url": "https://docs.geonode.com/", "formats": [ "markdown" ], "render_js": true, "same_domain_only": true, "include_subdomains": false, "proxy": { "country": "us" }, "wait_config": { "wait_for": 3000 } } ``` > For detailed information about these parameters, see the dedicated guides for output formats, JavaScript rendering, proxy configuration, and browser wait configuration. *** ## Complete Example The following request combines multiple crawl configuration options. ```json { "url": "https://docs.geonode.com/", "depth": 2, "limit": 5, "formats": [ "markdown" ], "render_js": true, "same_domain_only": true, "include_subdomains": false, "proxy": { "country": "us" }, "wait_config": { "wait_for": 3000 } } ``` If the request is accepted, the API returns a response similar to the following: ```json { "job_id": "7344caad-297d-4eac-99ac-198a8c8e6233", "url": "https://docs.geonode.com/", "status": "queued", "status_url": "/v1/crawl/7344caad-297d-4eac-99ac-198a8c8e6233", "estimated_pages": 5 } ``` Creating a crawl returns a job immediately. Use the returned job\_id or status\_url to monitor progress and retrieve results. *** ## Next Steps Continue to **Understanding Crawl Results** to learn how crawl results are structured and how to interpret the returned data. # Understanding Crawl Results (/docs/scraper-api/guides/crawl/03_understanding-crawl-results) After creating a crawl job, use the job ID to retrieve its current status and results. The response includes information about the crawl job, the configuration used, crawl statistics, and the extracted content for each crawled page. *** ## Get Crawl Results Retrieve the current status and results of a crawl job. ```http GET /v1/crawl/{job_id} ``` Once the crawl completes, the response includes the extracted pages and their associated metadata. B["Crawl Configuration"] A --> C["Statistics"] A --> D["Token Summary"] A --> E["Results"] E --> F["Page 1"] E --> G["Page 2"] E --> H["Page N"] `} /> *** ## Job Information The top-level response provides general information about the crawl job. | Field | Description | | -------------- | ------------------------------------ | | `job_id` | Unique identifier for the crawl job. | | `url` | Starting URL used for the crawl. | | `status` | Current status of the crawl job. | | `created_at` | Time the crawl job was created. | | `completed_at` | Time the crawl job completed. | Example: ```json { "job_id": "9293f045-0523-409e-a4e9-874abc663a9e", "url": "https://docs.geonode.com/", "status": "completed", "created_at": "2026-07-23T16:55:51.603677Z", "completed_at": "2026-07-23T16:56:06.735093Z" } ``` *** ## Crawl Configuration The `crawl_config` object shows the configuration that was used when the crawl job was created. | Field | Description | | -------------------- | -------------------------------------------------------------- | | `render_js` | Indicates whether JavaScript rendering was enabled. | | `formats` | Output formats requested for extracted content. | | `same_domain_only` | Indicates whether crawling was limited to the starting domain. | | `include_subdomains` | Indicates whether subdomains were included. | | `proxy` | Proxy configuration used during the crawl. | | `wait_config` | Browser wait configuration, if specified. | Example: ```json { "crawl_config": { "render_js": false, "formats": [ "markdown" ], "same_domain_only": true, "include_subdomains": false, "proxy": { "country": null, "type": "residential" }, "wait_config": null } } ``` The returned configuration reflects the settings used for the crawl job. *** ## Crawl Statistics The response includes statistics that summarize the crawl. | Field | Description | | ----------------- | -------------------------------------------- | | `total_pages` | Total number of pages included in the crawl. | | `completed_pages` | Number of pages successfully processed. | | `failed_pages` | Number of pages that failed to process. | | `cancelled_pages` | Number of pages that were cancelled. | Example: ```json { "total_pages": 5, "completed_pages": 5, "failed_pages": 0, "cancelled_pages": 0 } ``` *** ## Token Summary The `token_summary` object provides information about token usage for the crawl job. | Field | Description | | ---------------------- | ----------------------------------- | | `tokens_charged_total` | Total tokens charged for the crawl. | | `tokens_reserved` | Tokens reserved for the crawl job. | Example: ```json { "token_summary": { "tokens_charged_total": 5, "tokens_reserved": 0 } } ``` *** ## Understanding the Results Array The `results` array contains one object for each crawled page. B["Page Result"] B --> C["Page Information"] B --> D["Extracted Data"] B --> E["Metadata"] B --> F["Links"] `} /> *** ## Page Information Each object in the `results` array describes a single crawled page. | Field | Description | | --------------- | ----------------------------------- | | `url` | URL of the crawled page. | | `parent_url` | URL where this page was discovered. | | `depth` | Crawl depth of the page. | | `status` | Crawl status for the page. | | `error_code` | Error code if the page failed. | | `error_message` | Error details if the page failed. | Example: ```json { "url": "https://docs.geonode.com/", "parent_url": null, "depth": 0, "status": "completed", "error_code": null, "error_message": null } ``` *** ## Extracted Data The `data` object contains the extracted content for the page. Depending on the requested output formats, it may include Markdown, HTML, or both. Example: ```json { "data": { "markdown": "# Welcome to Geonode\n\nFind exactly what you need...", "html": null } } ``` *** ## Page Metadata Each page includes metadata describing the crawl operation. | Field | Description | | ---------------- | ------------------------------------------- | | `http_status` | HTTP response status received for the page. | | `duration_ms` | Time taken to process the page. | | `tokens_charged` | Tokens charged for processing the page. | Example: ```json { "metadata": { "http_status": 200, "duration_ms": 957, "tokens_charged": 1 } } ``` *** ## Discovered Links The `links` array contains the URLs discovered on the crawled page. Example: ```json { "links": [ "https://docs.geonode.com/", "https://docs.geonode.com/docs/api-reference", "https://docs.geonode.com/docs/guides", "https://docs.geonode.com/docs/scraper-api", "https://docs.geonode.com/docs/changelog", "https://geonode.com/contact" ] } ``` *** ## Complete Response Example The following example shows a completed crawl job response. ```json { "job_id": "9293f045-0523-409e-a4e9-874abc663a9e", "url": "https://docs.geonode.com/", "status": "completed", "crawl_config": { "render_js": false, "formats": [ "markdown" ], "same_domain_only": true, "include_subdomains": false, "proxy": { "country": null, "type": "residential" }, "wait_config": null }, "token_summary": { "tokens_charged_total": 5, "tokens_reserved": 0 }, "total_pages": 5, "completed_pages": 5, "failed_pages": 0, "cancelled_pages": 0, "created_at": "2026-07-23T16:55:51.603677Z", "completed_at": "2026-07-23T16:56:06.735093Z", "results": [ { "url": "https://docs.geonode.com/", "parent_url": null, "depth": 0, "status": "completed" } ] } ``` *** ## Next Steps Now that you understand the structure of crawl results, continue to **Managing Crawl Jobs** to learn how to list crawl jobs, monitor their progress, and retrieve specific jobs. # Managing Crawl Jobs (/docs/scraper-api/guides/crawl/04_managing-crawl-jobs) Crawl jobs are processed asynchronously. After creating a crawl job, you can retrieve its status, monitor its progress, or list previous crawl jobs. This guide explains how to work with crawl jobs throughout their lifecycle. *** ## List Crawl Jobs Retrieve a list of your crawl jobs. ```http GET /v1/crawl/jobs ``` The response includes your crawl jobs along with their current status, crawl statistics, configuration, and pagination information. Example: ```json { "jobs": [ { "job_id": "453bd7d7-5355-4d6d-a38e-d9e7eb218c3f", "url": "string", "status": "queued", "total_pages": 0, "completed_pages": 0, "failed_pages": 0, "config": { "render_js": true, "formats": [ "markdown" ], "same_domain_only": true, "include_subdomains": true, "proxy": { "country": "string", "type": "datacenter" }, "wait_config": { "wait_until": "commit", "wait_for": "string", "wait_timeout": 30000 } }, "created_at": "2019-08-24T14:15:22Z", "completed_at": "2019-08-24T14:15:22Z" } ], "page": 0, "page_size": 0, "page_count": 0 } ``` *** ## Retrieve a Crawl Job Retrieve the latest status and results for a specific crawl job. ```http GET /v1/crawl/{job_id} ``` Once the crawl completes, this endpoint also returns the extracted page results. B["Queued"] B --> C["Running"] C --> D["Completed"] `} /> *** ## Job Status Each crawl job includes a `status` field that indicates its current state. | Status | Description | | ----------- | -------------------------------------------------------- | | `queued` | The crawl job has been accepted and is waiting to start. | | `running` | The crawler is actively processing pages. | | `completed` | The crawl has finished and the results are available. | Example: ```json { "job_id": "453bd7d7-5355-4d6d-a38e-d9e7eb218c3f", "status": "queued" } ``` *** ## Monitor Crawl Progress Each job includes progress information that can be used to track the crawl. | Field | Description | | ----------------- | -------------------------------------------- | | `total_pages` | Total number of pages included in the crawl. | | `completed_pages` | Number of pages successfully processed. | | `failed_pages` | Number of pages that failed during crawling. | Example: ```json { "total_pages": 20, "completed_pages": 15, "failed_pages": 1 } ``` Retrieve the job periodically to monitor its progress until the status becomes `completed`. *** ## Crawl Configuration Each listed job includes the configuration used when the crawl was created. | Field | Description | | -------------------- | -------------------------------------------------------------- | | `render_js` | Indicates whether JavaScript rendering was enabled. | | `formats` | Requested output formats. | | `same_domain_only` | Indicates whether crawling was limited to the starting domain. | | `include_subdomains` | Indicates whether subdomains were included. | | `proxy` | Proxy configuration used for the crawl. | | `wait_config` | Browser wait configuration used during extraction. | *** ## Pagination The response includes pagination information for the job list. | Field | Description | | ------------ | --------------------------------- | | `page` | Current page number. | | `page_size` | Number of jobs returned per page. | | `page_count` | Total number of available pages. | Example: ```json { "page": 0, "page_size": 10, "page_count": 3 } ``` *** ## Typical Workflow B B --> C C --> D D --> E E --> F `} /> *** ## Next Steps If you no longer need an active crawl job, continue to **Cancelling Crawl Jobs** to learn how to stop a running crawl. # Cancelling Crawl Jobs (/docs/scraper-api/guides/crawl/05_cancelling-crawl-jobs) If you no longer need a crawl job to continue processing, you can cancel it using its job ID. Cancelling a crawl job stops scheduling new crawl pages while allowing any pages already being processed to finish. *** ## Cancel a Crawl Job Cancel an existing crawl job. ```http DELETE /v1/crawl/{job_id} ``` If the cancellation request is accepted, the API returns information about the crawl job, including its current status, a URL for checking job status, and the number of pages that are still being processed. Example: ```json { "job_id": "453bd7d7-5355-4d6d-a38e-d9e7eb218c3f", "status": "queued", "status_url": "string", "in_flight_pages": 0 } ``` *** ## Response Fields | Field | Description | | ----------------- | ---------------------------------------------------------------------------------------------- | | `job_id` | Unique identifier of the crawl job. | | `status` | Current status of the crawl job. | | `status_url` | Endpoint that can be used to retrieve the latest status of the crawl job. | | `in_flight_pages` | Number of queued or processing pages that are still completing after the cancellation request. | *** ## What Happens After Cancellation? B B --> C C --> D D --> E `} /> After the cancellation request is accepted, you can continue monitoring the crawl job by retrieving its latest status using the `status_url` or the job ID. *** ## Unable to Cancel a Job If a crawl job cannot be cancelled in its current state, the API returns a `409 Conflict` response. For example, attempting to cancel a job that has already completed returns: ```json { "code": "INVALID_STATE", "message": "Crawl job 9293f045-0523-409e-a4e9-874abc663a9e is already completed and cannot be cancelled", "retryable": false } ``` ### Error Fields | Field | Description | | ----------- | -------------------------------------------------------- | | `code` | Machine-readable error code. | | `message` | Description of why the cancellation request failed. | | `retryable` | Indicates whether retrying the same request may succeed. | *** ## Possible Responses | Status Code | Description | | ---------------------- | ------------------------------------------------------- | | `202 Accepted` | Cancellation request accepted successfully. | | `401 Unauthorized` | Authentication failed. | | `404 Not Found` | The specified crawl job was not found. | | `409 Conflict` | The crawl job cannot be cancelled in its current state. | | `422 Validation Error` | The request contains invalid parameters. | *** # Handling Crawl API Errors (/docs/scraper-api/guides/crawl/06_error-handling) The Crawl API uses standard HTTP status codes together with structured JSON error responses. The exact response depends on the endpoint and the reason the request could not be completed. *** ## Standard Error Response Most Crawl endpoints return the following error format: ```json { "code": "string", "message": "string", "correlation_id": "string", "retryable": false } ``` ### Fields | Field | Description | | ---------------- | --------------------------------------------------- | | `code` | Machine-readable error code. | | `message` | Human-readable description of the error. | | `correlation_id` | Request correlation identifier when available. | | `retryable` | Indicates whether retrying the request may succeed. | *** ## HTTP Status Codes ### `POST /v1/crawl` | Status Code | Description | | ------------------------- | ----------------------------------- | | `202 Accepted` | Crawl job accepted successfully. | | `401 Unauthorized` | Authentication failed. | | `402 Payment Required` | Insufficient balance. | | `422 Validation Error` | The request failed validation. | | `429 Too Many Requests` | Request rate limit exceeded. | | `503 Service Unavailable` | Service is temporarily unavailable. | *** ### `GET /v1/crawl/jobs` | Status Code | Description | | ---------------------- | ---------------------------------- | | `200 OK` | Crawl jobs retrieved successfully. | | `401 Unauthorized` | Authentication failed. | | `422 Validation Error` | Invalid query parameters. | *** ### `GET /v1/crawl/{job_id}` | Status Code | Description | | ---------------------- | --------------------------------- | | `200 OK` | Crawl job retrieved successfully. | | `401 Unauthorized` | Authentication failed. | | `404 Not Found` | Crawl job was not found. | | `422 Validation Error` | Invalid request parameters. | *** ### `DELETE /v1/crawl/{job_id}` | Status Code | Description | | ---------------------- | ------------------------------------------------------- | | `202 Accepted` | Cancellation request accepted. | | `401 Unauthorized` | Authentication failed. | | `404 Not Found` | Crawl job was not found. | | `409 Conflict` | The crawl job cannot be cancelled in its current state. | | `422 Validation Error` | Invalid request parameters. | *** ## Validation Errors Requests that fail validation return an `HTTPValidationError` response. The response contains a `detail` array describing one or more validation issues. ```json { "detail": [ { "loc": [ "string" ], "msg": "string", "type": "string" } ] } ``` ### Validation Fields | Field | Description | | ------ | ------------------------------------ | | `loc` | Location of the validation error. | | `msg` | Description of the validation error. | | `type` | Validation error type. | *** ## Example: Invalid Job State If a crawl job cannot be cancelled because of its current state, the API returns a `409 Conflict` response. Example: ```json { "code": "INVALID_STATE", "message": "Crawl job 9293f045-0523-409e-a4e9-874abc663a9e is already completed and cannot be cancelled", "retryable": false } ``` ### Error Fields | Field | Description | | ----------- | --------------------------------------------------- | | `code` | Machine-readable error code. | | `message` | Explains why the request could not be completed. | | `retryable` | Indicates whether retrying the request may succeed. | *** ## Summary The Crawl API returns structured error responses together with standard HTTP status codes. Validation errors include additional details about invalid request parameters, while endpoint-specific errors provide information about why an operation could not be completed. Reviewing both the HTTP status code and the response body can help identify and resolve issues more efficiently. # Understanding Extraction (/docs/scraper-api/guides/extraction/01_understanding_extraction) import { Step, Steps } from "fumadocs-ui/components/steps"; import { Tabs, Tab } from "fumadocs-ui/components/tabs"; import { FileText, CodeXml } from "lucide-react"; The Extraction API converts webpages into clean, structured content that can be consumed by applications, AI systems, search pipelines, and automation workflows. Instead of downloading a webpage and manually parsing raw HTML, you can send a URL and receive the extracted content in a format that is easier to process. ## How It Works #### Submit a URL Send the URL of the webpage you want to extract. #### The Page Is Processed Geonode fetches the webpage and extracts the primary content. #### Output Is Generated The extracted content is returned in one or more supported output formats. #### Use the Result Store, analyze, search, or process the extracted content in your application. ## Extraction Endpoints The Extraction API consists of three endpoints. | Endpoint | Purpose | | -------------------------- | ----------------------------------------------------------- | | `POST /v1/extract` | Extract Markdown and/or HTML from a webpage. | | `GET /v1/extract/jobs` | List and filter previous extraction jobs. | | `GET /v1/extract/{job_id}` | Retrieve the status or result of a specific extraction job. | Most extraction workflows begin with `POST /v1/extract`. The remaining endpoints are primarily used to monitor and retrieve asynchronous extraction jobs. ## Processing Modes The Extraction API supports both synchronous and asynchronous processing. The request remains open until extraction is complete. The extracted content is returned directly in the response. The request immediately returns a `job_id`. The extraction continues in the background and the result can be retrieved later using `GET /v1/extract/{job_id}`. ## Extraction Workflow ### Synchronous ```text POST /v1/extract ↓ Extraction completes ↓ Content returned ``` ### Asynchronous ```text POST /v1/extract ↓ job_id returned ↓ GET /v1/extract/{job_id} ↓ Content returned ``` ## Output Formats The Extraction API supports Markdown, HTML, or both formats in a single request. Markdown HTML Both The extracted content is returned in the `data.markdown` field. ```json { "formats": ["markdown"] } ``` Markdown returns the extracted content as plain text with lightweight formatting. The extracted content is returned in the `data.html` field. ```json { "formats": ["html"] } ``` HTML returns the extracted content with a structure closer to the original webpage. The extracted content is returned in both the `data.markdown` and `data.html` fields. ```json { "formats": ["markdown", "html"] } ``` Both Markdown and HTML are returned in the same response. ## Next Steps Continue to **Your First Extraction** to send your first extraction request and retrieve content from a webpage. # Your First Extraction (/docs/scraper-api/guides/extraction/02_your_first_extraction) In this guide, you'll send your first extraction request and retrieve the content of a webpage. ## What You'll Build By the end of this guide, you'll be able to: * Send an extraction request * Extract content from a webpage * Understand the response structure * Access the extracted content ## Send Your First Request Use the following request: ```bash title="request.sh" curl -X POST "https://scraper.geonode.io/v1/extract" \ -H "X-Api-Key: YOUR_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "url": "https://example.com" }' ``` ## Request Breakdown | Field | Description | | ----- | ------------------------------------ | | `url` | The webpage to extract content from. | Since no `formats` field is provided, the API returns HTML by default. ## Understanding the Response A successful request returns the extracted content and metadata. ```json title="request.json" { "data": { "html": "..." }, "metadata": { "url": "http://example.com/", "render_js": false, "http_status": 200, "formats": ["html"], "processing_mode": "sync" }, "tokens_charged": 1 } ``` ## Access the Extracted Content The extracted page content is available in: ```text data.html ``` The metadata section contains additional information about the extraction, including: * Target URL * HTTP status * Output format * Processing mode * Extraction duration ## Success If you received a response similar to the example above, your first extraction was successful. Your API key is working, the Extraction API is accessible, and you're ready to start working with different output formats and extraction options. ## Next Steps Continue to **Working With Output Formats** to learn how to return Markdown, HTML, or both formats in a single request. # Link Extraction (/docs/scraper-api/guides/extraction/03_extracting_links) By default, the Extraction API returns the extracted page content. If you also need links found on the page, enable link extraction using the `extract_links` option. ## Enable Link Extraction Set `extract_links` to `true` in the extraction request. ### Request ```bash title="request.sh" curl -X POST "https://scraper.geonode.io/v1/extract" \ -H "X-Api-Key: YOUR_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "url": "https://quotes.toscrape.com/", "formats": ["markdown"], "extract_links": true }' ``` ## Response When link extraction is enabled, the response includes a `links` field inside `data`. ```json title="response.json" { "data": { "markdown": "...", "html": null, "links": [ "https://quotes.toscrape.com/login", "https://quotes.toscrape.com/author/Albert-Einstein", "https://www.goodreads.com/quotes" ] } } ``` ## Access Extracted Links The extracted links are available in: ```text title="response-path.txt" data.links ``` Each item in the array contains a URL discovered on the extracted page. ## When to Use Link Extraction Enable `extract_links` when you need to: * Collect links from a webpage * Discover related pages referenced by the content * Build URL lists for further processing * Analyze page relationships ## Important `extract_links` returns links found on the extracted page. It does not crawl those links or recursively discover additional pages. For example: ```text Page A ├─ Link B ├─ Link C └─ Link D ``` The API returns links B, C, and D. It does not visit those pages automatically. {/* ## Link Extraction vs Map API | Task | Recommended API | |--------|--------| | Extract content and links from a page | Extraction API | | Discover URLs across a website | Map API | | Crawl multiple pages | Map API | | Build a site inventory | Map API | Use the Map API when you need large-scale URL discovery before deciding which pages to extract. */} ## Success You now know how to return links alongside extracted content in a single extraction request. ## Next Steps Continue to **Extraction Workflows** to learn how different extraction options can be combined in real-world scenarios. # Checking Job Status (/docs/scraper-api/guides/extraction/04_checking_job_status) When you run an extraction in asynchronous mode, the API immediately returns a job ID instead of waiting for the extraction to finish. You can use that job ID to check the current status of the extraction and retrieve the result once processing is complete. ## How It Works The asynchronous extraction workflow follows these steps: 1. Start an extraction using `processing_mode: "async"`. 2. Receive a `job_id`. 3. Poll the job status endpoint. 4. Retrieve the extracted content when the job is completed. ## Get an Extraction Job Use the following endpoint to retrieve the current status of an extraction job. ```http GET /v1/extract/{job_id} ``` Replace `{job_id}` with the value returned by your extraction request. ## Example Request ```bash curl -X GET "https://scraper.geonode.io/v1/extract/4844831a-a222-4cac-b5e6-7e3f2dd07b48" \ -H "X-Api-Key: YOUR_API_KEY" ``` ## Response While Processing A job may still be running when you check its status. During this stage, the extracted content is not yet available. ```json { "job_id": "4844831a-a222-4cac-b5e6-7e3f2dd07b48", "status": "processing", "created_at": "2026-05-26T10:30:00Z", "completed_at": null, "data": null, "metadata": null, "error": null, "tokens_charged": null } ``` ### What This Means | Field | Description | | ---------------- | ---------------------------------------------- | | `status` | Current state of the extraction job | | `created_at` | Time the job was created | | `completed_at` | `null` until processing finishes | | `data` | Extracted content, available after completion | | `metadata` | Extraction details, available after completion | | `error` | Error information if the job fails | | `tokens_charged` | Token usage after processing completes | ## Response After Completion Once the extraction finishes successfully, the response includes the extracted content and metadata. ```json { "job_id": "4844831a-a222-4cac-b5e6-7e3f2dd07b48", "status": "completed", "created_at": "2026-05-26T10:30:00Z", "completed_at": "2026-05-26T10:30:04Z", "data": { "markdown": "# Example Page Content" }, "metadata": { "url": "https://docs.python.org/3/library/json.html", "render_js": false, "http_status": 200, "duration_ms": 631, "formats": ["markdown"], "processing_mode": "async" }, "error": null, "tokens_charged": 1 } ``` ## Job Status Values The API can return the following job statuses. | Status | Description | | ------------ | ------------------------------------------------- | | `queued` | The job has been accepted and is waiting to start | | `processing` | The extraction is currently running | | `completed` | The extraction finished successfully | | `failed` | The extraction could not be completed | | `cancelled` | The job was cancelled before completion | ## Starting an Async Extraction To use this endpoint, first create an asynchronous extraction job. ```bash curl -X POST "https://scraper.geonode.io/v1/extract" \ -H "X-Api-Key: YOUR_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "url": "https://docs.python.org/3/library/json.html", "formats": ["markdown"], "processing_mode": "async" }' ``` ## Async Job Response The extraction endpoint returns a job ID that can be used for polling. ```json { "job_id": "4844831a-a222-4cac-b5e6-7e3f2dd07b48", "status": "queued", "status_url": "/v1/extract/4844831a-a222-4cac-b5e6-7e3f2dd07b48", "estimated_tokens": 1 } ``` ## Polling for Results A common pattern is to periodically check the job status until the extraction completes. ```text POST /v1/extract ↓ Receive job_id ↓ GET /v1/extract/{job_id} ↓ status = processing ↓ GET /v1/extract/{job_id} ↓ status = completed ↓ Read extracted content ``` ## Next Step Now that you can monitor individual extraction jobs, the next guide explains how to view and filter multiple extraction jobs using the Jobs endpoint. # Listing Extraction Jobs (/docs/scraper-api/guides/extraction/05_listing_extraction_jobs) Use the Jobs endpoint to view extraction jobs associated with your account. This endpoint is useful when you need to find a previous job, retrieve a job ID, monitor running jobs, or review extraction history. ## List Extraction Jobs Use the following endpoint to retrieve extraction jobs. ```http GET /v1/extract/jobs ``` ### Request ```bash title="request.sh" curl -X GET "https://scraper.geonode.io/v1/extract/jobs" \ -H "X-Api-Key: YOUR_API_KEY" ``` ### Response ```json title="response.json" { "jobs": [ { "job_id": "0131f784-9037-4fdd-af9e-b8d445fe2d5f", "status": "completed", "url": "https://geonode.com/", "created_at": "2026-06-13T18:14:54.926869Z", "execution_time": 2297, "output": [ "html" ] } ], "page": 1, "page_size": 100, "page_count": 1 } ``` ## Job Fields Each job contains summary information about an extraction request. | Field | Description | | ---------------- | ----------------------------------------- | | `job_id` | Unique identifier for the extraction job | | `status` | Current status of the extraction | | `url` | URL that was extracted | | `created_at` | Time the job was created | | `execution_time` | Processing time in milliseconds | | `output` | Output formats returned by the extraction | ## Filter Jobs ### Filter by Status Retrieve jobs with a specific status. ```bash title="request.sh" curl -X GET "https://scraper.geonode.io/v1/extract/jobs?status=completed" \ -H "X-Api-Key: YOUR_API_KEY" ``` Supported values: ```text queued processing completed failed cancelled ``` ### Filter by Output Format Retrieve jobs that generated a specific output format. ```bash title="request.sh" curl -X GET "https://scraper.geonode.io/v1/extract/jobs?output=html" \ -H "X-Api-Key: YOUR_API_KEY" ``` Supported values: ```text html markdown ``` ### Filter by URL Retrieve jobs created for a specific URL. ```bash title="request.sh" curl -X GET "https://scraper.geonode.io/v1/extract/jobs?url=https://geonode.com/" \ -H "X-Api-Key: YOUR_API_KEY" ``` ### Filter by Job ID Retrieve a specific job from the results. ```bash title="request.sh" curl -X GET "https://scraper.geonode.io/v1/extract/jobs?job_id=0131f784-9037-4fdd-af9e-b8d445fe2d5f" \ -H "X-Api-Key: YOUR_API_KEY" ``` ### Filter by Date Range Retrieve jobs created within a specific time period. ```bash title="request.sh" curl -X GET "https://scraper.geonode.io/v1/extract/jobs?start_date=2026-06-01&end_date=2026-06-30" \ -H "X-Api-Key: YOUR_API_KEY" ``` ## Pagination Use pagination when working with a large number of extraction jobs. ### Request ```bash title="request.sh" curl -X GET "https://scraper.geonode.io/v1/extract/jobs?page=1&page_size=5" \ -H "X-Api-Key: YOUR_API_KEY" ``` ### Pagination Fields | Field | Description | | ------------ | -------------------------------- | | `page` | Current page number | | `page_size` | Number of jobs returned per page | | `page_count` | Total number of available pages | ## Retrieve Full Job Details The Jobs endpoint returns summary information. To retrieve the complete extraction result, use: ```http GET /v1/extract/{job_id} ``` Example: ```bash title="request.sh" curl -X GET "https://scraper.geonode.io/v1/extract/0131f784-9037-4fdd-af9e-b8d445fe2d5f" \ -H "X-Api-Key: YOUR_API_KEY" ``` ## Common Use Cases Use this endpoint to: * Find a lost job ID * Review extraction history * Monitor running jobs * View completed extractions * Search jobs by URL * Retrieve recent extraction activity ## Success You now know how to list, search, and filter extraction jobs. ## Next Steps Continue to **Extraction Workflows** to learn how extraction features can be combined in real-world scenarios. # Extraction Workflows (/docs/scraper-api/guides/extraction/06_extraction_workflows) The Extraction API supports a variety of options that can be combined depending on your use case. This guide demonstrates common extraction workflows and the recommended settings for each scenario. ## Standard Website Extraction Use the default extraction settings when working with traditional websites that do not rely heavily on JavaScript. ### Request ```bash title="request.sh" curl -X POST "https://scraper.geonode.io/v1/extract" \ -H "X-Api-Key: YOUR_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "url": "https://example.com" }' ``` ### Best For * Blogs * News websites * Documentation sites * Static webpages *** ## AI and RAG Workflows Markdown is often the preferred format when content will be processed by AI systems. ### Request ```bash title="request.sh" curl -X POST "https://scraper.geonode.io/v1/extract" \ -H "X-Api-Key: YOUR_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "url": "https://docs.python.org/3/library/json.html", "formats": ["markdown"] }' ``` ### Best For * Vector databases * RAG pipelines * Embedding generation * Knowledge bases * LLM applications *** ## JavaScript-Powered Websites Some websites render content in the browser and require JavaScript execution before extraction. ### Request ```bash title="request.sh" curl -X POST "https://scraper.geonode.io/v1/extract" \ -H "X-Api-Key: YOUR_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "url": "https://geonode.com/", "render_js": true }' ``` ### Best For * React applications * Next.js websites * Vue applications * Single-page applications (SPAs) *** ## Geo-Targeted Extraction Use a proxy configuration when content changes based on visitor location. ### Request ```bash title="request.sh" curl -X POST "https://scraper.geonode.io/v1/extract" \ -H "X-Api-Key: YOUR_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "url": "https://example.com", "proxy": { "country": "DE", "type": "residential" } }' ``` ### Best For * Local search results * Country-specific pricing * Regional content * Localized webpages *** ## Content and Link Extraction Extract page content and collect links from the same request. ### Request ```bash title="request.sh" curl -X POST "https://scraper.geonode.io/v1/extract" \ -H "X-Api-Key: YOUR_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "url": "https://quotes.toscrape.com/", "formats": ["markdown"], "extract_links": true }' ``` ### Best For * URL discovery * Content analysis * Link collection * Website research *** ## Large Async Extractions Use asynchronous processing for pages that may take longer to extract. ### Start the Extraction ```bash title="request.sh" curl -X POST "https://scraper.geonode.io/v1/extract" \ -H "X-Api-Key: YOUR_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "url": "https://geonode.com/", "render_js": true, "processing_mode": "async" }' ``` ### Response ```json title="response.json" { "job_id": "4844831a-a222-4cac-b5e6-7e3f2dd07b48", "status": "queued", "status_url": "/v1/extract/4844831a-a222-4cac-b5e6-7e3f2dd07b48" } ``` ### Check Job Status ```bash title="request.sh" curl -X GET "https://scraper.geonode.io/v1/extract/4844831a-a222-4cac-b5e6-7e3f2dd07b48" \ -H "X-Api-Key: YOUR_API_KEY" ``` ### Best For * Large pages * Slow websites * Background processing * High-volume workflows *** ## Recommended Settings | Scenario | Recommended Configuration | | ------------------------ | -------------------------- | | Standard webpage | Default request | | AI and RAG workflows | `formats: ["markdown"]` | | JavaScript websites | `render_js: true` | | Country-specific content | `proxy` | | Link discovery | `extract_links: true` | | Long-running extractions | `processing_mode: "async"` | ## Success You now know how to combine extraction features for common real-world workflows. ## Next Steps Continue to **Best Practices** to learn how to improve extraction performance, reliability, and efficiency. # Best Practices (/docs/scraper-api/guides/extraction/07_best_practices) The following recommendations can help improve extraction results and reduce unnecessary processing. ## Choose the Right Output Format Use the output format that matches your use case. | Format | Best For | | ---------- | ---------------------------------------------------------- | | `markdown` | AI workflows, RAG pipelines, indexing, and text processing | | `html` | Preserving page structure and rendering content | | Both | Applications that require both formats | Requesting only the formats you need can reduce response size. ## Enable JavaScript Rendering Only When Needed JavaScript rendering increases extraction time because the page must be rendered before content can be extracted. Use: ```json title="request.json" { "render_js": true } ``` only for websites that depend on client-side rendering. Common examples include: * React * Next.js * Vue * Single-page applications (SPAs) ## Use Asynchronous Processing for Large Workloads For long-running extractions, use asynchronous processing. ```json title="request.json" { "processing_mode": "async" } ``` This prevents request timeouts and allows your application to continue processing while extraction runs in the background. ## Use Geo-Targeting Only When Required Proxy routing may increase processing time. Only specify a country when content differs by location. ```json title="request.json" { "proxy": { "country": "DE", "type": "residential" } } ``` ## Reuse Job IDs When using asynchronous extraction: 1. Create the extraction job once. 2. Store the returned `job_id`. 3. Poll the job status endpoint. Avoid creating duplicate extraction jobs for the same request. ## Use Link Extraction Only When Needed ```json title="request.json" { "extract_links": true } ``` Enable link extraction only when you need URLs from the page. This keeps responses smaller and easier to process. ## Monitor Job Status Before retrieving extraction results, check the job status. ```http GET /v1/extract/{job_id} ``` Wait until the status becomes: ```text completed ``` before processing the result. ## Store Extracted Content If content does not change frequently, consider storing extraction results instead of repeatedly extracting the same page. This can reduce costs and improve performance. ## Success You now know the recommended practices for building reliable extraction workflows. ## Next Steps Continue to **Common Errors** to learn how to troubleshoot common extraction issues. # Common Errors (/docs/scraper-api/guides/extraction/08_common_errors) The following issues are commonly encountered when working with the Extraction API. ## Missing API Key Requests must include the `X-Api-Key` header. ### Incorrect ```bash title="request.sh" curl -X POST "https://scraper.geonode.io/v1/extract" ``` ### Correct ```bash title="request.sh" curl -X POST "https://scraper.geonode.io/v1/extract" \ -H "X-Api-Key: YOUR_API_KEY" ``` ## Invalid URL The `url` field must contain a valid URL. ### Incorrect ```json title="request.json" { "url": "example" } ``` ### Correct ```json title="request.json" { "url": "https://example.com" } ``` ## Job Still Processing When using asynchronous extraction, the result may not be available immediately. ### Response ```json title="response.json" { "status": "processing", "data": null } ``` Wait until the job status becomes: ```text completed ``` before attempting to use the extracted content. ## Job Not Found A job may not exist or may belong to a different account. ### Request ```http GET /v1/extract/{job_id} ``` Verify that the job ID is correct and that it was created using the same API key. ## Empty or Unexpected Results Some websites require JavaScript rendering before content becomes available. Try enabling: ```json title="request.json" { "render_js": true } ``` This is common for: * React applications * Next.js websites * Vue applications * Single-page applications ## Geo-Targeted Content Is Different Than Expected Some websites return different content based on location. Specify a country explicitly: ```json title="request.json" { "proxy": { "country": "US", "type": "residential" } } ``` ## Custom Headers Not Applied Headers must be passed inside the `headers` object. ### Correct ```json title="request.json" { "headers": { "Accept-Language": "en-US,en;q=0.9" } } ``` Do not place your Geonode API key inside the `headers` object. Authentication must use: ```text X-Api-Key ``` ## Link Extraction Does Not Crawl Websites `extract_links` only returns links found on the extracted page. ```json title="request.json" { "extract_links": true } ``` It does not visit discovered links or recursively crawl a website. For large-scale URL discovery, use the Map API. ## Using JavaScript Rendering Unnecessarily JavaScript rendering increases extraction time. ```json title="request.json" { "render_js": true } ``` Enable it only when a website requires client-side rendering. ## Need More Help? If you continue to experience issues: * Verify the request payload * Verify the target URL * Check job status for asynchronous requests * Review response metadata * Contact the Geonode support team ## Success You now understand the most common Extraction API issues and how to resolve them. ## What's Next You have completed the Extraction guides. Continue to the next API section to learn about additional scraping and data collection capabilities. # Listing Jobs (/docs/scraper-api/guides/jobs/01_listing_jobs) The Scraper API allows you to retrieve previously created jobs. Listing jobs is useful for monitoring activity, reviewing completed requests, or locating a specific job ID. ## List Jobs Use the list jobs endpoint provided by the API. For example: ```http GET /v1/extract/jobs ``` or ```http GET /v1/batch/jobs ``` or ```http GET /v1/crawl/jobs ``` The response returns a paginated collection of jobs. ## Filter Results Most job endpoints support filtering and pagination. Common query parameters include: | Parameter | Description | | ------------ | ------------------------------------------- | | `status` | Return only jobs with a specific status. | | `start_date` | Return jobs created after a specific date. | | `end_date` | Return jobs created before a specific date. | | `page` | Page number to retrieve. | | `page_size` | Number of jobs returned per page. | Use these filters to quickly locate the jobs you need. ## Pagination Large job histories are split into pages. Increase or decrease `page_size` depending on how many results you want returned in each request. ## Common Use Cases Listing jobs is useful for: * Viewing recent activity * Finding completed jobs * Monitoring failed jobs * Reviewing processing history * Building dashboards or reporting tools # Checking Job Status (/docs/scraper-api/guides/jobs/02_checking_job_status) Some Scraper API endpoints process requests asynchronously. Instead of returning the final result immediately, the API creates a job and returns a unique job ID. You can use that job ID to check the progress of the request. Common job statuses include: | Status | Description | | ------------ | -------------------------------------------------------- | | `queued` | The job is waiting to be processed. | | `processing` | The request is currently being processed. | | `completed` | The job finished successfully and results are available. | | `failed` | The request could not be completed. | ## Check a Job Use the endpoint associated with your API to retrieve the latest status of a job. For example: ```http GET /v1/extract/jobs/{job_id} ``` or ```http GET /v1/batch/jobs/{job_id} ``` or ```http GET /v1/crawl/jobs/{job_id} ``` The response contains the current status and any available results. ## Typical Workflow ```text Create Request │ ▼ Receive Job ID │ ▼ Check Job Status │ ▼ Completed │ ▼ Read Results ``` Continue checking the job until its status changes to `completed` or `failed`. ## When to Use Checking job status is useful when: * Processing large requests * Crawling multiple pages * Running batch operations * Using asynchronous processing modes # Output Formats (/docs/scraper-api/guides/making-requests/02_output_formats) import { Tabs, Tab } from "fumadocs-ui/components/tabs"; import { FileText, CodeXml } from "lucide-react"; The Extraction API can return content in Markdown, HTML, or both formats in a single request. All examples in this guide use the following endpoint: ```http POST /v1/extract ``` The `formats` field controls which output formats are returned by the extraction request. ```json { "formats": ["markdown"] } ``` If the `formats` field is omitted, the API returns HTML by default. ## Output Formats Markdown HTML Both Markdown returns the extracted content as clean, readable text. ### Request ```bash title="request.sh" curl -X POST "https://scraper.geonode.io/v1/extract" \ -H "X-Api-Key: YOUR_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "url": "https://example.com", "formats": ["markdown"] }' ``` ### Response ```json title="response.json" { "data": { "markdown": "# Example Domain..." } } ``` The extracted content is available in the `data.markdown` field. Common use cases: * AI and LLM workflows * Search indexing * Knowledge bases * Text processing pipelines HTML returns the extracted content with a structure closer to the original webpage. ### Request ```bash curl -X POST "https://scraper.geonode.io/v1/extract" \ -H "X-Api-Key: YOUR_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "url": "https://example.com", "formats": ["html"] }' ``` ### Response ```json { "data": { "html": "..." } } ``` The extracted content is available in the `data.html` field. Common use cases: * Rendering content in applications * Preserving page structure * Working with HTML elements * Content transformation workflows Request both Markdown and HTML when your application needs both formats from the same extraction. ### Request ```bash curl -X POST "https://scraper.geonode.io/v1/extract" \ -H "X-Api-Key: YOUR_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "url": "https://example.com", "formats": ["markdown", "html"] }' ``` ### Response ```json { "data": { "markdown": "# Example Domain...", "html": "..." } } ``` The extracted content is available in both the `data.markdown` and `data.html` fields. ## Choosing the Right Format | Format | Best For | | -------- | -------------------------------------------------------------------- | | Markdown | AI workflows, search indexing, knowledge bases, and text processing | | HTML | Preserving page structure and rendering content | | Both | Applications that need both representations from a single extraction | ## Success You now know how to control the format returned by the Extraction API. Whether you need Markdown, HTML, or both, you can choose the format that best fits your workflow. ## Next Steps Continue to **Extracting JavaScript Websites** to learn how to extract content from pages that rely on client-side rendering. # JavaScript Rendering (/docs/scraper-api/guides/making-requests/03_javascript-rendering) Many modern websites load content after the initial page request using JavaScript. When this happens, a standard extraction may return incomplete content because the page has not finished rendering. To handle these websites, enable JavaScript rendering with the `render_js` option. ## When to Use JavaScript Rendering Enable JavaScript rendering when: * Important content is missing from the extraction result * Content appears after the page loads * The website relies on client-side rendering ## Enable JavaScript Rendering Set `render_js` to `true` in your extraction request. ```bash title="request.sh" curl -X POST "https://scraper.geonode.io/v1/extract" \ -H "X-Api-Key: YOUR_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "url": "https://geonode.com/", "render_js": true }' ``` ### Request Breakdown | Field | Description | | ----------- | -------------------------------------------------------- | | `url` | The webpage to extract content from. | | `render_js` | Renders the page in a browser before extracting content. | ## Response A successful response includes the rendered page content and extraction metadata. ```json title="response.json" { "data": { "html": "..." }, "metadata": { "url": "https://geonode.com/", "render_js": true, "http_status": 200 } } ``` Notice that `metadata.render_js` is set to `true`, confirming that browser rendering was used during extraction. ## Things to Consider JavaScript rendering provides more complete extraction results for dynamic websites, but it may: * Take longer than a standard extraction * Use additional browser resources * Be unnecessary for static websites For static websites, leave `render_js` disabled for faster extraction. ## Success You can now extract content from websites that rely on JavaScript to display their content. ## Next Steps Continue to **Processing Modes** to learn the difference between synchronous and asynchronous extraction requests. # Waiting for Dynamic Content (/docs/scraper-api/guides/making-requests/04_waiting_for_dynamic_content) Modern websites often load content after the initial page load using JavaScript. If content appears a few seconds later, extracting immediately may return incomplete results. The `wait_config` parameter gives you control over when the extraction should begin. ## How wait\_config Works When a `wait_config` is provided, the extraction process follows this order: ```text wait_until ↓ wait_for ↓ wait_timeout ↓ Extract Content ``` This allows you to wait for page events, specific elements, or additional delays before extraction starts. ## wait\_until The `wait_until` option controls which browser lifecycle event must complete before moving to the next step. ### commit Wait until the browser receives the response headers and commits the navigation. ```json { "url": "https://example.com", "wait_config": { "wait_until": "commit" } } ``` Best for: * Very fast extractions * Cases where you only need the initial response ### domcontentloaded Wait until the HTML is parsed and the DOM is ready. ```json { "url": "https://example.com", "wait_config": { "wait_until": "domcontentloaded" } } ``` Best for: * Most websites * Pages where content is already present in the HTML > This is the default value when `wait_until` is not provided. ### load Wait until the page and all resources are fully loaded. ```json { "url": "https://example.com", "wait_config": { "wait_until": "load" } } ``` Best for: * Pages that depend on images or external scripts * Slower websites that require additional loading time ### networkidle Wait until there is no network activity for 500 milliseconds. ```json { "url": "https://example.com", "wait_config": { "wait_until": "networkidle" } } ``` Best for: * Single-page applications (SPA) * React, Vue, Angular, and Next.js websites * Dynamic content loaded through API requests ## wait\_for The `wait_for` option waits until a specific element appears on the page before extraction starts. ### Using a CSS Selector ```json { "url": "https://example.com", "wait_config": { "wait_for": ".product-grid" } } ``` Extraction begins only after an element matching `.product-grid` is found. ### Using XPath ```json { "url": "https://example.com", "wait_config": { "wait_for": "//div[@class='product-grid']" } } ``` XPath expressions must start with: ```text // ``` or ```text xpath= ``` Otherwise, the value is treated as a CSS selector. ### Common Examples | Use Case | Selector | | ------------------ | ----------------- | | Product listing | `.product-grid` | | Search results | `.search-results` | | Article content | `article` | | Table data | `table` | | Loading completion | `.loaded` | ## wait\_timeout The `wait_timeout` option adds an additional delay after all previous waits complete. ```json { "url": "https://example.com", "wait_config": { "wait_for": ".product-grid", "wait_timeout": 5000 } } ``` In this example: 1. Wait for `.product-grid` 2. Wait an additional 5 seconds 3. Extract content ### When to Use wait\_timeout Use this option when content continues updating after the target element appears. Common examples include: * Infinite scrolling pages * Late-loading advertisements * Client-side rendering delays * Dynamic dashboards ## Complete Example ```json { "url": "https://example.com", "formats": ["markdown"], "wait_config": { "wait_until": "networkidle", "wait_for": ".product-grid", "wait_timeout": 3000 } } ``` This configuration: 1. Waits until network activity stops 2. Waits for `.product-grid` to appear 3. Waits an additional 3 seconds 4. Extracts the content ## Browser Rendering Behavior When a non-null `wait_config` is provided, browser rendering is automatically enabled if `render_js` is not explicitly specified. Example: ```json { "url": "https://example.com", "wait_config": { "wait_until": "networkidle" } } ``` The extraction will automatically use browser rendering. ## Invalid Configuration The following request is rejected: ```json { "url": "https://example.com", "render_js": false, "wait_config": { "wait_until": "networkidle" } } ``` This is considered ambiguous because `wait_config` requires browser rendering while `render_js` explicitly disables it. ## Best Practices * Use `domcontentloaded` for most websites. * Use `networkidle` for JavaScript-heavy applications. * Use `wait_for` when a specific element contains the data you need. * Use `wait_timeout` only when additional rendering time is required. * Avoid excessive delays, as they increase extraction time and cost. ## Next Steps Now that you understand how to wait for dynamic content, learn how to work with synchronous and asynchronous extraction workflows. # Processing Modes (/docs/scraper-api/guides/making-requests/05_processing_modes) import { Tabs, Tab } from "fumadocs-ui/components/tabs"; The Extraction API supports two processing modes: * `sync` for immediate results * `async` for background processing Use the `processing_mode` field to control how extraction requests are handled. ```json { "processing_mode": "sync" } ``` If `processing_mode` is omitted, the API uses `sync` mode by default. ## Processing Modes Synchronous mode waits for extraction to complete before returning a response. ### Request ```bash title="request.sh" curl -X POST "https://scraper.geonode.io/v1/extract" \ -H "X-Api-Key: YOUR_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "url": "https://example.com", "processing_mode": "sync" }' ``` ### Response ```json title="response.json" { "data": { "html": "...", "markdown": null }, "metadata": { "url": "http://example.com/", "http_status": 200, "processing_mode": "sync" }, "tokens_charged": 1 } ``` The request remains open until extraction is complete and the content is returned in the response. Asynchronous mode immediately creates an extraction job and returns a job ID. ### Request ```bash title="request.sh" curl -X POST "https://scraper.geonode.io/v1/extract" \ -H "X-Api-Key: YOUR_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "url": "https://geonode.com/", "processing_mode": "async" }' ``` ### Response ```json title="response.json" { "job_id": "32561cfc-4d87-4a46-af4a-a10e5f3168b9", "status": "queued", "status_url": "/v1/extract/32561cfc-4d87-4a46-af4a-a10e5f3168b9", "estimated_tokens": 1 } ``` The extraction continues in the background while your application continues running. Use the returned `job_id` or `status_url` to retrieve the extraction result later. ## Choosing a Processing Mode | Mode | Best For | | ----- | -------------------------------------------------------------------- | | Sync | Interactive applications, quick extractions, and immediate results | | Async | Background processing, large workloads, and long-running extractions | ## Success You now know how to choose between synchronous and asynchronous extraction requests. Synchronous mode returns the extracted content immediately, while asynchronous mode returns a job ID that can be used to retrieve the result later. ## Next Steps Continue to **Checking Job Status** to learn how to retrieve the status and results of an asynchronous extraction job using `GET /v1/extract/{job_id}`. # Proxy and Geo-Targeting (/docs/scraper-api/guides/making-requests/06_proxy_and_geo_targeting) The Scraper API can route extraction requests through Geonode proxies. If you do not provide a `proxy` object, the API uses residential proxies by default and automatically determines the most appropriate routing location when possible. Geo-targeting is useful when a website returns different content, language, pricing, availability, or search results based on the visitor's location. ## Configure a Proxy Use the `proxy` object to control the country and proxy type used for the extraction request. ### Request ```bash title="request.sh" curl -X POST "https://scraper.geonode.io/v1/extract" \ -H "X-Api-Key: YOUR_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "url": "https://example.com", "proxy": { "country": "US", "type": "residential" } }' ``` ## Country Codes The `proxy.country` field uses standard ISO 3166-1 alpha-2 country codes. | Country | Code | | -------------- | ---- | | United States | `US` | | United Kingdom | `GB` | | Germany | `DE` | | Pakistan | `PK` | | Canada | `CA` | | France | `FR` | For the complete list of supported country codes: [https://en.wikipedia.org/wiki/ISO\_3166-1\_alpha-2](https://en.wikipedia.org/wiki/ISO_3166-1_alpha-2) ## Proxy Types ### Residential Routes requests through residential IP addresses. ```json title="request.json" { "proxy": { "type": "residential" } } ``` ### Datacenter Routes requests through datacenter IP addresses. ```json title="request.json" { "proxy": { "type": "datacenter" } } ``` ### Mix Allows the API to use a combination of available proxy networks. ```json title="request.json" { "proxy": { "type": "mix" } } ``` ## Proxy Configuration Reference | Field | Type | Description | | --------------- | -------------- | --------------------------------------------------------------------------------------------------- | | `proxy.country` | string or null | Two-letter ISO country code such as `US`, `GB`, `DE`, or `PK`. | | `proxy.type` | string | Proxy type. Supported values are `residential`, `datacenter`, and `mix`. Defaults to `residential`. | ## Verify the Applied Proxy The proxy configuration used for the request is returned in the response metadata. ### Response ```json title="response.json" { "data": { "html": "..." }, "metadata": { "url": "http://example.com/", "proxy": { "country": "US", "type": "residential" }, "processing_mode": "sync" }, "tokens_charged": 1 } ``` You can inspect `metadata.proxy.country` and `metadata.proxy.type` to verify which proxy configuration was applied during extraction. ## Success You now know how to control the country and proxy network used for extraction requests. ## Next Steps Continue to **Using Custom Headers** to learn how to send additional HTTP headers with extraction requests. # Using Custom Headers (/docs/scraper-api/guides/making-requests/07_using_custom_headers) Use the `headers` object to send custom HTTP headers with the extraction request. ## Important ```json title="request.json" { "url": "https://example.com", "headers": { "Accept-Language": "en-US,en;q=0.9" } } ``` Do not put your Geonode API key inside `headers`. Authentication belongs in the `X-Api-Key` request header sent to the Scraper API. ## Send Custom Headers Use the `headers` object to include one or more HTTP headers in your extraction request. ### Request ```bash title="request.sh" curl -X POST "https://scraper.geonode.io/v1/extract" \ -H "X-Api-Key: YOUR_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "url": "https://geonode.com/", "headers": { "Accept-Language": "en-US,en;q=0.9", "User-Agent": "Mozilla/5.0" } }' ``` ## Common Headers | Header | Purpose | | ----------------- | ------------------------------------------------- | | `Accept-Language` | Request content in a specific language | | `User-Agent` | Identify the browser or client making the request | | `Referer` | Indicate the page that initiated the request | | `Cookie` | Send session or authentication cookies | ## Multiple Headers You can send multiple headers in a single request. ```json title="request.json" { "url": "https://geonode.com/", "headers": { "Accept-Language": "en-US,en;q=0.9", "User-Agent": "Mozilla/5.0" } } ``` ## Success You now know how to send custom HTTP headers during extraction requests. ## Next Steps Continue to **Link Extraction** to learn how to return links found on a webpage along with its extracted content. # Understanding Map (/docs/scraper-api/guides/map/00_understanding_map) The Map API helps you discover URLs across a website without extracting the content of each page. Instead of downloading and processing every page, the Map API discovers URLs from a website and returns them as a structured list. This makes it useful for exploring websites, planning scraping workflows, and identifying the pages you want to process later. *** ## How the Map API Works The Map API starts with a single website URL and discovers additional URLs from the website. During discovery, the Map API can collect URLs from: * Website sitemaps. * Links discovered in HTML pages. Each discovered URL includes information about how it was found, such as `sitemap` or `html`. *** ## Map vs Crawl vs Extraction Although these APIs work together, they solve different problems. | Feature | Map | Crawl | Extraction | | -------------------------------- | :-: | :---: | :--------: | | Discover website URLs | ✓ | ✓ | — | | Extract page content | — | ✓ | ✓ | | Process a single page | — | — | ✓ | | Process multiple pages | ✓ | ✓ | — | | Return a list of discovered URLs | ✓ | — | — | Use the Map API to discover pages first. Once you've identified the pages you need, use the Extraction API to extract individual pages or the Crawl API to process larger sections of the website. *** ## Common Use Cases The Map API is useful for: * Building a list of URLs from a website. * Discovering documentation, blog, or support pages. * Preparing URLs before extraction or crawling. * Exploring the content available on a website. * Identifying pages for further processing or analysis. *** ## When Should You Use the Map API? Choose the Map API when your goal is to discover **where content exists**, rather than extracting the content itself. | Goal | Recommended API | | ----------------------------------------- | --------------- | | Discover available pages | Map API | | Extract content from one page | Extraction API | | Extract content from many connected pages | Crawl API | *** ## Next Steps Now that you understand what the Map API does, continue to **Map Workflows** to learn how to use the Map endpoint as part of a complete URL discovery workflow. # Map Workflows (/docs/scraper-api/guides/map/01_map_workflows) The Map API is typically the first step in a scraping workflow. It helps you discover URLs from a website so you can decide which pages to process next. Instead of extracting page content, the Map API returns a list of discovered URLs that can be used with other Scraper API endpoints. *** ## Typical Workflow Most applications use the Map API as part of the following workflow. B[POST /v1/map] B-->C[Review Discovered URLs] C-->D[Extract Selected Pages] C-->E[Crawl Website] C-->F[Export URL List] `} /> The workflow starts with a website URL. After reviewing the discovered URLs, you can decide whether to extract specific pages, crawl larger sections of the website, or export the URL list for further processing. *** ## Step 1 — Submit a Website Start by sending a request to the Map endpoint. ```http POST /v1/map ``` Provide the website URL that you want to discover. The Map API returns a response containing the discovered URLs. *** ## Step 2 — Review the Results The response includes a list of discovered URLs. Each URL also indicates how it was discovered, such as: * `html` * `sitemap` Review the results before deciding which pages should be processed further. *** ## Step 3 — Choose Your Next Step After reviewing the discovered URLs, choose the workflow that best fits your application. B[Extract Individual Pages] A-->C[Crawl Multiple Pages] A-->D[Export URL Inventory] `} /> Each option serves a different purpose: | Next Step | When to Use | | -------------- | ---------------------------------------------------------------------- | | Extraction API | Process individual pages and extract structured content. | | Crawl API | Process larger sections of a website. | | Export URLs | Save the discovered URLs for reporting, analysis, or another workflow. | *** ## Example Workflow The following example shows how the Map API fits into a typical scraping pipeline. B[Map API] -->C[Discovered URLs] -->D[Filter Documentation Pages] D-->E[Extraction API] `} /> Rather than processing every page, the application first discovers all available URLs, filters the pages it needs, and then extracts content only from the relevant pages. *** ## Building a Processing Pipeline Many applications use the Map API as the first stage of a larger workflow. B[Discovered URLs] B-->C[Filter URLs] C-->D[Extraction API] C-->E[Crawl API] C-->F[Store Results] `} /> Separating URL discovery from content extraction gives you greater control over what your application processes. *** ## Best Practices * Start with the Map API before extracting or crawling a large website. * Review the discovered URLs before processing them. * Use the `search` parameter to narrow the discovered URLs when appropriate. * Use the Extraction API for individual pages. * Use the Crawl API when processing larger sections of a website. * Export or store discovered URLs if they will be reused later. *** ## End-to-End Workflow | Step | Action | Outcome | | ---- | ----------------------------------- | ----------------------------- | | 1 | Submit a website to the Map API | Discover available URLs | | 2 | Review the discovered URLs | Identify relevant pages | | 3 | Filter the results if needed | Reduce unnecessary processing | | 4 | Extract or crawl the selected pages | Process the website content | *** ## Next Steps Now that you understand how the Map API fits into a complete workflow, continue to **Your First Map** to create your first request and explore the request and response in detail. # Your First Map (/docs/scraper-api/guides/map/02_your_first_map) In this guide, you'll use the Map endpoint to discover URLs under a website. By the end of this guide, you'll know how to: * Send your first Map request. * Understand the request body. * Read the response. * Work with the discovered URLs. *** ## Before You Begin Before using the Map API, make sure you have: * A valid Geonode API key. * The Scraper API base URL. * A website URL to map. The Map API discovers URLs by combining sitemap parsing with HTML link extraction from the provided website. *** ## Step 1 — Send a Map Request Send a POST request to the Map endpoint. ```http POST /v1/map ``` The minimum request only requires a website URL. ```json { "url": "https://geonode.com" } ``` B["POST /v1/map"] -->C["Map API"] -->D["Discovered URLs"] `} /> *** ## Step 2 — Understand the Request The Map endpoint accepts the following request fields. | Field | Required | Description | | ------------------------- | -------- | --------------------------------------------------------------------------- | | `url` | Yes | Base URL to discover links from. | | `search` | No | Filters discovered URLs using a case-insensitive substring match. | | `include_subdomains` | No | Includes common sibling subdomains during discovery. Default: `true`. | | `ignore_query_parameters` | No | Removes query parameters when normalizing discovered URLs. Default: `true`. | For example: ```json { "url": "https://geonode.com", "include_subdomains": true, "ignore_query_parameters": true } ``` *** ## Step 3 — Review the Response A successful request returns a response similar to: ```json { "success": true, "links": [ { "url": "https://geonode.com/docs", "source": "html" }, { "url": "https://geonode.com/blog", "source": "sitemap" } ] } ``` The response contains three fields. | Field | Description | | --------- | ----------------------------------------------------- | | `success` | Indicates whether the request completed successfully. | | `links` | List of discovered URLs. | | `warning` | Optional non-fatal advisory message. | *** ## Understanding Discovered Links Each discovered URL contains two values. | Field | Description | | -------- | -------------------------------------------------------- | | `url` | The discovered page URL. | | `source` | Where the URL was discovered from (`html` or `sitemap`). | B["success"] A --> C["links"] C --> D["url"] C --> E["source"] E --> F["html"] E --> G["sitemap"] `} /> *** ## Filtering Results Use the optional `search` field to return only URLs containing specific text. For example: ```json { "url": "https://geonode.com", "search": "docs" } ``` The `search` parameter performs a case-insensitive substring match against the discovered URLs. It does not query a search engine. *** ## Working with the Results After receiving the response, you can: * Review the discovered website structure. * Select specific pages for extraction. * Pass URLs to the Crawl API. * Build your own processing workflow. B["Review"] A --> C["Extract Pages"] A --> D["Crawl Website"] `} /> *** ## Complete Workflow B["POST /v1/map"] -->C["Receive Response"] -->D["Review Links"] -->E["Use URLs"] `} /> *** ## Next Steps Now that you've created your first Map request, continue to **Configuring Map Requests** to learn how to customize URL discovery using the available request parameters. # Configuring Map Requests (/docs/scraper-api/guides/map/03_configuring_map_requests) The Map API provides several optional request parameters that let you control how URLs are discovered and returned. This guide explains when to use each option and how it affects the mapping results. *** ## Request Parameters The Map endpoint accepts the following parameters. | Field | Required | Description | | ------------------------- | -------- | --------------------------------------------------------------------------- | | `url` | Yes | The website to map. | | `search` | No | Returns only discovered URLs that contain the specified text. | | `include_subdomains` | No | Includes URLs from supported subdomains during discovery. Default: `true`. | | `ignore_query_parameters` | No | Removes query parameters when normalizing discovered URLs. Default: `true`. | *** ## Basic Request The minimum request only requires the website URL. ```json { "url": "https://geonode.com" } ``` B["Map API"] -->C["Discover URLs"] `} /> *** ## Filter Results with `search` Use the `search` parameter to return only URLs containing a specific keyword. ```json { "url": "https://geonode.com", "search": "docs" } ``` The search performs a case-insensitive substring match against the discovered URLs. For example, searching for: ```text docs ``` may return URLs such as: ```text https://docs.geonode.com https://docs.geonode.com/docs/api-reference ``` If no URLs match your search term, the request still succeeds. Example response: ```json { "success": true, "links": [], "metadata": { "url": "https://geonode.com/", "duration_ms": 4208, "links_count": 0 }, "warning": "No results found. If you targeted a sub-path, try mapping the base domain for broader coverage." } ``` An empty `links` array does not indicate an error. It simply means that no discovered URLs matched the search value. *** ## Include Subdomains By default, the Map API includes supported subdomains during discovery. You can control this behavior using `include_subdomains`. ```json { "url": "https://geonode.com", "include_subdomains": true } ``` Set the value to `false` if you only want to map the primary domain. ```json { "url": "https://geonode.com", "include_subdomains": false } ``` Use this option when: * Discovering documentation hosted on subdomains. * Mapping an organization's entire website. * Limiting discovery to the primary domain only. *** ## Ignore Query Parameters Many websites generate multiple URLs that differ only by query parameters. For example: ```text /products?page=1 /products?page=2 /products?sort=newest ``` When `ignore_query_parameters` is enabled, the Map API normalizes these URLs to reduce duplicates. ```json { "url": "https://geonode.com", "ignore_query_parameters": true } ``` Disable this option if query parameters represent unique pages that you want to keep. ```json { "url": "https://geonode.com", "ignore_query_parameters": false } ``` *** ## Combining Parameters You can combine multiple options in the same request. ```json { "url": "https://geonode.com", "search": "docs", "include_subdomains": true, "ignore_query_parameters": true } ``` B["Include Subdomains"] B --> C["Ignore Query Parameters"] C --> D["Apply Search Filter"] D --> E["Return Matching URLs"] `} /> *** ## Best Practices * Always provide the base website URL. * Use `search` to reduce the number of returned URLs. * Enable `include_subdomains` when mapping websites that host content across multiple subdomains. * Leave `ignore_query_parameters` enabled unless query parameters represent unique content. * Combine parameters to return only the URLs that are relevant to your application. *** ## Next Steps Now that you know how to configure mapping requests, continue to **Understanding Map Results** to learn how to interpret the response returned by the Map API. # Understanding Map Results (/docs/scraper-api/guides/map/04_understanding_map_results) After a successful mapping request, the Map API returns a structured response containing the discovered URLs and information about the mapping operation. Understanding this response helps you decide which URLs to extract, crawl, or process further. *** ## Response Structure A successful response contains the following top-level fields. B["success"] A --> C["links"] C --> D["url"] C --> E["source"] A --> F["metadata"] F --> G["url"] F --> H["duration_ms"] F --> I["links_count"] A --> J["warning (optional)"] `} /> *** ## Example Response ```json { "success": true, "links": [ { "url": "https://docs.geonode.com", "source": "sitemap" }, { "url": "https://docs.geonode.com/docs/api-reference", "source": "sitemap" } ], "metadata": { "url": "https://geonode.com/", "duration_ms": 0, "links_count": 202 } } ``` *** ## Response Fields | Field | Description | | ---------- | -------------------------------------------------------------------------------- | | `success` | Indicates whether the mapping request completed successfully. | | `links` | List of discovered URLs. | | `metadata` | Information about the completed mapping request. | | `warning` | Optional message returned when the request succeeds but requires your attention. | *** ## Understanding `success` The `success` field indicates whether the request completed successfully. ```json { "success": true } ``` A value of `true` means the Map API successfully processed the request and returned a response. *** ## Understanding `links` The `links` array contains every URL discovered during the mapping process. Each item contains: | Field | Description | | -------- | --------------------------- | | `url` | The discovered page URL. | | `source` | How the URL was discovered. | Example: ```json { "url": "https://docs.geonode.com/docs/api-reference", "source": "sitemap" } ``` The `source` field helps you understand where each URL originated. Possible values include: | Value | Description | | --------- | -------------------------------------------------------- | | `sitemap` | The URL was discovered from a website sitemap. | | `html` | The URL was discovered by following links in HTML pages. | The same response may contain URLs discovered from both `sitemap` and `html` sources. *** ## Understanding `metadata` The `metadata` object provides information about the mapping request itself. ```json { "metadata": { "url": "https://geonode.com/", "duration_ms": 0, "links_count": 202 } } ``` | Field | Description | | ------------- | ------------------------------------------------------------ | | `url` | The website that was mapped. | | `duration_ms` | Time taken to complete the mapping request, in milliseconds. | | `links_count` | Total number of discovered URLs returned. | This information can be useful for monitoring request performance and understanding the size of the mapping result. *** ## Understanding `warning` The `warning` field is optional. It appears when the request succeeds but there is additional information that may help you improve the results. For example: ```json { "success": true, "links": [], "metadata": { "url": "https://geonode.com/", "duration_ms": 4208, "links_count": 0 }, "warning": "No results found. If you targeted a sub-path, try mapping the base domain for broader coverage." } ``` In this example: * The request completed successfully. * No matching URLs were found. * The warning suggests mapping the base domain instead of a sub-path. A `warning` does not indicate that the request failed. Always check the `success` field before determining whether the request was successful. *** ## Empty Results Sometimes a successful request may not return any discovered URLs. Example: ```json { "success": true, "links": [], "metadata": { "links_count": 0 } } ``` This usually means: * No URLs matched the request. * The `search` filter did not match any discovered URLs. * The mapped location did not contain discoverable pages. *** ## What Can You Do with the Results? Once you've received the response, you can use the discovered URLs in several ways. B["Review URLs"] B --> C["Extraction API"] B --> D["Crawl API"] B --> E["Export Results"] B --> F["Store in Database"] `} /> Typical next steps include: * Extract structured content from selected pages. * Crawl larger sections of the website. * Export the discovered URLs for reporting or analysis. * Store the URLs for future processing. *** ## Best Practices * Always check the `success` field before processing the response. * Handle an empty `links` array as a valid response. * Review the `warning` field when it is present. * Use `links_count` to understand the size of the mapping result. * Use the `source` field to understand how each URL was discovered. *** ## Next Steps Now that you understand the Map response, continue to **Map Best Practices** to learn recommended approaches for building efficient URL discovery workflows. # Retrieving Map Jobs (/docs/scraper-api/guides/map/05_retrieving_map_jobs) After creating a mapping job, you can retrieve it later instead of creating a new request. The Map API provides two endpoints for this: * `GET /v1/map/jobs` — List your previous mapping jobs. * `GET /v1/map/{job_id}` — Retrieve the complete details of a specific mapping job. These endpoints are useful for reviewing previous jobs, recovering after an application restart, and accessing discovered URLs. *** ## Retrieval Workflow The following diagram shows how both endpoints work together. B["Receive job_id"] B-->C["GET /v1/map/jobs"] C-->D["Select a Job"] D-->E["GET /v1/map/:job_id"] E-->F["Review Discovered URLs"] `} /> *** ## When to Use Each Endpoint | Endpoint | Use Case | | ---------------------- | ------------------------------------------------- | | `GET /v1/map/jobs` | View all of your previous mapping jobs. | | `GET /v1/map/{job_id}` | Retrieve the complete details of one mapping job. | *** ## List Previous Mapping Jobs Retrieve a paginated list of your mapping jobs. ```http GET /v1/map/jobs ``` A successful response contains a list of jobs together with pagination information. Each job includes: | Field | Description | | -------------- | ---------------------------------------------------- | | `job_id` | Unique identifier of the mapping job. | | `url` | Website that was mapped. | | `status` | Current job status. | | `links_count` | Number of discovered URLs. | | `duration_ms` | Time required to complete the job. | | `search` | Search filter used when the job was created, if any. | | `error_code` | Error code if the job failed. Otherwise `null`. | | `created_at` | Time when the job started. | | `completed_at` | Time when the job finished. | The response also includes: | Field | Description | | ------------ | --------------------------------- | | `page` | Current page number. | | `page_size` | Number of jobs returned per page. | | `page_count` | Total number of available pages. | *** ## Finding the Right Job Most applications identify a job by checking: * The website URL. * The job status. * The creation time. * The optional search filter. B["Review URL"] -->C["Check Status"] -->D["Select job_id"] -->E["Retrieve Details"] `} /> *** ## Retrieve a Specific Mapping Job Once you have a `job_id`, retrieve the complete job information. ```http GET /v1/map/{job_id} ``` This endpoint returns: * Job information. * Mapping configuration. * Statistics. * Every discovered URL. *** ## Understanding the Response A completed mapping job contains several sections. B["Job Information"] A-->C["Configuration"] A-->D["Statistics"] A-->E["Discovered Links"] `} /> ### Job Information These fields describe the mapping job itself. | Field | Description | | -------------- | ---------------------------------- | | `job_id` | Unique job identifier. | | `url` | Website that was mapped. | | `status` | Current status of the mapping job. | | `created_at` | Time when the job was created. | | `completed_at` | Time when the job completed. | *** ### Mapping Configuration These values show how the mapping job was executed. | Field | Description | | ------------------------- | ---------------------------------------------------------------------- | | `search` | Search filter applied during URL discovery. | | `include_subdomains` | Indicates whether subdomains were included. | | `ignore_query_parameters` | Indicates whether query parameters were ignored when discovering URLs. | *** ### Statistics The response also includes useful information about the completed job. | Field | Description | | ---------------- | --------------------------------------------- | | `links_count` | Total number of discovered URLs. | | `duration_ms` | Time taken to complete the mapping job. | | `tokens_charged` | Number of API tokens charged for the request. | *** ### Discovered Links The `links` array contains every URL discovered during the mapping process. Each entry contains: | Field | Description | | -------- | --------------------------------------------------- | | `url` | The discovered page URL. | | `source` | Where the URL was discovered (`html` or `sitemap`). | Example: ```json { "url": "https://docs.geonode.com", "source": "sitemap" } ``` *** ## Recovering Previous Jobs If your application loses the original `job_id`, you don't need to create another mapping request. Instead: 1. Call `GET /v1/map/jobs`. 2. Find the required job. 3. Copy its `job_id`. 4. Retrieve the complete results with `GET /v1/map/{job_id}`. B["GET /v1/map/jobs"] -->C["Locate job_id"] -->D["GET /v1/map/:job_id"] -->E["Continue Processing"] `} /> This approach helps avoid creating duplicate mapping jobs. *** ## Best Practices * Save the returned `job_id` whenever you create a mapping job. * Use `GET /v1/map/jobs` to locate previous jobs if the identifier is unavailable. * Check the job `status` before using its results. * Review `links_count` to understand how many URLs were discovered. * Reuse completed mapping jobs whenever possible instead of creating duplicate requests. *** ## Next Steps Now that you know how to retrieve previous mapping jobs and inspect their results, continue to **Best Practices** to learn recommendations for building efficient and reliable mapping workflows. # Map Best Practices (/docs/scraper-api/guides/map/06_map_best_practices) The Map API is often the first step in a scraping workflow. Following a few best practices can help you discover relevant URLs, reduce unnecessary processing, and build more efficient applications. *** ## Start with the Base Domain Whenever possible, begin mapping from the root of the website. For example: ```text https://example.com ``` instead of: ```text https://example.com/docs/getting-started ``` Starting from the base domain gives the Map API a broader view of the website and increases the chances of discovering all relevant pages. B["More Discovered URLs"] C["Sub-page"] -->D["Limited Discovery"] `} /> *** ## Use `search` to Reduce Results If you're only interested in a specific section of a website, use the `search` parameter. For example: ```json { "url": "https://geonode.com", "search": "docs" } ``` This reduces the number of returned URLs and makes it easier to work with the results. Typical search values include: * `docs` * `blog` * `api` * `support` The `search` parameter performs a case-insensitive substring match against discovered URLs. *** ## Configure Subdomain Discovery Appropriately The `include_subdomains` option controls whether supported subdomains are included during URL discovery. Enable it when your content is distributed across multiple subdomains, such as: ```text docs.example.com blog.example.com support.example.com ``` Disable it if you only want to discover pages from the primary domain. *** ## Decide How to Handle Query Parameters Many websites generate multiple URLs that differ only by query parameters. For example: ```text /products?page=1 /products?page=2 /products?sort=newest ``` When `ignore_query_parameters` is enabled, these URLs are normalized during discovery. Disable this option only if the query parameters represent unique content that should be treated as separate pages. *** ## Review Results Before Processing The Map API discovers URLs—it does not extract page content. Review the discovered URLs before deciding what to process next. B["Review URLs"] B-->C["Extract Selected Pages"] B-->D["Crawl Website"] B-->E["Export URL List"] `} /> This approach helps reduce unnecessary extraction and crawling. *** ## Reuse Existing Mapping Jobs If you have already mapped a website, retrieve the existing job instead of creating another mapping request. B["GET /v1/map/jobs"] -->C["Select job_id"] -->D["GET /v1/map/:job_id"] -->E["Continue Processing"] `} /> Reusing existing jobs helps avoid duplicate requests and allows your application to continue working with previously discovered URLs. *** ## Handle Empty Results Gracefully A successful request may return an empty `links` array. For example: ```json { "success": true, "links": [], "warning": "No results found. If you targeted a sub-path, try mapping the base domain for broader coverage." } ``` This does not indicate that the request failed. Instead, consider: * Mapping the base domain instead of a sub-path. * Using a less restrictive `search` value. * Verifying that the target website contains discoverable pages. *** ## Build a Complete Mapping Workflow The Map API works best as the first stage of a larger scraping pipeline. B["Map API"] -->C["Review URLs"] C-->D["Extraction API"] C-->E["Crawl API"] D-->F["Store Results"] E-->F `} /> This workflow allows you to discover URLs first and then process only the pages that are relevant to your application. *** ## Best Practices Checklist Before deploying your application, make sure you: * Start mapping from the base domain whenever possible. * Use `search` to narrow large result sets. * Enable subdomain discovery only when needed. * Configure query parameter handling based on your use case. * Review discovered URLs before extraction or crawling. * Reuse existing mapping jobs whenever possible. * Handle empty results as a valid response. * Build your workflow around URL discovery before content extraction. *** ## Next Steps You now understand how to use the Map API effectively. Continue to the **Map API Reference** for detailed endpoint documentation, request schemas, response fields, and additional examples. # Overview (/docs/scraper-api/guides/mcp/00_overview) AI assistants are changing how developers build software. Instead of switching between documentation, browsers, terminals, and APIs, developers can now ask their AI assistant to perform tasks such as writing code, debugging applications, analyzing data, and interacting with external services. However, AI assistants have one important limitation—they cannot access live websites, external APIs, or other tools on their own. If you are an AI coding agent onboarding to Geonode, start with the [Agent Onboarding skill](/agent-onboarding/SKILL.md). ## What is MCP? The **Model Context Protocol (MCP)** is an open standard that enables AI assistants to securely connect to external tools, APIs, and data sources. Instead of relying only on the information available in a conversation, an AI assistant can use MCP to retrieve live data, execute actions, and return results in real time. This allows AI assistants to become active participants in your workflow rather than simple conversational assistants. ## Why use MCP instead of manually calling APIs? Without MCP, interacting with external services often requires switching between multiple tools. You typically need to: * Find the documentation. * Write an API request. * Send the request. * Review the response. * Copy the results back into your AI assistant. With MCP, your AI assistant performs these tasks for you. You simply describe what you need, and the assistant communicates with the appropriate service behind the scenes. For example, instead of manually calling an API, you can simply ask: * Extract the content from this product page. * Crawl the official React documentation. * Process these 500 URLs. * Check whether my extraction job has finished. The AI assistant automatically selects the appropriate MCP tool to complete your request. ## What is Geonode MCP? Geonode MCP connects your AI assistant directly to Geonode's web scraping platform. Once connected, your assistant can retrieve live web content, extract structured data, crawl entire websites, process batch extraction jobs, and monitor long-running tasks without leaving your editor or chat. Instead of managing API requests yourself, you simply describe the task, and your assistant uses Geonode's scraping capabilities on your behalf. ## What can you do with Geonode MCP? With Geonode MCP, your AI assistant can: * Extract content from individual web pages. * Crawl websites and documentation portals. * Process hundreds or thousands of URLs in batch jobs. * Retrieve structured content in Markdown or HTML. * Render JavaScript-powered websites. * Route requests through residential proxies with geo-targeting. * Include custom request headers when required. * Monitor the progress of long-running extraction and crawl jobs. ## How Geonode MCP Works ```text You │ ▼ AI Assistant (Cursor, Claude, Windsurf, etc.) │ ▼ Geonode MCP Server │ ▼ Geonode Scraper API │ ▼ Target Website │ ▼ Structured Results │ ▼ AI Assistant Response ``` ## Try without an account You can point an MCP client at `https://scraper.geonode.io/mcp` with **no API key** and use the `extract` tool under free anonymous limits. See [Use Without an API Key](/docs/scraper-api/guides/mcp/01_use-without-an-api-key) for the endpoint, client config, limits, and how to upgrade. ## Supported Clients Geonode MCP supports the following AI assistants and development environments: * [Claude Desktop](/docs/scraper-api/guides/mcp/03_claude-desktop) * [Claude Code](/docs/scraper-api/guides/mcp/02_claude-code) * [Cursor](/docs/scraper-api/guides/mcp/04_cursor) * Windsurf * Smithery * [Docker](/docs/scraper-api/guides/mcp/05_docker-mcp) * [Visual Studio Code](/docs/scraper-api/guides/mcp/08_visual-studio-code) * [Codex](/docs/scraper-api/guides/mcp/09_codex) Each client has its own setup guide with step-by-step installation instructions. ## Next Steps To try Geonode MCP with no signup, start with [Use Without an API Key](/docs/scraper-api/guides/mcp/01_use-without-an-api-key). For full access (all tools, async jobs, JavaScript rendering, and plan limits), continue to [Before You Start](/docs/scraper-api/guides/mcp/01_before-you-start), then follow the setup guide for your preferred client. # Before You Start (/docs/scraper-api/guides/mcp/01_before-you-start) Before configuring Geonode MCP with full access, make sure you have everything you need. This guide covers the API key, authentication, endpoint, and the tools your AI assistant can access once connected. You can use Geonode MCP without an API key for small jobs. See [Use Without an API Key](/docs/scraper-api/guides/mcp/01_use-without-an-api-key). This page covers authenticated setup for the full tool set. ## Prerequisites Before continuing, you'll need: * A Geonode account. * A valid Geonode API key. * An MCP-compatible client such as Claude Desktop, Claude Code, Cursor, Windsurf, or Docker. ## Get Your API Key Authenticated requests through the Geonode MCP Server use an API key. To create or view your API key: 1. Sign in to your Geonode dashboard. 2. Open **API Keys**. 3. Copy an existing key or create a new one. 4. Store your API key securely. Never share your API key or commit it to a public repository. Anyone with your API key can make requests on your behalf. ## MCP Endpoint Configure your MCP client to connect to the following endpoint: ```text https://scraper.geonode.io/mcp ``` ## Authentication Geonode MCP authenticates requests using your API key. Most MCP clients send the key through the `X-Api-Key` HTTP header. Some clients cannot configure custom headers. In those cases, you can provide the key as the `api_key` argument when calling a tool. Using the `X-Api-Key` header is the recommended approach, and all setup guides in this documentation use this method. Replace `YOUR_API_KEY` with your own Geonode API key in every configuration example throughout this documentation. ## Available MCP Tools Once connected, your AI assistant can automatically use the following Geonode tools. ### Content Extraction | Tool | Description | | --------- | -------------------------------------------------------------------------------------------------------------------------------------------------- | | `extract` | Extract content from a single web page in Markdown or HTML. Supports JavaScript rendering, residential proxies, geo-targeting, and custom headers. | | `job` | Retrieve the result of a previously submitted asynchronous extraction job. | | `jobs` | List your extraction jobs and filter them by status, URL, date, or output format. | ### Batch Processing | Tool | Description | | -------------- | ------------------------------------------------------------------- | | `batch` | Extract up to 1,000 URLs in a single asynchronous batch job. | | `batch_status` | Monitor the progress of a batch job and retrieve paginated results. | | `cancel_batch` | Cancel a running batch job. | ### Website Crawling | Tool | Description | | -------------- | ---------------------------------------------------------------------------------------------------------------------------- | | `crawl` | Crawl an entire website starting from a seed URL. Supports breadth-first crawling, configurable depth, and domain filtering. | | `crawl_status` | Check the progress of a crawl job and retrieve discovered pages. | | `cancel_crawl` | Cancel a running crawl job. | ### Usage Statistics | Tool | Description | | ------------ | ---------------------------------------------------------------------------------- | | `statistics` | Retrieve usage statistics such as extraction count, success rate, and token usage. | ## How Your AI Uses These Tools You don't need to manually choose which tool to use. Instead, simply describe the task, and your AI assistant selects the appropriate Geonode tool automatically. For example: | You ask | AI uses | | ----------------------------------------------- | ----------------------- | | "Extract the content from this page." | `extract` | | "Process these 500 URLs." | `batch` | | "Crawl this documentation website." | `crawl` | | "Check whether my extraction job has finished." | `job` or `batch_status` | | "Show my extraction statistics." | `statistics` | ## Next Steps You're now ready to configure Geonode MCP with an API key. To try first without a key, see [Use Without an API Key](/docs/scraper-api/guides/mcp/01_use-without-an-api-key). Choose the setup guide for your preferred client: * [Claude Desktop](/docs/scraper-api/guides/mcp/03_claude-desktop) * [Claude Code](/docs/scraper-api/guides/mcp/02_claude-code) * [Cursor](/docs/scraper-api/guides/mcp/04_cursor) * Windsurf * Smithery * [Docker MCP](/docs/scraper-api/guides/mcp/05_docker-mcp) # Use Without an API Key (/docs/scraper-api/guides/mcp/01_use-without-an-api-key) You can connect Geonode MCP to your AI assistant and start extracting web pages without creating an account or adding an API key. Point your MCP client at the Geonode endpoint, and your assistant can use the `extract` tool right away. This path is ideal for trying the service and for small jobs. When you need more capacity or more tools, add an API key to the same configuration. * How to connect without an API key * What works in anonymous mode * Free rate and volume limits * How to upgrade to authenticated access *** ## MCP Endpoint Configure your MCP client with this endpoint: ```text https://scraper.geonode.io/mcp ``` The server uses MCP over streamable HTTP. To confirm the endpoint is reachable, run: ```bash curl https://scraper.geonode.io/mcp/health ``` A healthy response looks like this: ```json {"status":"ok","transport":"streamable-http","endpoint":"/mcp"} ``` Always use `https://scraper.geonode.io/mcp`. Other hostnames, including preproduction ones, are not supported for public use and will fail. *** ## Connect Your Client You do not need a key, token, or custom header for anonymous access. If your client asks for authentication, leave it empty. ### Claude Code ```bash claude mcp add --transport http geonode https://scraper.geonode.io/mcp ``` ### Cursor, Claude Desktop, and other JSON-based clients ```json { "mcpServers": { "geonode": { "url": "https://scraper.geonode.io/mcp" } } } ``` After you save the configuration, restart your client if required, then open a new chat and ask your assistant to extract a page. *** ## What You Can Do Without a Key Without an API key, Geonode MCP exposes one tool: `extract`. It fetches a URL and returns the page content. You can ask your assistant something like: ```text Extract the content from https://example.com ``` Behind the scenes, the assistant calls `extract` on that URL. Anonymous access is intentionally limited compared with a keyed connection: | Capability | Without a key | With an API key | | -------------------- | ---------------------- | ---------------------------------------------------------------- | | Tools | `extract` only | `extract`, `map`, `search`, `crawl`, `batch`, and job management | | Mode | Synchronous only | Synchronous and asynchronous jobs | | JavaScript rendering | Not available | Available with `render_js` | | Custom headers | `Accept-Language` only | Any headers you need | | Country selection | Not available | Available | If you request a feature that anonymous access does not support, the server returns a clear message about which limit applied, instead of failing silently. *** ## Free Limits Anonymous use is rate limited by network address: | Limit | Value | | --------------- | ------------------------- | | Requests | 5 per minute | | Concurrency | 1 request at a time | | Download volume | 100 MB of content per day | If several people share the same office or provider network, they may share a larger group allowance for that network prefix. Limits then apply to the group rather than to each person alone. If requests keep getting refused, a short cooling-off period starts and can grow if the refusals continue. Wait for it to clear; the limits reset on their own. Anonymous traffic is metered but not billed. No payment method is required. Requests that are refused are not billed, with or without an API key. *** ## Upgrade to an API Key When you need the full tool set, async jobs, JavaScript rendering, country targeting, or higher limits, create a Geonode account, get an API key, and add it to the same MCP configuration: ```json { "mcpServers": { "geonode": { "url": "https://scraper.geonode.io/mcp", "headers": { "X-Api-Key": "YOUR_API_KEY" } } } } ``` Replace `YOUR_API_KEY` with your Geonode API key. As soon as the key is recognized, the available tools expand and your plan rate limits apply. * Plans and pricing: [geonode.com/pricing](https://geonode.com/pricing) * Full authenticated setup: [Before You Start](/docs/scraper-api/guides/mcp/01_before-you-start) Then follow the setup guide for your client, such as [Claude Code](/docs/scraper-api/guides/mcp/02_claude-code), [Claude Desktop](/docs/scraper-api/guides/mcp/03_claude-desktop), or [Cursor](/docs/scraper-api/guides/mcp/04_cursor). *** ## Next Steps 1. Connect your client to `https://scraper.geonode.io/mcp` with no key. 2. Ask your assistant to extract a public page. 3. When you outgrow the free limits, add an API key and continue with [Before You Start](/docs/scraper-api/guides/mcp/01_before-you-start). # Claude Code (/docs/scraper-api/guides/mcp/02_claude-code) Claude Code has built-in support for the Model Context Protocol (MCP), allowing it to securely connect to external tools such as the Geonode Scraper API. Once configured, Claude can automatically extract webpages, crawl websites, process multiple URLs, and retrieve structured content using natural language prompts. Make sure you've completed the **Before You Start** guide and have: * A Geonode API key * The Geonode MCP endpoint * Claude Code installed * Claude Code authenticated with your Anthropic account *** ## Step 1 — Open your terminal Open your terminal and verify that Claude Code is installed. Launch Claude Code by running: ```bash claude ``` If Claude Code starts successfully, you're ready to configure the Geonode MCP server. Claude Code Home *** ## Step 2 — Register the Geonode MCP Server Run the following command to register the Geonode MCP server with Claude Code. ```bash claude mcp add \ --transport http \ geonode-scraper \ https://scraper.geonode.io/mcp \ --header "X-Api-Key:YOUR_API_KEY" ``` Replace `YOUR_API_KEY` with your Geonode API key. Claude Code saves this configuration locally and automatically registers the MCP server. Register the MCP Server *** ## Step 3 — Verify the Connection To verify that the server has been registered successfully, run: ```bash claude mcp list ``` You should see output similar to: ```text geonode-scraper: https://scraper.geonode.io/mcp (HTTP) - Connected ``` If the status is **Connected**, Claude Code can communicate with the Geonode MCP server. Verify the Connection *** ## Step 4 — Start Claude Code Launch Claude Code. ```bash claude ``` Claude automatically loads all configured MCP servers when it starts. Launch Claude Code *** ## Step 5 — Use Geonode with Natural Language You don't need to manually invoke MCP tools. Simply describe the task, and Claude automatically chooses the appropriate Geonode MCP tool. ### Extract a web page ```text Extract the content from https://geonode.com ``` ### Summarize a page ```text Read https://geonode.com and summarize the homepage. ``` ### Crawl a documentation website ```text Crawl https://docs.geonode.com and tell me what documentation sections exist. ``` ### Process multiple pages ```text Extract these URLs and summarize each page. https://example.com/page1 https://example.com/page2 https://example.com/page3 ``` ### Check extraction statistics ```text Show my Geonode extraction statistics. ``` Claude automatically calls the appropriate MCP tool, waits for the response, and presents the results directly in the conversation. Using Geonode MCP in Claude Code *** ## How Claude Code Uses MCP When you submit a prompt, Claude analyzes your request and automatically selects the correct Geonode MCP tool. | Your Prompt | MCP Tool | | -------------------------------- | ------------ | | "Extract this page." | `extract` | | "Crawl this documentation site." | `crawl` | | "Process these URLs." | `batch` | | "Show my extraction statistics." | `statistics` | There is no need to manually select or invoke MCP tools. Claude Code handles tool selection, parameter mapping, execution, and result processing automatically. *** ## FAQs Verify that: * The server URL is `https://scraper.geonode.io/mcp`. * Your `X-Api-Key` header contains a valid Geonode API key. * Your internet connection is active. Run the following command to verify the connection: ```bash claude mcp list ``` If necessary, remove and add the server again. Authentication failures usually occur when: * The API key is incorrect. * The `X-Api-Key` header is missing. * The API key has expired or been revoked. Generate a new API key and register the server again. If Claude cannot reach the MCP server: * Verify the server URL. * Check your internet connection. * Confirm that the Geonode service is available. Then run: ```bash claude mcp list ``` to confirm the server status. Run: ```bash claude mcp list ``` Claude displays every configured MCP server along with its current connection status. Run: ```bash claude mcp remove geonode-scraper ``` This removes the Geonode MCP configuration from Claude Code. # Claude Desktop (/docs/scraper-api/guides/mcp/03_claude-desktop) import { Accordion, Accordions } from 'fumadocs-ui/components/accordion'; Claude Desktop supports the Model Context Protocol (MCP), so you can use Geonode scraping tools directly in chat. Geonode authenticates with a static `X-Api-Key` header. In Claude Desktop, the working setup today is **Settings → Developer → Edit Config** with `mcp-remote`. **Add custom connector** only accepts a server URL and optional OAuth. It has no field for static headers, so it cannot authenticate to Geonode today. Use **Developer → Edit Config** instead. Custom connector support will be improved later. Make sure you've completed the **Before You Start** guide and have: * A Geonode API key * The Geonode MCP endpoint: `https://scraper.geonode.io/mcp` * Claude Desktop installed * **Node.js 18+** installed (required by `mcp-remote`) *** ## Step 1 — Open Settings Open Claude Desktop. Click your profile menu in the bottom-left corner, then select **Settings**. Open Claude Desktop Settings *** ## Step 2 — Open Developer → Edit Config In Settings: 1. Under **Desktop app**, open **Developer**. 2. Find **Local MCP servers**. 3. Click **Edit Config**. Open Developer Edit Config This opens `claude_desktop_config.json`. You can also edit the file directly: | OS | Path | | ----------- | ----------------------------------------------------------------- | | **macOS** | `~/Library/Application Support/Claude/claude_desktop_config.json` | | **Windows** | `%APPDATA%\Claude\claude_desktop_config.json` | Claude Desktop config file *** ## Step 3 — Add the Geonode MCP server Add the Geonode server under the top-level `mcpServers` key. If the file already has other settings, keep them and place `mcpServers` **alongside** those top-level keys — do not nest it inside another object. ```json { "mcpServers": { "geonode-scraper": { "command": "npx", "args": [ "mcp-remote", "https://scraper.geonode.io/mcp", "--header", "X-Api-Key:${GEONODE_API_KEY}" ], "env": { "GEONODE_API_KEY": "YOUR_API_KEY" } } } } ``` Replace `YOUR_API_KEY` with your Geonode API key. Geonode MCP config in claude_desktop_config.json Write `X-Api-Key:${GEONODE_API_KEY}` with **no space** after the colon. Claude Desktop splits `args` on spaces. A space after the colon drops the key value and authentication fails with `401`. `mcpServers` must sit at the root of `claude_desktop_config.json`, next to any existing keys such as preferences. Do not nest `mcpServers` inside another object. *** ## Step 4 — Fully restart Claude Desktop 1. Save `claude_desktop_config.json`. 2. **Quit Claude Desktop completely** — close the process, not only the window. 3. Open Claude Desktop again. 4. Start a **new chat**. A partial window close is not enough. MCP servers load on a full app restart. *** ## Step 5 — Verify the connection In a new chat, ask Claude to call a Geonode tool. For example: ```text Call the geonode scraper statistics endpoint and tell me the HTTP status and a short summary of the response. ``` If the connection works: * Claude uses the Geonode MCP tools * You do **not** need to pass `api_key` in the chat prompt * A successful statistics call returns **HTTP 200** Verified Geonode MCP statistics call *** ## Step 6 — Start using Geonode MCP Once connected, describe tasks in natural language. Claude selects the Geonode tool automatically. ### Extract a web page ```text Extract the content from https://geonode.com ``` ### Summarize a page ```text Read https://geonode.com and summarize the homepage. ``` ### Crawl a documentation website ```text Crawl https://docs.geonode.com and tell me what documentation sections exist. ``` ### Check extraction statistics ```text Show my Geonode extraction statistics. ``` *** ## Important setup notes | Requirement | Why it matters | | --------------------------------------------- | ------------------------------------------------------------- | | Use **Edit Config**, not Add custom connector | Custom connector cannot send static `X-Api-Key` headers today | | Keep `mcpServers` at the **top level** | Nested config is ignored | | No space in `X-Api-Key:${GEONODE_API_KEY}` | Spaces break header parsing and cause `401` | | **Node.js 18+** | Required by `npx mcp-remote` | | Full app restart + new chat | MCP servers load only after a complete restart | *** ## FAQs Not for Geonode right now. **Add custom connector** only supports a server URL and optional OAuth. It has no static header field, so it cannot send `X-Api-Key`. Use **Developer → Edit Config** with `mcp-remote` instead. Check these first: * `X-Api-Key:${GEONODE_API_KEY}` has **no space** after the colon * `GEONODE_API_KEY` in `env` is set to a valid Geonode API key * You fully quit and restarted Claude Desktop * You are testing in a **new chat** Verify that: * `mcpServers` is a top-level key in `claude_desktop_config.json` * Node.js 18+ is installed (`node -v`) * The JSON file is valid * Claude Desktop was fully quit and reopened * You opened a new chat after restarting No. After Edit Config is set up correctly, the API key is sent through the `X-Api-Key` header by `mcp-remote`. You do not need to include the key in prompts. | OS | Path | | ----------- | ----------------------------------------------------------------- | | **macOS** | `~/Library/Application Support/Claude/claude_desktop_config.json` | | **Windows** | `%APPDATA%\Claude\claude_desktop_config.json` | You can also open it from **Settings → Developer → Edit Config**. # Cursor (/docs/scraper-api/guides/mcp/04_cursor) import { Accordion, Accordions } from 'fumadocs-ui/components/accordion'; Cursor supports remote MCP servers over HTTP, allowing your AI assistant to securely connect to Geonode's scraping platform. Once connected, Cursor can automatically extract web pages, crawl websites, process batch jobs, and retrieve structured content based on your prompts. Make sure you've completed the **Before You Start** guide and have: * A Geonode API key * The Geonode MCP endpoint * Cursor installed ## Step 1 — Open MCP Settings Open Cursor and navigate to: Settings → Tools & MCPs Under Home MCP Servers, click ew MCP Server Open MCP Settings *** ## Step 2 — Create a New MCP Server Configure the server with the following information: | Field | Value | | -------------- | -------------------------------- | | **Name** | `geonode-scraper` | | **Type** | `URL` | | **Server URL** | `https://scraper.geonode.io/mcp` | Under **Headers**, add: | Key | Value | | ----------- | -------------- | | `X-Api-Key` | `YOUR_API_KEY` | Replace `YOUR_API_KEY` with your Geonode API key. *** ## Step 3 — Save the Configuration Click **Add MCP**. Cursor creates the server configuration automatically. If you prefer editing the configuration manually, the generated `mcp.json` looks like this: ```json { "mcpServers": { "geonode-scraper": { "url": "https://scraper.geonode.io/mcp", "headers": { "X-Api-Key": "YOUR_API_KEY" } } } } ``` Add Custom MCP *** ## Step 4 — Verify the Connection Return to **Settings → Tools & MCPs**. If the connection is successful, you'll see: * A green status indicator. * The **geonode-scraper** server. * All available Geonode tools. The available tools include: * `extract` * `job` * `jobs` * `statistics` * `batch` * `batch_status` * `cancel_batch` * `crawl` * `crawl_status` * `cancel_crawl` Verify Custom MCP *** ## Step 5 — Start Using Geonode MCP Open a new **Agent** or **Composer** chat. You don't need to manually select a tool. Simply describe the task in natural language, and Cursor automatically chooses the appropriate Geonode MCP tool. For example: ### Extract a web page ```text Extract the content from https://geonode.com ``` ### Summarize a page ```text Read https://geonode.com and summarize the homepage. ``` ### Crawl a documentation website ```text Crawl https://docs.geonode.com and tell me what documentation sections exist. ``` ### Process multiple pages ```text Extract these URLs and summarize each page. https://example.com/page1 https://example.com/page2 https://example.com/page3 ``` ### Check extraction statistics ```text Show my Geonode extraction statistics. ``` *** ## How it works When you send a prompt, Cursor analyzes your request and automatically calls the appropriate Geonode MCP tool. For example: | Your Prompt | MCP Tool | | -------------------------------- | ------------ | | "Extract this page." | `extract` | | "Crawl this documentation site." | `crawl` | | "Process these URLs." | `batch` | | "Show my extraction statistics." | `statistics` | You never need to call these tools manually—Cursor handles tool selection and parameter mapping for you. *** ## Viewing Tool Calls During execution, Cursor displays each MCP tool invocation in the chat. You can expand each tool call to inspect: * The tool that was executed. * The request parameters. * The response returned by the Geonode MCP Server. This can be useful for debugging, understanding how Cursor interprets your prompts, or learning which MCP tool was used for a particular request. Ran Extract in geonode-scraper *** ## FAQs Verify the following: * The server URL is correct. * Your Geonode API key is valid. * The MCP server configuration has been saved. If the server still doesn't appear, restart Cursor and try again. Ensure your `X-Api-Key` header contains a valid Geonode API key. If you've recently generated a new key, update your MCP configuration and reconnect the server. Open **Settings → Tools & MCPs** and verify: * The server has a green status indicator. * The connection is active. * The list of Geonode tools is displayed. If the tools are missing, refresh the MCP configuration or reconnect the server. # Docker MCP (/docs/scraper-api/guides/mcp/05_docker-mcp) Docker provides two ways to use the Geonode MCP Server. * **Docker Desktop** provides a graphical interface for managing MCP servers and AI clients. * **Docker CLI** runs a lightweight MCP bridge directly from the terminal. Choose the option that best fits your workflow. Make sure you've completed the **Before You Start** guide and have: * A Geonode API key * The Geonode MCP endpoint * Docker installed *** # Option 1 — Docker Desktop Docker Desktop includes the **MCP Toolkit**, which allows you to manage MCP servers and connect them to supported AI clients. ### Step 1 — Install Docker Desktop Download and install Docker Desktop for your operating system. Launch Docker Desktop and ensure the Docker Engine is running. ### Step 2 — Open the MCP Toolkit From the Docker Desktop sidebar, open: **Models → MCP Toolkit** This is where Docker manages MCP servers and connected AI clients. Docker Desktop MCP Toolkit ### Step 3 — Connect Your AI Client Open the **Clients** tab. Choose the AI client you want to use. For example: * Cursor * Continue.dev * Codex * Gemini CLI * Goose Click **Connect** next to your client. Docker configures the client to communicate with the MCP Toolkit. Connect Your AI Client ### Step 4 — Restart Your AI Client After connecting the client, restart it. Once restarted, Docker automatically exposes the configured MCP servers to the client. ### Step 5 — Verify the Connection Open your AI client's MCP settings. You should see: * a green status indicator * the configured MCP server * the available Geonode tools Verify the Connection *** # Option 2 — Docker CLI If you prefer using the terminal, Docker can run a lightweight MCP bridge without any additional installation. ### Step 1 — Verify Docker Open a terminal. Run: ```bash docker --version ``` If Docker is installed correctly, the installed version is displayed. ### Step 2 — Start the MCP Bridge Run: ```bash docker run -i --rm \ -e GEONODE_API_KEY=YOUR_API_KEY \ node:20-alpine \ npx -y mcp-remote https://scraper.geonode.io/mcp \ --header "X-Api-Key:${GEONODE_API_KEY}" ``` Replace `YOUR_API_KEY` with your Geonode API key. This command: * starts a temporary Docker container * downloads `mcp-remote` * connects to the Geonode MCP Server * creates a local MCP bridge over STDIO ### Step 3 — Verify the Bridge When the bridge starts successfully, you'll see output similar to: ```text Using transport strategy: http-first Connected to remote server using StreamableHTTPClientTransport Local STDIO server running Proxy established successfully Press Ctrl+C to exit ``` The bridge is now ready to receive requests. Verify the Connection ### Step 4 — Configure Your AI Client Configure your AI client to use the local Docker bridge. Once connected, the client automatically gains access to the available Geonode MCP tools. *** # Start Using Geonode MCP You can now ask your AI assistant to perform tasks such as: #### Extract a web page ```text Extract the content from https://geonode.com ``` #### Summarize a page ```text Summarize https://geonode.com ``` #### Crawl a website ```text Crawl https://docs.geonode.com ``` #### Process multiple pages ```text Extract and summarize these URLs: https://example.com/page1 https://example.com/page2 https://example.com/page3 ``` #### View extraction statistics ```text Show my Geonode extraction statistics. ``` *** # Stopping the Docker Bridge To stop the Docker bridge, press: ```text Ctrl + C ``` Since the container is started with the `--rm` flag, Docker automatically removes it after it exits. *** # FAQ Ensure Docker Desktop is installed and the Docker Engine is running before starting the bridge. Verify that your Geonode API key is correct. Confirm that Docker has internet access and that the Geonode MCP endpoint is reachable. Restart your AI client after configuring the MCP bridge, then verify the MCP server is enabled in the client's settings. Press **Ctrl +C** in the terminal. Docker automatically removes the temporary container because the command uses the `--rm` option. # Devin AI (/docs/scraper-api/guides/mcp/07_devin-ai) Devin.ai supports the Model Context Protocol (MCP), allowing AI agents to securely connect to external tools like the Geonode Scraper API. Once connected, Cascade can automatically extract web pages, crawl websites, process multiple URLs, and retrieve structured content based on your prompts. Windsurf has been acquired by Devin AI. Recent versions use the Devin branding throughout the application, while older releases may still display Windsurf. The MCP configuration process is the same for both. Make sure you've completed the **Before You Start** guide and have: * A Geonode API key * The Geonode MCP endpoint * A Devin.ai installation with MCP support. *** ## Step 1 — Open the MCP Marketplace Open **Devin.ai**. Navigate to: **Settings → MCP Marketplace** From the MCP Marketplace, click **Add custom MCP**. Open the MCP Marketplace *** ## Step 2 — Add the Geonode MCP Server Configure the MCP server using the following information. | Field | Value | | -------------- | -------------------------------- | | **Name** | `geonode-scraper` | | **Server URL** | `https://scraper.geonode.io/mcp` | Under **Headers**, add the following authentication header. | Header | Value | | ----------- | -------------- | | `X-Api-Key` | `YOUR_API_KEY` | Replace `YOUR_API_KEY` with your Geonode API key. Add the Geonode MCP Server *** ## Step 3 — Connect the Server Click **Connect**. If the configuration is correct, Devin.ai establishes a connection to the Geonode MCP server. The status changes from **Not Connected** to **Connected**. *** ## Step 4 — Verify the Configuration After connecting, open the MCP server details to verify that: * The server is connected. * Your API key has been saved. * The Geonode MCP server appears in the list of installed MCP servers. If the server is shown as **Connected**, the setup is complete. Connected Server Details *** ## Step 5 — Start Using Geonode MCP Open a new **Cascade** conversation. You don't need to manually choose an MCP tool. Simply describe the task in natural language, and Devin.ai automatically selects the appropriate Geonode MCP tool. For example: ### Extract a web page ```text Extract the content from https://geonode.com ``` ### Summarize a page ```text Read https://geonode.com and summarize the homepage. ``` ### Crawl a documentation website ```text Crawl https://docs.geonode.com and tell me what documentation sections exist. ``` ### Process multiple pages ```text Extract these URLs and summarize each page. https://example.com/page1 https://example.com/page2 https://example.com/page3 ``` ### Check extraction statistics ```text Show my Geonode extraction statistics. ``` *** ## How Devin.ai Uses MCP When you send a prompt, Devin.ai analyzes your request and automatically calls the appropriate Geonode MCP tool. For example: | Your Prompt | MCP Tool | | ------------------------------ | ------------ | | Extract this page. | `extract` | | Crawl this documentation site. | `crawl` | | Process these URLs. | `batch` | | Show my extraction statistics. | `statistics` | You never need to manually invoke these tools—Cascade handles tool selection and parameter mapping automatically. *** ## FAQ's Verify that: * The server URL is `https://scraper.geonode.io/mcp`. * Your `X-Api-Key` header contains a valid Geonode API key. * Your internet connection is active. After making changes, try connecting again. Authentication failures usually occur when: * The API key is incorrect. * The `X-Api-Key` header is missing. * The API key has expired or been revoked. Generate a new API key if necessary and reconnect the server. If the MCP server cannot be reached: * Verify that the server URL is correct. * Check your network connection. * Confirm that the Geonode service is available. Then reconnect the MCP server. Ensure that: * The MCP server status is **Connected**. * The server appears in the installed MCP servers list. * Devin.ai has finished loading the MCP configuration. If necessary, reconnect the MCP server or restart Devin.ai. # Visual Studio Code (/docs/scraper-api/guides/mcp/08_visual-studio-code) Visual Studio Code supports the **Model Context Protocol (MCP)**, allowing **GitHub Copilot Agent Mode** to securely connect to external tools such as the Geonode Scraper API. Once configured, GitHub Copilot can automatically extract web pages, crawl websites, process multiple URLs, and retrieve structured content directly from your editor. Make sure you've completed the **Before You Start** guide and have: * A Geonode API key. * The Geonode MCP endpoint. * Visual Studio Code installed. * GitHub Copilot installed and signed in. * Agent Mode enabled. *** ## Step 1 — Open the MCP Servers Panel Open **Visual Studio Code** and open the **GitHub Copilot Chat** panel. Click the **Settings (⚙️)** icon, then open **Agent Customizations** and select **MCP Servers** from the left navigation. This page lets you add and manage MCP servers for your workspace. Open the MCP Servers panel *** ## Step 2 — Add a New MCP Server Click the **+** button in the upper-right corner of the **MCP Servers** page. From the available options, select: ```text HTTP (HTTP or Server-Sent Events) ``` This option allows Visual Studio Code to connect to a remote MCP server over HTTP. Select the HTTP MCP server type *** ## Step 3 — Enter the Geonode MCP Endpoint When prompted, enter the Geonode MCP endpoint: ```text https://scraper.geonode.io/mcp ``` Press **Enter** to save the server configuration. Enter the Geonode MCP endpoint *** ## Step 4 — Verify the Connection After adding the endpoint, Visual Studio Code automatically attempts to connect to the MCP server. If the connection is successful, the server appears under **Workspace** with a **Running** status. Verify the MCP server is running If the Geonode MCP server displays **Running**, the connection has been established successfully and GitHub Copilot can begin using the available MCP tools. *** ## Step 5 — Review the Generated Configuration Visual Studio Code automatically creates an **mcp.json** file in your workspace. This file stores your MCP server configuration and can be updated later if your server configuration changes. Generated mcp.json configuration The Geonode MCP server requires authentication using your Geonode API key. Configure authentication according to your deployment before using the server. *** ## Step 6 — Start Using Geonode MCP Once the MCP server is connected, open the **GitHub Copilot Chat** panel and make sure **Agent** mode is selected. You can now interact with the Geonode MCP server using natural language. GitHub Copilot automatically selects the appropriate MCP tool based on your request. For example: ### Extract a web page ```text Extract the content from https://geonode.com ``` ### Summarize a page ```text Read https://geonode.com and summarize the homepage. ``` ### Crawl a website ```text Crawl https://docs.geonode.com and summarize its documentation. ``` ### Process multiple URLs ```text Extract and summarize the following URLs: https://example.com/page1 https://example.com/page2 https://example.com/page3 ``` Using Geonode MCP with GitHub Copilot *** ## How GitHub Copilot Uses MCP When you submit a prompt in **Agent Mode**, GitHub Copilot analyzes your request and automatically determines whether an MCP tool is required. If your request involves extracting web content, crawling websites, or processing multiple URLs, GitHub Copilot invokes the appropriate Geonode MCP tool behind the scenes and returns the results directly in the chat. You don't need to manually choose or invoke MCP tools. *** ## FAQs Verify that: * The MCP endpoint is correct. * Your Geonode API key has been configured. * Your internet connection is active. After making any changes, restart the MCP server or reload Visual Studio Code. If the server does not display a **Running** status, verify that the endpoint and authentication have been configured correctly. Ensure that: * GitHub Copilot is signed in. * Agent Mode is enabled. * The Geonode MCP server is running. Then submit your request again. Open the **MCP Servers** panel, remove the server configuration, and save your changes. # Codex (/docs/scraper-api/guides/mcp/09_codex) The Codex desktop application supports the **Model Context Protocol (MCP)**, allowing it to securely connect to external tools such as the Geonode Scraper API. Once connected, Codex can automatically extract web pages, crawl websites, process multiple URLs, and retrieve structured content using natural language prompts. Before continuing, make sure you've completed the **Before You Start** guide and have: * A Geonode API key. * The Geonode MCP endpoint. * The Codex desktop application installed. * Internet access. *** ## Step 1 — Open MCP Settings Open the **Codex** desktop application. Click the **Settings** icon in the upper-right corner. Open Codex Settings *** ## Step 2 — Open the MCP Servers Page From the Settings menu: 1. Select **Plugins** from the left sidebar. 2. Open the **MCPs** tab. 3. Click **Add server**. This opens the configuration page for adding a custom MCP server. Open the MCP Servers page *** ## Step 3 — Configure the GeoNode MCP Server Enter the following information: ### Name You can use any descriptive name, for example: ```text geonode-server ``` ### Type Select: ```text Streamable HTTP ``` ### URL Enter the GeoNode MCP endpoint: ```text https://scraper.geonode.io/mcp ``` ### Authentication Under **Headers**, add the following header: | Key | Value | | ----------- | -------------------- | | `X-API-Key` | Your Geonode API key | After completing the configuration, click **Save**. Configure the GeoNode MCP server *** ## Step 4 — Verify the Connection After saving the configuration, return to the MCP Servers page. If the connection is successful, the GeoNode MCP server appears in the server list and is enabled. Verify the GeoNode MCP server If the GeoNode MCP server appears in the server list and is enabled, Codex is successfully connected and ready to use the available MCP tools. *** ## Step 5 — Test the MCP Server Open a new chat in Codex and ask it to list the available GeoNode MCP tools. For example: ```text List all available tools from the GeoNode MCP server. ``` If the integration is configured correctly, Codex will display the available MCP tools exposed by the GeoNode server. List the available GeoNode MCP tools *** ## Step 6 — Start Using GeoNode MCP Once the MCP server is connected, you can interact with the GeoNode Scraper API using natural language. For example: ### Extract a web page ```text Extract the content from https://geonode.com ``` ### Summarize a page ```text Read https://geonode.com and summarize the homepage. ``` ### Crawl a website ```text Crawl https://docs.geonode.com and summarize its documentation. ``` ### Process multiple URLs ```text Extract and summarize the following URLs: https://example.com/page1 https://example.com/page2 https://example.com/page3 ``` Codex automatically selects the appropriate GeoNode MCP tool based on your request. *** ## How Codex Uses MCP When you submit a prompt, Codex analyzes your request to determine whether an MCP tool is required. If your request involves extracting web content, crawling websites, processing multiple URLs, or retrieving statistics, Codex automatically invokes the appropriate GeoNode MCP tool and returns the results directly in the conversation. You don't need to manually choose or invoke MCP tools. *** ## FAQs Verify that: * The MCP endpoint is correct. * The `X-API-Key` header contains a valid Geonode API key. * Your internet connection is active. After making changes, save the configuration and try connecting again. Ensure that: * The server configuration has been saved. * The server URL is correct. * The server is enabled from the MCP Servers page. Verify that: * The GeoNode MCP server is enabled. * Your API key is valid. * The MCP server connection is active. Then try your prompt again. Authentication failures usually occur when: * The API key is incorrect. * The `X-API-Key` header is missing. * The API key has expired or been revoked. Update the API key in the server configuration and save the changes. Open **Settings → Plugins → MCPs**, select the GeoNode MCP server, and choose **Uninstall** or disable it from the server list. # Help & FAQ (/docs/scraper-api/guides/mcp/help-and-faq) import { Accordion, Accordions } from 'fumadocs-ui/components/accordion'; ## Troubleshooting * **No tools appear:** check that Node.js is installed (for `mcp-remote` setups), restart the client fully, and re-check the endpoint URL. * **401 / authentication errors:** check the `X-Api-Key` value and that the key is active in your Geonode dashboard. Keyless access does not require a key; see [Use Without an API Key](/docs/scraper-api/guides/mcp/01_use-without-an-api-key). * **Header not recognized:** with `mcp-remote`, keep the format `X-Api-Key:${VAR}` with no space after the colon. * **Only `extract` appears:** that is expected without an API key. Add `X-Api-Key` to unlock the full tool set. *** ## FAQs Any MCP-compatible client. Claude Desktop, Cursor, Claude Code, and Windsurf have direct support; Smithery and Docker MCP help you deploy and manage connections. No, not to try. Point your client at `https://scraper.geonode.io/mcp` with no key to use `extract` under free anonymous limits. Add an API key when you need more tools, async jobs, JavaScript rendering, or higher limits. See [Use Without an API Key](/docs/scraper-api/guides/mcp/01_use-without-an-api-key). Only what you ask it to scrape. With a key, that runs under your Geonode account and plan limits. Without a key, anonymous rate and volume limits apply. Not for keyless trial use. For authenticated access you need a Geonode API key and a plan that covers your usage. On the client side, some assistants require their own paid tier to turn on MCP. # Search Workflows (/docs/scraper-api/guides/search/01_search_overview) The Search API lets you submit a search query and receive search results. Each search returns a unique `job_id`, which can be used to retrieve the full details of a completed search job. ## Complete Search Workflow The following diagram shows how the Search API endpoints work together. A search starts with a query and returns search results together with a unique job ID. *** ## Typical Workflow Most applications follow these steps when working with Search. ### Step 1 — Submit a Search Start by submitting a search query. ```http POST /v1/search ``` The request requires a `query` field. ```json { "query": "web scraping" } ``` You can also provide optional search settings such as: * `locale` * `page` * `safe` * `time_range` The `query` can contain between 1 and 1,000 characters. The available values and limits for the optional fields are covered in [Search Parameters](/docs/scraper-api/guides/search/03_search_parameters). *** ### Step 2 — Receive the Search Response A successful search returns a `SearchResponse`. The response includes: * `job_id` * `query` * `page` * `attempts` * `results` * `suggestions` * `spelling_correction` The `job_id` uniquely identifies the search job. *** ### Step 3 — Review the Results The `results` array contains the search results returned for the requested page. Each result contains: * `position` * `title` * `url` It can also include: * `snippet` * `thumbnail` * `displayed_url` * `source_host` The `position` field represents the 1-based rank of the result on the page. *** ## Finding Previous Search Jobs Search responses include a `job_id` that identifies the search. You can list search jobs for the authenticated user using: ```http GET /v1/search/jobs ``` The Search Jobs endpoint supports optional filters for: * Search query * Job status * Start date * End date * Page * Page size The `page_size` can be between 1 and 100 and defaults to 10. *** ## Retrieving Search Job Details After you have a search `job_id`, retrieve the full details of a completed search job using: ```http GET /v1/search/{job_id} ``` The response provides the search configuration and its results. It can include: * `job_id` * `query` * `locale` * `page` * `safe` * `time_range` * `status` * `results` * `results_count` * `suggestions` * `spelling_correction` * `attempts` * `final_url` * `duration_ms` * `tokens_charged` * `block_reason` * `error_code` * `error_message` * `created_at` * `completed_at` The endpoint description specifically defines this operation as retrieving the full details and results for a **completed search job**. *** ## Pagination The Search API supports result pagination through the `page` request parameter. The page number must be between **1 and 20**. For example: ```json { "query": "web scraping", "page": 2 } ``` The response returns the page that was requested in the `page` field. *** ## Search Options The Search API provides additional request options for controlling the search. ### Locale Use `locale` to specify a language code, language-region tag, or `all`. Examples supported by the OpenAPI schema include: ```text en en-US all ``` ### Safe Search Use `safe` to select a safe-search level. Supported values are: ```text off moderate strict ``` The default is `off`. ### Time Range Use `time_range` to restrict results to a recency window. Supported values are: ```text day week month year ``` ### Result Page Use `page` to request a specific result page from 1 through 20. These options can be combined with the required `query` field in the same request. *** ## Common Workflow Patterns ### Standard Search This is the basic Search workflow. *** ### Search with Pagination Use the `page` parameter when you need to retrieve another result page. *** ### Find and Retrieve a Previous Search Use this workflow when you need to find a previous search job and retrieve its details. *** ## Best Practices * Store the `job_id` returned by the Search API if you need to retrieve the search later. * Use `page` to request additional result pages. * Use `locale` when you need to specify the search locale. * Use `safe` when you need a specific safe-search level. * Use `time_range` when you need to restrict results to a specific recency window. * Use the Search Jobs endpoint when you need to find previous search jobs. * Use the Search Job Details endpoint to retrieve the full details and results of a completed search job. *** ## Next Steps You now understand the main Search API workflow. Continue with: * [Getting Started](/docs/scraper-api/guides/search/02_getting_started) — Make your first Search API request. # Getting Started with Search (/docs/scraper-api/guides/search/02_getting_started) In this guide, you'll make your first Search API request and see how the API returns search results. ## Before You Begin Before making a Search request, make sure you have: * A Geonode API key * The Scraper API base URL If you haven't completed the initial setup, see [Quick Start Guide](/docs/scraper-api/quick-start). ## Make Your First Search Create a search by sending a `POST` request to the Search endpoint. ```http POST /v1/search ``` For your first search, you only need to provide a search query. ### Example Request ```bash curl -X POST "https://scraper.geonode.io/v1/search" \ -H "X-Api-Key: YOUR_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "query": "web scraping best practice" }' ``` You can also send the request body as JSON: ```json { "query": "web scraping best practice" } ``` ## Example Response A successful request returns the search results together with a unique job ID. ```json { "job_id": "a478199c-54d2-4884-ba8d-d6d7568614e9", "query": "web scraping best practice", "page": 1, "attempts": 2, "results": [ { "position": 1, "title": "Web Scraping Best Practices in 2026", "url": "https://www.scrapingbee.com/blog/web-scraping-best-practices/", "snippet": "Web scraping is the automated process of retrieving data from websites and transforming raw HTML or other web data into structured formats for analysis or use. Whether you are working on a small web scraping project or managing large-scale data collection activities, choosing the right web scraping tool and following best practices is essential. In this article, I'll walk you through the best ...", "thumbnail": null, "displayed_url": "www.scrapingbee.com", "source_host": "www.scrapingbee.com" } ], "suggestions": [], "spelling_correction": null } ``` The response contains: | Field | Description | | --------------------- | ------------------------------------------------ | | `job_id` | Unique identifier for the search job. | | `query` | The search query submitted to the API. | | `page` | The result page returned by the search. | | `attempts` | Number of attempts made for the search. | | `results` | Search results returned for the query. | | `suggestions` | Search suggestions returned by the API. | | `spelling_correction` | Spelling correction information, when available. | The `results` array contains the webpages returned for the search query. You can learn more about the individual result fields in [Understanding Search Results](/docs/scraper-api/guides/search/04_understanding_search_results). ## What Happens Next? Your first Search request is now complete. The next guides explain how to customize and work with Search: * [Search Parameters](/docs/scraper-api/guides/search/03_search_parameters) — Learn how to configure your search request. # Search Parameters (/docs/scraper-api/guides/search/03_search_parameters) The Search API provides several parameters to control what results are returned. The `query` parameter is required, while `locale`, `page`, `safe`, and `time_range` are optional. ## query The `query` parameter contains the search query you want to submit. It is the only required parameter. ### Requirements * Type: `string` * Minimum length: `1` * Maximum length: `1000` ### Example ```json { "query": "web scraping best practice" } ``` Every Search request must include `query`. The value must contain between 1 and 1,000 characters. *** ## locale The `locale` parameter specifies the locale for the search. It accepts: * A language code, such as `en` * A language-region tag, such as `en-US` * `all` ### Example ```json { "query": "web scraping", "locale": "en-US" } ``` The `locale` parameter is optional. *** ## page The `page` parameter specifies which result page to fetch. ### Requirements * Type: `integer` * Minimum: `1` * Maximum: `20` * Default: `1` ### Example ```json { "query": "web scraping", "page": 2 } ``` The OpenAPI defines example values of `1`, `2`, and `20`. The `page` value must be between `1` and `20`. *** ## safe The `safe` parameter controls the safe-search filtering level. The accepted values are: | Value | Description | | ---------- | ---------------------------------- | | `off` | Safe-search filtering is disabled. | | `moderate` | Moderate safe-search filtering. | | `strict` | Strict safe-search filtering. | The default value is `off`. ### Example ```json { "query": "web scraping", "safe": "moderate" } ``` The OpenAPI only defines these three values for `safe`. Use only `off`, `moderate`, or `strict` for the `safe` parameter. *** ## time\_range The `time_range` parameter restricts results to a recency window. The accepted values are: | Value | Recency window | | ------- | -------------- | | `day` | Day | | `week` | Week | | `month` | Month | | `year` | Year | ### Example ```json { "query": "web scraping", "time_range": "week" } ``` The parameter is optional and can be set to `day`, `week`, `month`, or `year`. *** ## Using Multiple Parameters You can combine the optional parameters with the required `query` parameter in a single request. For example: ```json { "query": "web scraping best practice", "locale": "en-US", "page": 2, "safe": "moderate", "time_range": "month" } ``` The same parameters can be sent to: ```http POST /v1/search ``` *** ## Complete Request Example The following request uses all available Search request parameters: ```bash curl -X POST "YOUR_SCRAPER_API_URL/v1/search" \ -H "X-Api-Key: YOUR_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "query": "web scraping best practice", "locale": "en-US", "page": 2, "safe": "moderate", "time_range": "month" }' ``` *** ## Parameter Summary | Parameter | Required | Type | Default | Accepted values / limits | | ------------ | -------- | --------- | ------- | -------------------------------------------- | | `query` | Yes | `string` | — | 1–1000 characters | | `locale` | No | `string` | — | Language code, language-region tag, or `all` | | `page` | No | `integer` | `1` | `1–20` | | `safe` | No | `string` | `off` | `off`, `moderate`, `strict` | | `time_range` | No | `string` | — | `day`, `week`, `month`, `year` | These are the complete request parameters defined by the current `SearchRequest` schema. *** ## What's Next? You now know how to configure a Search request. Continue with [Understanding Search Results](/docs/scraper-api/guides/search/04_understanding_search_results) to learn how the API structures the returned search results. # Understanding Search Results (/docs/scraper-api/guides/search/04_understanding_search_results) The Search API returns a `SearchResponse` containing the submitted query, the returned result page, search results, and additional search information. ## Search Response A successful Search request returns the following top-level fields: | Field | Description | | --------------------- | ----------------------------------------- | | `job_id` | Unique job identifier for this search. | | `query` | Search query that was submitted. | | `page` | Result page that was returned. | | `attempts` | Number of upstream attempts made. | | `results` | Search results returned by the API. | | `suggestions` | Query suggestions returned by the engine. | | `spelling_correction` | Spelling correction applied to the query. | The OpenAPI defines `job_id`, `query`, `page`, `attempts`, `results`, and `suggestions` as required response fields. `spelling_correction` can be a string or `null`. ## Example Response The following is a real response from the Search API: ```json { "job_id": "a478199c-54d2-4884-ba8d-d6d7568614e9", "query": "web scraping best practice", "page": 1, "attempts": 2, "results": [ { "position": 1, "title": "Web Scraping Best Practices in 2026", "url": "https://www.scrapingbee.com/blog/web-scraping-best-practices/", "snippet": "Web scraping is the automated process of retrieving data from websites and transforming raw HTML or other web data into structured formats for analysis or use. Whether you are working on a small web scraping project or managing large-scale data collection activities, choosing the right web scraping tool and following best practices is essential. In this article, I'll walk you through the best ...", "thumbnail": null, "displayed_url": "www.scrapingbee.com", "source_host": "www.scrapingbee.com" } ], "suggestions": [], "spelling_correction": null } ``` The response above uses the fields defined by the `SearchResponse` and `SearchHitModel` schemas. ## Search Results The `results` field is an array of `SearchHitModel` objects. Each search result contains three required fields: * `position` * `title` * `url` It can also contain: * `snippet` * `thumbnail` * `displayed_url` * `source_host` ### Result Fields | Field | Required | Description | | --------------- | -------- | ------------------------------------------ | | `position` | Yes | 1-based rank of this result on the page. | | `title` | Yes | Result title. | | `url` | Yes | Result target URL. | | `snippet` | No | Result preview text. | | `thumbnail` | No | Thumbnail image URL, if available. | | `displayed_url` | No | Human-readable URL as shown by the engine. | | `source_host` | No | Hostname the result was served from. | The OpenAPI defines `position`, `title`, and `url` as the required fields for each result. `thumbnail`, `displayed_url`, and `source_host` can be `null`. ## Working With a Result A result can be accessed from the `results` array. For example, the first result in the response is: ```json { "position": 1, "title": "Web Scraping Best Practices in 2026", "url": "https://www.scrapingbee.com/blog/web-scraping-best-practices/", "snippet": "Web scraping is the automated process of retrieving data from websites and transforming raw HTML or other web data into structured formats for analysis or use.", "thumbnail": null, "displayed_url": "www.scrapingbee.com", "source_host": "www.scrapingbee.com" } ``` The `position` identifies the result's 1-based rank on the returned page. The `title` and `url` identify the result, while the remaining fields provide additional result information when available. ## Suggestions The `suggestions` field contains query suggestions returned by the engine. In the example response, no suggestions were returned: ```json { "suggestions": [] } ``` The OpenAPI defines `suggestions` as an array of strings. ## Spelling Correction The `spelling_correction` field contains the spelling correction applied to the query. It can contain a string or `null`. In the example response: ```json { "spelling_correction": null } ``` The OpenAPI defines this field as either a string or `null`. ## Complete Response Structure The Search response can be represented by the following structure: ```text SearchResponse ├── job_id ├── query ├── page ├── attempts ├── results[] │ ├── position │ ├── title │ ├── url │ ├── snippet │ ├── thumbnail │ ├── displayed_url │ └── source_host ├── suggestions[] └── spelling_correction ``` This structure follows the `SearchResponse` and `SearchHitModel` schemas defined in the OpenAPI specification. ## What's Next? You now understand the structure of a Search API response. Continue with [Pagination and Filters](/docs/scraper-api/guides/search/05_pagination_and_filters) to learn how to work with result pages and the available Search request filters. # Pagination and Filters (/docs/scraper-api/guides/search/05_pagination_and_filters) The Search API provides pagination and optional filters that you can include in the search request. You can use `page` to request a specific result page and use `locale`, `safe`, and `time_range` to configure the search request. ## Pagination Use the `page` parameter to specify which result page to fetch. The OpenAPI defines the following limits: * Minimum: `1` * Maximum: `20` * Default: `1` For example, to request the second result page: ```json { "query": "web scraping best practice", "page": 2 } ``` The `page` value must be an integer between `1` and `20`. ### Request a Specific Page You can request any page from `1` through `20`. ```bash curl -X POST "YOUR_SCRAPER_API_URL/v1/search" \ -H "X-Api-Key: YOUR_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "query": "web scraping best practice", "page": 2 }' ``` The returned `page` field identifies the result page that was returned. *** ## Locale Use the `locale` parameter to specify the locale for the search. The OpenAPI accepts: * A language code, such as `en` * A language-region tag, such as `en-US` * `all` For example: ```json { "query": "web scraping", "locale": "en-US" } ``` The `locale` parameter is optional. *** ## Safe Search Use the `safe` parameter to specify the safe-search level. The supported values are: | Value | Description | | ---------- | ------------------------------------ | | `off` | Safe-search level set to `off`. | | `moderate` | Safe-search level set to `moderate`. | | `strict` | Safe-search level set to `strict`. | The default value is `off`. ### Example ```json { "query": "web scraping", "safe": "moderate" } ``` Only `off`, `moderate`, and `strict` are accepted values for this parameter. *** ## Time Range Use the `time_range` parameter to restrict results to a recency window. The supported values are: | Value | | ------- | | `day` | | `week` | | `month` | | `year` | For example: ```json { "query": "web scraping", "time_range": "week" } ``` The `time_range` parameter is optional. *** ## Combining Parameters You can include multiple optional parameters in the same Search request. For example: ```json { "query": "web scraping best practice", "locale": "en-US", "page": 2, "safe": "moderate", "time_range": "month" } ``` This request uses: * `locale` to specify `en-US` * `page` to request page `2` * `safe` with the `moderate` value * `time_range` with the `month` value All four parameters are defined as optional in the `SearchRequest` schema. *** ## Complete Request The following example combines all available Search request options: ```bash curl -X POST "YOUR_SCRAPER_API_URL/v1/search" \ -H "X-Api-Key: YOUR_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "query": "web scraping best practice", "locale": "en-US", "page": 2, "safe": "moderate", "time_range": "month" }' ``` ### Parameter Summary | Parameter | Required | Type | Default | Values / Limits | | ------------ | -------- | ------------------ | ------- | -------------------------------------------- | | `query` | Yes | `string` | — | 1–1000 characters | | `locale` | No | `string` or `null` | — | Language code, language-region tag, or `all` | | `page` | No | `integer` | `1` | `1–20` | | `safe` | No | `string` | `off` | `off`, `moderate`, `strict` | | `time_range` | No | `string` or `null` | — | `day`, `week`, `month`, `year` | These are the request parameters defined by the current `SearchRequest` schema. *** ## What's Next? You now know how to paginate Search results and configure the available Search filters. Continue with [Search Jobs](/docs/scraper-api/guides/search/06_search_jobs) to learn how to list and retrieve search jobs. # Search Jobs (/docs/scraper-api/guides/search/06_search_jobs) Every Search request returns a unique `job_id`. You can use this ID to retrieve the details of a completed search job, or use the Search Jobs endpoint to find previous searches. ## Search Job Workflow The following diagram shows how the Search Job endpoints work together. ## List Search Jobs Use the Search Jobs endpoint to list search jobs for the authenticated user. ```http GET /v1/search/jobs ``` The endpoint supports optional filtering and pagination. ### Example Request ```bash curl "YOUR_SCRAPER_API_URL/v1/search/jobs" \ -H "X-Api-Key: YOUR_API_KEY" ``` ### Example Response The following is a real response from the Search API: ```json { "jobs": [ { "job_id": "a478199c-54d2-4884-ba8d-d6d7568614e9", "query": "web scraping best practice", "status": "completed", "results_count": 10, "duration_ms": 1800, "block_reason": null, "error_code": null, "created_at": "2026-08-16T17:44:07.447657Z", "completed_at": "2026-08-16T17:44:09.673985Z" }, { "job_id": "ed9f8595-be2a-4d3e-873b-1a1c52620a0e", "query": "https://scraper.geonode.io", "status": "completed", "results_count": 10, "duration_ms": 5077, "block_reason": null, "error_code": null, "created_at": "2026-08-16T17:43:29.525681Z", "completed_at": "2026-08-16T17:43:35.053054Z" } ], "page": 1, "page_size": 10, "page_count": 1 } ``` The response contains: | Field | Description | | ------------ | --------------------------------- | | `jobs` | Search jobs on the current page. | | `page` | Current page number. | | `page_size` | Number of jobs returned per page. | | `page_count` | Total number of pages available. | These fields are defined by the `SearchJobsResponse` schema. ## Search Job Fields Each item in the `jobs` array is a `SearchListItemResponse`. The available fields are: | Field | Description | | --------------- | ---------------------------------------------- | | `job_id` | Unique job identifier. | | `query` | Search query that was submitted. | | `status` | Job status. | | `results_count` | Number of results returned. | | `duration_ms` | Execution time in milliseconds. | | `block_reason` | Upstream block classification for failed jobs. | | `error_code` | Machine-readable error code for failed jobs. | | `created_at` | Time when the search job was created. | | `completed_at` | Time when the search job finished. | The OpenAPI defines `job_id`, `query`, `status`, and `created_at` as required fields. The remaining fields can be absent or nullable according to the schema. ## Filter Search Jobs You can filter the jobs returned by `GET /v1/search/jobs`. The available filters are: | Parameter | Description | | ------------ | --------------------------------------------- | | `query` | Filter by search query using a partial match. | | `status` | Filter by job status. | | `start_date` | Filter jobs created on or after this date. | | `end_date` | Filter jobs created on or before this date. | For example, to filter jobs by query: ```bash curl "YOUR_SCRAPER_API_URL/v1/search/jobs?query=web%20scraping" \ -H "X-Api-Key: YOUR_API_KEY" ``` To filter by status: ```bash curl "YOUR_SCRAPER_API_URL/v1/search/jobs?status=completed" \ -H "X-Api-Key: YOUR_API_KEY" ``` The OpenAPI defines these four filters for the Search Jobs endpoint. ## Paginate Search Jobs The Search Jobs endpoint supports pagination using: * `page` * `page_size` `page` specifies the page number and defaults to `1`. `page_size` specifies the number of results per page. It must be between `1` and `100` and defaults to `10`. For example: ```bash curl "YOUR_SCRAPER_API_URL/v1/search/jobs?page=1&page_size=20" \ -H "X-Api-Key: YOUR_API_KEY" ``` The response includes `page`, `page_size`, and `page_count` so you can determine the current page and the total number of available pages. The `page_size` parameter accepts values from `1` through `100`. ## Retrieve Search Job Details Once you have a `job_id`, use the Search Job endpoint to retrieve the full details and results for a completed search job. ```http GET /v1/search/{job_id} ``` For example: ```bash curl "YOUR_SCRAPER_API_URL/v1/search/a478199c-54d2-4884-ba8d-d6d7568614e9" \ -H "X-Api-Key: YOUR_API_KEY" ``` ### Example Response The following is a real response for the search job used in the examples above: ```json { "job_id": "a478199c-54d2-4884-ba8d-d6d7568614e9", "query": "web scraping best practice", "locale": null, "page": 1, "safe": "off", "time_range": null, "status": "completed", "results": [ { "position": 1, "title": "Web Scraping Best Practices in 2026", "url": "https://www.scrapingbee.com/blog/web-scraping-best-practices/", "snippet": "Web scraping is the automated process of retrieving data from websites and transforming raw HTML or other web data into structured formats for analysis or use. Whether you are working on a small web scraping project or managing large-scale data collection activities, choosing the right web scraping tool and following best practices is essential. In this article, I'll walk you through the best ...", "thumbnail": null, "displayed_url": "www.scrapingbee.com", "source_host": "www.scrapingbee.com" }, { "position": 2, "title": "10 Best Sample Websites for Web Scraping Practice in 2026", "url": "https://thunderbit.com/blog/best-web-scraping-test-sites", "snippet": "Practice web scraping on the best sample sites for all skill levels. Thunderbit helps automate extraction, handle complex sites, and streamline data workflows.", "thumbnail": null, "displayed_url": "thunderbit.com", "source_host": "thunderbit.com" }, { "position": 3, "title": "11 Web Scraping Best Practices for Reliable Data Collection", "url": "https://scrapfly.io/blog/posts/web-scraping-best-practices", "snippet": "11 web scraping best practices for 2026: robots.txt, rate limiting, hidden APIs, proxy rotation, retries, validation, and monitoring, with working code.", "thumbnail": null, "displayed_url": "scrapfly.io", "source_host": "scrapfly.io" } ], "results_count": 10, "suggestions": [], "spelling_correction": null, "attempts": 2, "final_url": "http://searxng:8080/search?q=web+scraping+best+practice&format=json&engines=duckduckgo&pageno=1&safesearch=0", "duration_ms": 1800, "tokens_charged": 1, "block_reason": null, "error_code": null, "error_message": null, "created_at": "2026-08-16T17:44:07.447657Z", "completed_at": "2026-08-16T17:44:09.673985Z" } ``` The `GET /v1/search/{job_id}` endpoint is documented as retrieving the full details and results for a completed search job. ## Search Job Details The detailed response contains the search configuration, status, results, and execution information. | Field | Description | | --------------------- | ---------------------------------------------- | | `job_id` | Unique job identifier. | | `query` | Search query that was submitted. | | `locale` | Locale requested for the search. | | `page` | Result page that was requested. | | `safe` | Safe-search level requested. | | `time_range` | Time range filter, if requested. | | `status` | Job status. | | `results` | Search results. | | `results_count` | Number of results returned. | | `suggestions` | Query suggestions returned by the engine. | | `spelling_correction` | Spelling correction applied to the query. | | `attempts` | Number of upstream attempts made. | | `final_url` | Final upstream URL that produced the results. | | `duration_ms` | Execution time in milliseconds. | | `tokens_charged` | Tokens charged for the job. | | `block_reason` | Upstream block classification for failed jobs. | | `error_code` | Machine-readable error code. | | `error_message` | Human-readable error description. | | `created_at` | Time when the search job was created. | | `completed_at` | Time when the search job finished. | These fields are defined by `SearchJobDetailResponse` in the OpenAPI. ## Job Status Search jobs use the `JobStatus` schema. The available statuses are: ```text queued processing completed failed cancelled ``` The Search Jobs response uses this status field for each listed search job. ## Complete Search Job Workflow ## What's Next? You now know how to list Search jobs, filter and paginate the job list, and retrieve the details of a completed search job. Continue with [Search Errors](/docs/scraper-api/guides/search/07_search_errors) to learn about the Search API error responses. # Search Errors (/docs/scraper-api/guides/search/07_search_errors) The Search API can return different error responses when a request cannot be completed. The response format depends on the type of error. The OpenAPI defines `SearchErrorResponse` for Search-specific errors and `ErrorResponse` for other API errors. ## Search Error Response Search-specific errors use the following response structure: ```json { "error": "error_code", "message": "Human-readable error description", "details": {} } ``` The `error` and `message` fields are required. The `details` field is optional and can contain additional error context. ### Error Fields | Field | Description | | --------- | ----------------------------------------- | | `error` | Machine-readable error code. | | `message` | Human-readable error description. | | `details` | Additional error context, when available. | The OpenAPI gives `no_results` and `upstream_blocked` as examples of the `error` field. *** ## Search-Specific Errors The Search endpoint documents two responses that use `SearchErrorResponse`. ### `422` — Unprocessable Entity The Search endpoint can return `422` with a `SearchErrorResponse`. ```http 422 Unprocessable Entity ``` This response uses the following structure: ```json { "error": "no_results", "message": "Human-readable error description", "details": {} } ``` The OpenAPI does not define a fixed message or `details` structure for this response, so use the values returned by the API. ### `502` — Bad Gateway The Search endpoint can return `502` with a `SearchErrorResponse`. ```http 502 Bad Gateway ``` The response uses the same Search-specific error structure: ```json { "error": "upstream_blocked", "message": "Human-readable error description", "details": {} } ``` The OpenAPI documents `upstream_blocked` as an example machine-readable error code. The exact `error`, `message`, and `details` values depend on the response returned by the API. Do not assume a specific message or details object. *** ## Other Search API Errors The Search endpoint also documents the following HTTP responses: | Status | Description | Response | | ------ | ------------------------------- | ---------------------------- | | `401` | Unauthorized | `ErrorResponse` | | `402` | Insufficient token balance | `ErrorResponse` | | `408` | Request timed out | `ErrorResponse` | | `422` | Unprocessable Entity | `SearchErrorResponse` | | `429` | Request throttled | No response schema specified | | `500` | Internal server error | `ErrorResponse` | | `502` | Bad Gateway | `SearchErrorResponse` | | `503` | Service temporarily unavailable | `ErrorResponse` | These responses and their associated schemas are defined on `POST /v1/search` in the OpenAPI specification. *** ## Handle Errors in Your Application Check the HTTP status code before processing a successful Search response. For Search-specific errors, inspect the returned `error`, `message`, and `details` fields. For example: ```python response = requests.post( "YOUR_SCRAPER_API_URL/v1/search", headers={ "X-Api-Key": "YOUR_API_KEY", "Content-Type": "application/json" }, json={ "query": "web scraping best practice" } ) if response.status_code == 200: data = response.json() else: error = response.json() print(error) ``` The exact error fields available depend on the response documented for that HTTP status. *** ## Search Error Summary The Search endpoint documents these HTTP responses: ```text 200 Successful Response 401 Unauthorized 402 Insufficient token balance 408 Request timed out 422 Unprocessable Entity 429 Request throttled 500 Internal server error 502 Bad Gateway 503 Service temporarily unavailable ``` The `422` and `502` responses use `SearchErrorResponse`. The other documented error responses use `ErrorResponse`, except `429`, for which the OpenAPI does not specify a response schema. *** ## What's Next? You now know the error responses documented for the Search API. For the complete request and response definitions, see the [Search API Reference](/docs/scraper-api/v1/reference). # Extract a JavaScript-Rendered Product Page (/docs/scraper-api/real-world/ecommerce/01_javascript_rendered_product_page) This real-world example shows how to extract product content from a JavaScript-rendered e-commerce page using the Scraper API. We will use the ScrapingCourse JavaScript Rendering page and compare the same extraction with JavaScript rendering disabled and enabled. ## Use Case Modern websites often use JavaScript to render or load content after the initial page response. For this example, we use: ```text https://www.scrapingcourse.com/javascript-rendering ``` The page contains multiple products with their names, prices, and product links. The extracted content includes products such as: * Chaz Kangeroo Hoodie — $52 * Teton Pullover Hoodie — $70 * Bruno Compete Hoodie — $63 * Frankie Sweatshirt — $60 * Hollister Backyard Sweatshirt — $52 * Stark Fundamental Hoodie — $42 * Hero Hoodie — $54 * Oslo Trek Hoodie — $42 * Abominable Hoodie — $69 * Mach Street Sweatshirt — $62 * Grayson Crewneck Sweatshirt — $64 * Ajax Full-Zip Sweatshirt — $69 ## Before You Start You need: * A Geonode API key * Python 3.9 or later * The `requests` package Store your API key in an environment variable: ```bash export GEONODE_SCRAPER_API_KEY="YOUR_API_KEY" ``` On Windows PowerShell: ```powershell $env:GEONODE_SCRAPER_API_KEY="YOUR_API_KEY" ``` Never put your Geonode API key directly in your source code or commit it to your repository. Store it in an environment variable or `.env` file instead. ## Install the Dependency Install the `requests` package: ```bash pip install requests ``` ## Case 1: Extract Without JavaScript Rendering First, send the request with JavaScript rendering disabled. ```python import os import requests api_key = os.environ["GEONODE_SCRAPER_API_KEY"] response = requests.post( "https://scraper.geonode.io/v1/extract", headers={ "X-Api-Key": api_key, "Content-Type": "application/json", }, json={ "url": "https://www.scrapingcourse.com/javascript-rendering", "formats": ["markdown"], "render_js": False, "processing_mode": "sync", }, ) response.raise_for_status() result = response.json() print(result["data"]["markdown"]) ``` ### Case 1 Result The actual extraction returned the following Markdown: ```markdown --- meta-viewport: width=device-width, initial-scale=1.0 title: JS Rendering Challenge to Learn Web Scraping - ScrapingCourse.com --- # JS Rendering ![](https://www.scrapingcourse.com/assets/images/challenge.svg) ## Challenge Enable JavaScript to see products [Chaz Kangeroo Hoodie Chaz Kangeroo Hoodie $52](https://scrapingcourse.com/ecommerce/product/chaz-kangeroo-hoodie) [Teton Pullover Hoodie Teton Pullover Hoodie $70](https://scrapingcourse.com/ecommerce/product/teton-pullover-hoodie) [Bruno Compete Hoodie Bruno Compete Hoodie $63](https://scrapingcourse.com/ecommerce/product/bruno-compete-hoodie) [Frankie Sweatshirt Frankie Sweatshirt $60](https://scrapingcourse.com/ecommerce/product/frankie-sweatshirt) [Hollister Backyard Sweatshirt Hollister Backyard Sweatshirt $52](https://scrapingcourse.com/ecommerce/product/hollister-backyard-sweatshirt) [Stark Fundamental Hoodie Stark Fundamental Hoodie $42](https://scrapingcourse.com/ecommerce/product/stark-fundamental-hoodie) [Hero Hoodie Hero Hoodie $54](https://scrapingcourse.com/ecommerce/product/hero-hoodie) [Oslo Trek Hoodie Oslo Trek Hoodie $42](https://scrapingcourse.com/ecommerce/product/oslo-trek-hoodie) [Abominable Hoodie Abominable Hoodie $69](https://scrapingcourse.com/ecommerce/product/abominable-hoodie) [Mach Street Sweatshirt Mach Street Sweatshirt $62](https://scrapingcourse.com/ecommerce/product/mach-street-sweatshirt) [Grayson Crewneck Sweatshirt Grayson Crewneck Sweatshirt $64](https://scrapingcourse.com/ecommerce/product/grayson-crewneck-sweatshirt) [Ajax Full-Zip Sweatshirt Ajax Full-Zip Sweatshirt $69](https://scrapingcourse.com/ecommerce/product/ajax-full-zip-sweatshirt) ``` ## Case 2: Extract With JavaScript Rendering Now run the same extraction with JavaScript rendering enabled. The only change is: ```json "render_js": true ``` ```python import os import requests api_key = os.environ["GEONODE_SCRAPER_API_KEY"] response = requests.post( "https://scraper.geonode.io/v1/extract", headers={ "X-Api-Key": api_key, "Content-Type": "application/json", }, json={ "url": "https://www.scrapingcourse.com/javascript-rendering", "formats": ["markdown"], "render_js": True, "processing_mode": "sync", }, ) response.raise_for_status() result = response.json() print(result["data"]["markdown"]) ``` ### Case 2 Result The actual extraction returned the following Markdown: ```markdown --- meta-viewport: width=device-width, initial-scale=1.0 title: JS Rendering Challenge to Learn Web Scraping - ScrapingCourse.com --- # JS Rendering ![](https://www.scrapingcourse.com/assets/images/challenge.svg) ## Challenge Enable JavaScript to see products [Chaz Kangeroo Hoodie Chaz Kangeroo Hoodie $52](https://scrapingcourse.com/ecommerce/product/chaz-kangeroo-hoodie) [Teton Pullover Hoodie Teton Pullover Hoodie $70](https://scrapingcourse.com/ecommerce/product/teton-pullover-hoodie) [Bruno Compete Hoodie Bruno Compete Hoodie $63](https://scrapingcourse.com/ecommerce/product/bruno-compete-hoodie) [Frankie Sweatshirt Frankie Sweatshirt $60](https://scrapingcourse.com/ecommerce/product/frankie-sweatshirt) [Hollister Backyard Sweatshirt Hollister Backyard Sweatshirt $52](https://scrapingcourse.com/ecommerce/product/hollister-backyard-sweatshirt) [Stark Fundamental Hoodie Stark Fundamental Hoodie $42](https://scrapingcourse.com/ecommerce/product/stark-fundamental-hoodie) [Hero Hoodie Hero Hoodie $54](https://scrapingcourse.com/ecommerce/product/hero-hoodie) [Oslo Trek Hoodie Oslo Trek Hoodie $42](https://scrapingcourse.com/ecommerce/product/oslo-trek-hoodie) [Abominable Hoodie Abominable Hoodie $69](https://scrapingcourse.com/ecommerce/product/abominable-hoodie) [Mach Street Sweatshirt Mach Street Sweatshirt $62](https://scrapingcourse.com/ecommerce/product/mach-street-sweatshirt) [Grayson Crewneck Sweatshirt Grayson Crewneck Sweatshirt $64](https://scrapingcourse.com/ecommerce/product/grayson-crewneck-sweatshirt) [Ajax Full-Zip Sweatshirt Ajax Full-Zip Sweatshirt $69](https://scrapingcourse.com/ecommerce/product/ajax-full-zip-sweatshirt) ``` ## Compare the Results You can compare both requests directly: ### Request ```json { "url": "https://www.scrapingcourse.com/javascript-rendering", "formats": ["markdown"], "render_js": false, "processing_mode": "sync" } ``` ### Result The captured result contains the page heading, the JavaScript challenge message, and the 12 product links with their prices. ### Request ```json { "url": "https://www.scrapingcourse.com/javascript-rendering", "formats": ["markdown"], "render_js": true, "processing_mode": "sync" } ``` ### Result The captured result also contains the page heading, the JavaScript challenge message, and the 12 product links with their prices. For this particular target and the captured runs, both cases returned the same Markdown content. The example therefore demonstrates how to enable JavaScript rendering, but the captured result does not establish that `render_js` is required for this page. ## Extracted Products The returned Markdown contains 12 product entries. | Product | Price | | ----------------------------- | ----: | | Chaz Kangeroo Hoodie | $52 | | Teton Pullover Hoodie | $70 | | Bruno Compete Hoodie | $63 | | Frankie Sweatshirt | $60 | | Hollister Backyard Sweatshirt | $52 | | Stark Fundamental Hoodie | $42 | | Hero Hoodie | $54 | | Oslo Trek Hoodie | $42 | | Abominable Hoodie | $69 | | Mach Street Sweatshirt | $62 | | Grayson Crewneck Sweatshirt | $64 | | Ajax Full-Zip Sweatshirt | $69 | ## Understanding the Response The extracted page content is available under: ```text data.markdown ``` You can use the returned Markdown directly or process it further in your application. For example: ```python markdown = result["data"]["markdown"] print(markdown) ``` The Scraper API returns page content rather than automatically converting the page into structured product fields. If your application needs fields such as: ```text name price url sku availability ``` you can parse the returned Markdown or HTML in your own application. The Scraper API returns extracted page content as Markdown or HTML. It does not automatically return e-commerce fields such as product name, price, or SKU as structured fields. ## When to Use JavaScript Rendering Use: ```json "render_js": true ``` when the content you need depends on JavaScript execution. Common examples include: * Product information loaded after the page opens * Client-side rendered product listings * Content populated dynamically by JavaScript * Pages where the initial HTML does not contain the content you need For pages where the required content is already available without browser rendering, you can use: ```json "render_js": false ``` This avoids browser rendering when it is not needed. ## What About `wait_config`? This example does not include `wait_config`. The focused Case 2 probe successfully completed with: ```json { "render_js": true, "processing_mode": "sync", "wait_config": null } ``` It reported 12 product links and 12 prices in the returned Markdown. If a JavaScript-rendered page requires an explicit browser wait condition, you can use `wait_config` to control when the browser should consider the page ready for extraction. ## Complete Example Once you understand the difference between the two requests, the JavaScript-rendered version can be reduced to this: ```python import os import requests API_KEY = os.environ["GEONODE_SCRAPER_API_KEY"] URL = "https://www.scrapingcourse.com/javascript-rendering" response = requests.post( "https://scraper.geonode.io/v1/extract", headers={ "X-Api-Key": API_KEY, "Content-Type": "application/json", }, json={ "url": URL, "formats": ["markdown"], "render_js": True, "processing_mode": "sync", }, ) response.raise_for_status() result = response.json() print(result["data"]["markdown"]) ``` ## Request Parameters | Parameter | Value | Purpose | | ----------------- | ----------------------------------------------------- | ----------------------------------------- | | `url` | `https://www.scrapingcourse.com/javascript-rendering` | Target page to extract | | `formats` | `["markdown"]` | Returns the extracted content as Markdown | | `render_js` | `true` | Enables JavaScript rendering | | `processing_mode` | `sync` | Waits for the extraction to complete | ## Next Steps Now that you can extract a page with JavaScript rendering, continue to **Compare a Product Page Across Locations** to extract the same product page from different countries with proxy geo-targeting. You can also explore: * Extract multiple product URLs with Batch * Discover product URLs with Map and extract them with Batch * Crawl a product or content section * Run longer extractions asynchronously # Compare a Product Page Across Locations (/docs/scraper-api/real-world/ecommerce/02_geo_targeted_product_page) Retail sites often change what they show based on where the request comes from. That is useful when you are checking localized messaging, market selectors, shipping copy, or regional pricing — but it is hard to do from a single office IP. This example uses the Scraper API to fetch the **same product URL** twice: once through a United States residential proxy, and once through a United Kingdom residential proxy. Then we compare the Markdown. ## What you will get By the end, you will have: * Two Markdown extracts of the same Nike product page * Confirmed `metadata.proxy.country` values for each run * A simple side-by-side check for location-specific content ## Target page ```text https://www.nike.com/t/air-force-1-07-mens-shoes-jBrhbr/CW2288-111 ``` Why this page works well for a demo: * It is a real public product detail page * It responds differently by exit country * In our captured UK run, Nike surfaced an explicit location banner ## Setup You need a Geonode API key and Python with `requests` installed. ```bash pip install requests export GEONODE_SCRAPER_API_KEY="YOUR_API_KEY" ``` PowerShell: ```powershell pip install requests $env:GEONODE_SCRAPER_API_KEY="YOUR_API_KEY" ``` Do not hardcode the key in source or commit it to git. Use an environment variable or a local `.env` file. ## The idea Keep everything identical except the country: ```json { "url": "https://www.nike.com/t/air-force-1-07-mens-shoes-jBrhbr/CW2288-111", "formats": ["markdown"], "render_js": true, "processing_mode": "sync", "proxy": { "country": "US", "type": "residential" } } ``` Then repeat the same request with `"country": "GB"`. That isolates geo-targeting as the only intentional variable. | Field | Why it is set | | ------------------------- | ---------------------------------------- | | `formats: ["markdown"]` | Easy to read and compare | | `render_js: true` | Nike's PDP is JS-heavy | | `processing_mode: sync` | Wait for the full result in one response | | `proxy.country` | Force US or GB exit | | `proxy.type: residential` | Better fit for consumer retail sites | ## Implementation ```python import json import os from pathlib import Path import requests API_KEY = os.environ["GEONODE_SCRAPER_API_KEY"] URL = "https://www.nike.com/t/air-force-1-07-mens-shoes-jBrhbr/CW2288-111" COUNTRIES = ["US", "GB"] OUT = Path("output") OUT.mkdir(exist_ok=True) def extract(country: str) -> dict: response = requests.post( "https://scraper.geonode.io/v1/extract", headers={ "X-Api-Key": API_KEY, "Content-Type": "application/json", }, json={ "url": URL, "formats": ["markdown"], "render_js": True, "processing_mode": "sync", "proxy": { "country": country, "type": "residential", }, }, timeout=300, ) response.raise_for_status() return response.json() results = {} for country in COUNTRIES: result = extract(country) markdown = result["data"]["markdown"] path = OUT / f"nike_af1_{country}.md" path.write_text(markdown, encoding="utf-8") results[country] = { "proxy": result["metadata"]["proxy"], "tokens_charged": result["tokens_charged"], "markdown_length": len(markdown), "has_uk_banner": "We think you are in United Kingdom" in markdown, "path": str(path), } print(country, results[country]) us = (OUT / "nike_af1_US.md").read_text(encoding="utf-8") gb = (OUT / "nike_af1_GB.md").read_text(encoding="utf-8") summary = { "target_url": URL, "markdown_differs": us != gb, "results": results, } (OUT / "geo_comparison.json").write_text( json.dumps(summary, indent=2), encoding="utf-8", ) print("markdown_differs:", summary["markdown_differs"]) ``` Run it: ```bash python geo_targeted_extract.py ``` You should get: * `output/nike_af1_US.md` * `output/nike_af1_GB.md` * `output/geo_comparison.json` ## What you get in the response A successful sync extract looks like this shape: ```json { "data": { "markdown": "# Nike Air Force 1 '07\n\n$115\n..." }, "metadata": { "url": "https://www.nike.com/t/air-force-1-07-mens-shoes-jBrhbr/CW2288-111", "render_js": true, "http_status": 200, "duration_ms": 9069, "formats": ["markdown"], "proxy": { "country": "US", "type": "residential" }, "processing_mode": "sync", "headers": {}, "wait_config": null }, "tokens_charged": 1 } ``` You mainly care about: | Field | Meaning | | ---------------------- | ------------------------------------ | | `data.markdown` | Clean page content for that country | | `metadata.proxy` | Country and proxy type actually used | | `metadata.http_status` | Status from the target site | | `metadata.duration_ms` | How long the extract took | | `tokens_charged` | Tokens billed for the request | You do **not** get structured product fields such as `name`, `price`, or `sku`. Those must be parsed from `data.markdown` if you need them. ## Captured US vs GB results From the live runs used for this example: | | US | GB | | ------------------------ | ------------- | ------------- | | HTTP status | `200` | `200` | | `tokens_charged` | `1` | `1` | | `metadata.proxy.country` | `US` | `GB` | | `metadata.proxy.type` | `residential` | `residential` | | `metadata.render_js` | `true` | `true` | | `metadata.duration_ms` | `9069` | `13864` | | Markdown length | `67817` | `42453` | | Markdown differs | yes | yes | Both returned the product: ```markdown # Nike Air Force 1 '07 $115 ``` The UK extract also included this location banner: ```markdown # We think you are in United Kingdom. Update your location? ``` That banner was not present in the captured US extract. Geo-targeting does not guarantee a currency change. In this captured run, both markets still showed `$115`. The useful proof is that the page reacted to exit country — here via Nike's location banner and different Markdown overall. ## How to judge success 1. **Proxy applied** — `metadata.proxy.country` matches what you requested 2. **Content extracted** — `data.markdown` contains the product page 3. **Meaningful difference** — the two Markdown files differ, or one country shows a clear market signal Do not judge only on price. Banners, market pickers, shipping text, and locale strings all count. ## When to use this pattern Use geo-targeted extract when location changes the page you care about: * Market localization QA * Competitor monitoring by region * Shipping / availability messaging checks * Detecting country-specific offers or banners If location does not matter, omit `proxy.country` and let the API use default routing. ## Next Once this pattern is working, natural follow-ons are: * Batch-extract a list of SKUs from one country * Map a storefront, filter product URLs, then batch-extract them * Combine `proxy.country` with custom headers such as `Accept-Language` # Cancel a Batch Job (/docs/scraper-api/v1/batch/cancel-batch-job) Stop scheduling new batch items and let in-flight child extractions drain # Get Batch Job Status (/docs/scraper-api/v1/batch/get-batch-job-status) Poll for the current status and partial results of a batch job # List batch jobs (/docs/scraper-api/v1/batch/list-batch-job) List batch jobs for the authenticated user with optional filtering and pagination # Start a Batch Job (/docs/scraper-api/v1/batch/start-batch-job) Queue a batch of URLs for asynchronous extraction # Cancel a Crawl Job (/docs/scraper-api/v1/crawl/cancel-crawl-job) Stop scheduling new crawl pages and let in-flight page work drain # Get Crawl Job Status (/docs/scraper-api/v1/crawl/get-crawl-job-status) Poll for the current status and results of a crawl job # List crawl jobs (/docs/scraper-api/v1/crawl/list-crawl-jobs) List crawl jobs for the authenticated user with optional filtering and pagination # Start a Crawl Job (/docs/scraper-api/v1/crawl/start-crawl-job) Crawl a website starting from a seed URL up to a given depth and page limit # Extract Content (/docs/scraper-api/v1/extraction/extract-content) Extract clean markdown or HTML from any URL with sync or async mode support # Get Extraction Job (/docs/scraper-api/v1/extraction/get-extraction-job) Poll for job progress and retrieve extraction results (async mode only) # List Extraction Jobs (/docs/scraper-api/v1/extraction/list-extraction-jobs) List and filter extraction jobs by job ID, URL, status, output format, and date range # Get map job details (/docs/scraper-api/v1/map/get-map-job) Retrieve the full details and discovered links for a completed map job # List map jobs (/docs/scraper-api/v1/map/list-map-jobs) List map jobs for the authenticated user with optional filtering and pagination. # Map URLs (/docs/scraper-api/v1/map/map-urls) Returns the list of URLs found under the given base URL by combining sitemap parsing with HTML link extraction from the seed page. The optional `search` parameter filters the discovered URLs by case-insensitive substring match — it does NOT query a search engine. # Get search job details (/docs/scraper-api/v1/search/get-search-job) Retrieve the full details and results for a completed search job # List search jobs (/docs/scraper-api/v1/search/list-search-jobs) List search jobs for the authenticated user with optional filtering and pagination {/* This file was generated by Fumadocs. Do not edit this file directly. Any changes should be made by running the generation command again. */} # Search (/docs/scraper-api/v1/search/start-search-job) Start a new search job {/* This file was generated by Fumadocs. Do not edit this file directly. Any changes should be made by running the generation command again. */ } {/* This file was generated by Fumadocs. Do not edit this file directly. Any changes should be made by running the generation command again. */} # Get Statistics (/docs/scraper-api/v1/statistics/get-statistics) This endpoint is documented but not yet available in production. The contract below reflects the planned behavior. Reach out via [support](https://geonode.com/contact) for early access or launch notification. Retrieve aggregated extraction statistics for a given date range # Health Check (/docs/scraper-api/v1/system/health-check) {/* This file was generated by Fumadocs. Do not edit this file directly. Any changes should be made by running the generation command again. */} # Get Concurrency Usage (/docs/scraper-api/v1/usage/get-concurrency-usage) Return the caller's live work-concurrency slot usage against their plan limit. # Create a Webhook (/docs/scraper-api/v1/webhooks/create-webhook) This endpoint is documented but not yet available in production. The contract below reflects the planned behavior. Reach out via [support](https://geonode.com/contact) for early access or launch notification. Create a webhook subscription for a given event type and return its generated signing secret # Delete a Webhook (/docs/scraper-api/v1/webhooks/delete-webhook) This endpoint is documented but not yet available in production. The contract below reflects the planned behavior. Reach out via [support](https://geonode.com/contact) for early access or launch notification. Permanently remove a webhook subscription # Get a Webhook (/docs/scraper-api/v1/webhooks/get-webhook) This endpoint is documented but not yet available in production. The contract below reflects the planned behavior. Reach out via [support](https://geonode.com/contact) for early access or launch notification. Retrieve a single webhook by its identifier # List Webhook Deliveries (/docs/scraper-api/v1/webhooks/list-webhook-deliveries) This endpoint is documented but not yet available in production. The contract below reflects the planned behavior. Reach out via [support](https://geonode.com/contact) for early access or launch notification. List webhook delivery attempts with optional status filter and pagination # List Webhooks (/docs/scraper-api/v1/webhooks/list-webhooks) This endpoint is documented but not yet available in production. The contract below reflects the planned behavior. Reach out via [support](https://geonode.com/contact) for early access or launch notification. List webhooks registered for the current user with pagination # Rotate Webhook Secret (/docs/scraper-api/v1/webhooks/rotate-webhook-secret) This endpoint is documented but not yet available in production. The contract below reflects the planned behavior. Reach out via [support](https://geonode.com/contact) for early access or launch notification. Generate a new signing secret for the webhook and invalidate the previous one # Update a Webhook (/docs/scraper-api/v1/webhooks/update-webhook) This endpoint is documented but not yet available in production. The contract below reflects the planned behavior. Reach out via [support](https://geonode.com/contact) for early access or launch notification. Partially update a webhook's url, description, event type, or active flag # Perform Bandwidth-Limited Proxy Session (/docs/proxies/api-reference/sticky-session/get-sticky-session/get-bandwidth-limited) Create a proxy session with bandwidth limitations to control the data transfer rate for your connection. This API requires Basic Authentication. Include the following in your request: * **Username**: `-session--limit-` * **Password**: `` * **Proxy**: `http://proxy.geonode.io:` * **Accept Header**: `application/json` ## Request ```bash curl --request GET \ -x "http://proxy.geonode.io:" \ --user "-session--limit-:" \ --url "http://ip-api.com/json" ``` ## Response ### 200 Success Bandwidth-limited proxy session executed successfully. #### Response Fields | Field | Type | Description | | ------------- | ------ | ---------------------------------------------------- | | `status` | string | The status of the request (e.g., "success") | | `country` | string | The full name of the country where the IP is located | | `countryCode` | string | The two-letter country code | | `region` | string | The region code | | `regionName` | string | The full name of the region | | `city` | string | The name of the city | | `zip` | string | The postal code associated with the IP | | `lat` | number | The latitude coordinate | | `lon` | number | The longitude coordinate | | `timezone` | string | The timezone of the IP location | | `isp` | string | The name of the Internet Service Provider | | `org` | string | The organization that owns the IP address | | `as` | string | The Autonomous System (AS) number | | `query` | string | The queried IP address | #### Example Response ```json { "status": "success", "country": "Algeria", "countryCode": "DZ", "region": "22", "regionName": "Sidi Bel Abbès", "city": "Sidi Bel Abbes", "zip": "22000", "lat": 34.8934, "lon": -0.6526, "timezone": "Africa/Algiers", "isp": "4 djaweb de AS fawri", "org": "", "as": "AS36947 Telecom Algeria", "query": "197.203.245.147" } ``` # Create a New Sticky Session (/docs/proxies/api-reference/sticky-session/get-sticky-session/get-create) Create a new sticky session using proxy credentials. Sticky sessions let you keep the same IP address longer, which is useful for tasks that need a stable connection, like managing social media accounts or web scraping. Knowing how to use sticky sessions can improve your proxy experience by making it more efficient and reliable. Geonode provides specific ports for sticky sessions, ensuring a persistent IP session: | Protocol | Port Range | | -------------- | ------------- | | **HTTP/HTTPS** | 10000 - 10900 | | **SOCKS5** | 12000 - 12010 | Use sticky ports when you need a stable IP for session-based activities. ## Request Make a request using a sticky session by including the session ID in your username string: ```bash curl -X "http://proxy.geonode.io:" \ --user "-country--session-:" \ --url "http://ip-api.com/json" \ --header "Accept: application/json" ``` ### Parameters The session ID can be any custom string (1-25 alphanumeric characters or underscores). As long as you keep sending requests with the same session ID, you will maintain the same proxy IP. ## Response ### 200 Success The response contains detailed geolocation information about the IP address assigned to your sticky session. #### Response Fields | Field | Type | Description | | --------------- | ------- | --------------------------------------------------------- | | `status` | string | The status of the request (e.g., "success") | | `continent` | string | The continent where the IP is located | | `continentCode` | string | The two-letter continent code | | `country` | string | The full name of the country | | `countryCode` | string | The two-letter country code (ISO 3166-1 alpha-2) | | `region` | string | The region code | | `regionName` | string | The full name of the region | | `city` | string | The name of the city | | `district` | string | The district name, if available | | `zip` | string | The postal code associated with the IP | | `lat` | number | The latitude coordinate | | `lon` | number | The longitude coordinate | | `timezone` | string | The timezone of the IP location | | `offset` | integer | The time offset in seconds from UTC | | `currency` | string | The currency code of the country | | `isp` | string | The name of the Internet Service Provider | | `org` | string | The organization that owns the IP address | | `as` | string | The Autonomous System (AS) number | | `asname` | string | The name associated with the AS number | | `mobile` | boolean | Indicates whether the connection is from a mobile network | | `proxy` | boolean | Indicates whether the IP is a known proxy | | `hosting` | boolean | Indicates whether the IP is from a hosting provider | | `query` | string | The queried IP address | #### Example Response ```json { "status": "success", "continent": "Africa", "continentCode": "AF", "country": "Algeria", "countryCode": "DZ", "region": "05", "regionName": "Batna", "city": "Batna City", "district": "", "zip": "05000", "lat": 35.5064, "lon": 6.0707, "timezone": "Africa/Algiers", "offset": 3600, "currency": "DZD", "isp": "4 djaweb de AS fawri", "org": "", "as": "AS36947 Telecom Algeria", "asname": "ALGTEL-AS", "mobile": false, "proxy": false, "hosting": false, "query": "XXX.XXX.XXX.XXX" } ``` # List All Active Sessions (/docs/proxies/api-reference/sticky-session/get-sticky-session/get-list-active) Retrieve a paginated list of all currently active proxy sessions for your account. This endpoint helps you monitor and manage your active connections. ## Request ```bash curl -X GET "https://app-api.geonode.com/api/sessions/proxies" \ -H "Authorization: Basic base64(username:password)" ``` ### Query Parameters | Parameter | Type | Required | Default | Description | | ---------- | ------- | -------- | ------- | ------------------------------------- | | `page` | integer | No | 1 | Page number for pagination | | `pageSize` | integer | No | 250 | Number of sessions per page (max 250) | ## Response ### 200 Success A list of active sessions retrieved successfully. #### Response Fields | Field | Type | Description | | -------------------------------------- | ------- | ----------------------------------------------------------- | | `sessions` | array | A list of currently active proxy sessions | | `sessions[].id` | string | A unique identifier for the session | | `sessions[].userSessionId` | string | The user-defined session identifier | | `sessions[].userId` | string | The Geonode user ID associated with the session | | `sessions[].port` | string | The port number assigned to the session | | `sessions[].domain` | string | The target domain being accessed in the session | | `sessions[].country` | string | The country code representing the location | | `sessions[].rotatingIntervalInSeconds` | number | The interval (in seconds) at which the session rotates | | `sessions[].durationInSeconds` | number | The total duration (in seconds) the session has been active | | `count` | integer | The number of active sessions returned in the current page | | `total` | integer | The total number of active sessions across all pages | | `page` | integer | The current page number | | `pageSize` | integer | The number of sessions per page | #### Example Response ```json { "sessions": [ { "id": "a7f7c7a3-aaaa-4908-b7bd-71091daadaa6", "userSessionId": "vzzzzd", "userId": "geonode_userid", "port": "10000", "domain": "ip-api.com", "country": "dz", "rotatingIntervalInSeconds": 179.32, "durationInSeconds": 2.785 } ], "count": 1, "total": 1, "page": 1, "pageSize": 250 } ``` # Configuring Proxy Session ID & Lifetime (/docs/proxies/api-reference/sticky-session/get-sticky-session/get-proxy-session) Control the duration of your proxy sessions by setting a custom session ID and lifetime. This allows you to maintain stable connections for extended periods, which is essential for tasks requiring persistent IP addresses. The `lifetime` parameter is measured in **minutes** and defines how long the session remains active: | Constraint | Value | | ----------- | ----------------------- | | **Minimum** | 3 minutes | | **Maximum** | 1440 minutes (24 hours) | | **Default** | 10 minutes | **Examples:** * `lifetime: 5` → 5 minutes * `lifetime: 30` → 30 minutes * `lifetime: 60` → 1 hour * `lifetime: 180` → 3 hours Your custom session ID must follow these rules: * **Format**: 1-25 alphanumeric characters or underscores (`_`) * **Regex pattern**: `^[A-Za-z0-9_]{1,25}$` * **Valid characters**: Letters (A-Z, a-z), digits (0-9), and underscores * **No spaces or special symbols** allowed **Examples of valid session IDs:** * `session123` * `my_session_01` * `ABC123` ## Request ```bash curl --request GET \ -x "http://proxy.geonode.io:" \ --user "-session--lifetime-:" \ --url "http://ip-api.com/json" ``` ## Response ### 200 Success Proxy session created successfully with the specified lifetime. #### Response Fields | Field | Type | Description | | --------------- | ------- | --------------------------------------------------------- | | `status` | string | The status of the request (e.g., "success") | | `continent` | string | The continent where the IP is located | | `continentCode` | string | The two-letter continent code | | `country` | string | The full name of the country | | `countryCode` | string | The two-letter country code | | `region` | string | The region code | | `regionName` | string | The full name of the region | | `city` | string | The name of the city | | `district` | string | The district name, if available | | `zip` | string | The postal code associated with the IP | | `lat` | number | The latitude coordinate | | `lon` | number | The longitude coordinate | | `timezone` | string | The timezone of the IP location | | `offset` | integer | The time offset in seconds from UTC | | `currency` | string | The currency code of the country | | `isp` | string | The name of the Internet Service Provider | | `org` | string | The organization that owns the IP address | | `as` | string | The Autonomous System (AS) number | | `asname` | string | The name associated with the AS number | | `mobile` | boolean | Indicates whether the connection is from a mobile network | | `proxy` | boolean | Indicates whether the IP is a known proxy | | `hosting` | boolean | Indicates whether the IP is from a hosting provider | | `query` | string | The queried IP address | #### Example Response ```json { "status": "success", "continent": "Africa", "continentCode": "AF", "country": "Algeria", "countryCode": "DZ", "region": "05", "regionName": "Batna", "city": "Batna City", "district": "", "zip": "05000", "lat": 35.5064, "lon": 6.0707, "timezone": "Africa/Algiers", "offset": 3600, "currency": "DZD", "isp": "4 djaweb de AS fawri", "org": "", "as": "AS36947 Telecom Algeria", "asname": "ALGTEL-AS", "mobile": false, "proxy": false, "hosting": false, "query": "192.168.1.1" } ``` # Release Sticky Session by Session ID and Port (/docs/proxies/api-reference/sticky-session/release-sticky-session/put-release-by-session-id-and-port) Release one or more sticky sessions by specifying their session IDs and ports. This allows you to free up resources and terminate specific proxy connections. ## Request ```bash curl -X PUT "https://app-api.geonode.com/api/sessions/release/proxies" \ -H "Authorization: Basic base64(username:password)" \ -H "Content-Type: application/json" \ -d '{"data":[{"sessionId":"random0001","port":10000}]}' ``` ### Request Body Parameters | Field | Type | Required | Description | | ------------------ | ------- | -------- | ------------------------------------------- | | `data` | array | Yes | Array containing session details | | `data[].sessionId` | string | Yes | The session ID to release | | `data[].port` | integer | Yes | The port number associated with the session | ### Example Request Body ```json { "data": [ { "sessionId": "random0001", "port": 10000 } ] } ``` ## Response ### 200 Success Session released successfully. #### Response Fields | Field | Type | Description | | --------- | ------- | ---------------------------------------------------- | | `success` | boolean | Indicates whether the session release was successful | #### Example Response ```json { "success": true } ``` # Release Proxy Session by Multiple Port (/docs/proxies/api-reference/sticky-session/release-sticky-session/put-release-proxy-session-by-multiple-port) Release sticky sessions running on multiple ports in a single request. This is useful when you need to free several proxy connections at once. ## Request ```bash curl -X PUT "https://monitor.geonode.com/sessions/release/proxies" \ -H "Authorization: Basic base64(username:password)" \ -H "Content-Type: application/json" \ -d '{"data":[{"port":10000},{"port":10001},{"port":10002}]}' ``` ### Request Body Parameters | Field | Type | Required | Description | | ------------- | ------- | -------- | -------------------------------------------- | | `data` | array | Yes | Array containing session details | | `data[].port` | integer | Yes | Port number of each proxy session to release | ### Example Request Body ```json { "data": [{ "port": 10000 }, { "port": 10001 }, { "port": 10002 }] } ``` ## Response ### 200 Success Successfully released the specified ports. #### Response Fields | Field | Type | Description | | --------- | ------- | ---------------------------------------------------- | | `success` | boolean | Indicates whether the session release was successful | #### Example Response ```json { "success": true } ``` # Release Proxy Session by Port (/docs/proxies/api-reference/sticky-session/release-sticky-session/put-release-proxy-session-by-port) Release a sticky session by specifying its port number. Use this when you need to free up a particular proxy connection without affecting other sessions. ## Request ```bash curl -X PUT "https://monitor.geonode.com/sessions/release/proxies" \ -H "Authorization: Basic base64(username:password)" \ -H "Content-Type: application/json" \ -d '{"data":[{"port":10001}]}' ``` ### Request Body Parameters | Field | Type | Required | Description | | ------------- | ------- | -------- | ----------------------------------------------- | | `data` | array | Yes | Array containing session details | | `data[].port` | integer | Yes | The port number of the proxy session to release | ### Example Request Body ```json { "data": [ { "port": 10001 } ] } ``` ## Response ### 200 Success Session released successfully. #### Response Fields | Field | Type | Description | | --------- | ------- | ---------------------------------------------------- | | `success` | boolean | Indicates whether the session release was successful | #### Example Response ```json { "success": true } ``` # Overview (/docs/proxies/api-reference/sticky-session/release-sticky-session/release-sticky-session) Release sticky sessions for specified Geonode services by session ID and port. This allows you to free up resources and terminate specific proxy connections. ## What is Session Release? Session release is the process of terminating active sticky sessions before they naturally expire. When you release a session, you're telling the Geonode proxy system to immediately free up the IP address and resources associated with that session. ## Why Release Sessions? Release sessions when you need to free up resources or terminate specific proxy connections before they naturally expire. ## How It Works To release a sticky session, you need to provide: * **Session ID**: The unique identifier for the session you want to release * **Port**: The port number associated with the session The system will immediately terminate the session and free up the associated resources. Once released, the IP address becomes available for other users, and you'll need to create a new session if you want to continue using a sticky connection. ## Available Operations This section provides endpoints for managing sticky session releases: * **[Release Sticky Session by Session ID and Port](/docs/proxies/api-reference/sticky-session/release-sticky-session/put-release-by-session-id-and-port)**: Release one or more specific sessions by providing their session IDs and ports Before releasing sessions, you can view all your active sessions using the [List All Active Sessions](/docs/proxies/api-reference/sticky-session/get-sticky-session/get-list-active) endpoint to identify which sessions you want to release. ## Best Practices Before releasing sessions, you can view all your active sessions to identify which ones you want to release. Once a session is released, it cannot be restored, so make sure you're ready to release it before executing the operation. Once a session is released, it cannot be restored. You'll need to create a new session if you want to continue using a sticky connection. Make sure you're ready to release a session before executing the release operation. # Port Configuration (/docs/proxies/getting-started/setup_and_configuration/advance-configuration/port-configuration) import SupportParagraph from "../../../../../snippets/support-paragraph.mdx"; The **Port Configuration** section in your Geonode dashboard lets you set up and manage proxy ports with precise geo-targeting and connection options. *** ## Step 1 — Open the Port Configuration page 1. Log in to your Geonode dashboard. 2. Select **Port Configuration** from the top navigation bar. Navigate to Port Configuration You’ll see configuration options on the right and a list of added locations in the center. *** ## Step 2 — Configure your port settings ### 1. Country targeting (required) * Open the **Countries** dropdown. * Select the country you want to target — this field is required for geo-targeting. Country Targeting ### 2. State targeting (optional) * Once a country is selected, the **States** dropdown becomes active. * Choose a specific state if needed. State Targeting ### 3. City targeting (optional) * After choosing a state, the **Cities** dropdown will be enabled. * Select a city for more precise targeting. ### 4. Protocol selection * Click the **HTTP/HTTPS** dropdown. * Choose a protocol: * **HTTP/HTTPS** — standard web traffic. * **SOCKS5** — advanced, more flexible option (if available). Protocol Selection → [Understanding Protocol Type for Proxy Configuration](/docs/proxies/getting-started/knowledge-base/protocol-type) ### 5. Session type * Select a session mode: * **Rotating** — IP changes periodically (best for scraping/automation). * **Sticky** — one IP per session (best for login or persistent tasks). Session Type → [Understanding Session Type for Proxy Configuration](/docs/proxies/getting-started/knowledge-base/session-type) ### 6. Port selection (required) * Click **Select Port** to choose from available ports. * This field is required to complete setup. Port Selection (Required) → [Understanding Port Usage for Proxy Configuration](/docs/proxies/getting-started/knowledge-base/proxy-usage) *** ## Step 3 — Add the configuration When all required fields are filled, the **Add** button becomes active.\ Click **Add** to save your configuration. Add the Configuration *** ## Step 4 — Review added locations After saving, all configured ports and locations appear in the list.\ You can view, edit, or delete configurations at any time. Review Added Locations *** ## Port in use When a port is assigned to a specific country, it can’t be reused elsewhere. Port in Use *** ## Final result You’re now ready to manage ports efficiently in Geonode. * **Rotating Example** For Rotating * **Sticky Example** For Sticky *** ## Troubleshooting * **Add button disabled:** make sure Country and Port are selected. * **State/City inactive:** these fields unlock only after a country is chosen. * **Connection issues:** verify your protocol and session settings. *** *** ## FAQs The button stays inactive until all required fields are selected. Choose both a country and a port before proceeding. {" "} No. Each port can only be assigned to one country. {" "} Rotating sessions periodically change your IP, ideal for scraping and automation. {" "} Yes, you can select either protocol during configuration. SOCKS5 support depends on your plan. State and city fields are only active after selecting a country. # Proxy Authorization in API Using Headers (/docs/proxies/getting-started/setup_and_configuration/advance-configuration/proxy-authentication-in-api) import SupportParagraph from "../../../../../snippets/support-paragraph.mdx"; *** API authorization ensures secure access to Geonode’s proxy services.\ This guide explains how to authenticate requests using the `Authorization` header. *** ## Step 1 — Get your API credentials Before setting up authentication, make sure you have your Geonode API credentials. → [How to access your Geonode API credentials](/docs/proxies/getting-started/prerequisites/access-credentials) *** ## Step 2 — Generate the Authorization header Geonode’s API uses **Basic Authentication**, which requires your username and password encoded in **Base64**. 1. Open the API documentation for the endpoint you want to test (for example, *Retrieve Usage Statistics*). 2. Click **Try it** on the API page. Try it 3. A popup will appear asking for your **username** and **password**. Enter them. 4. The system automatically generates the `Authorization` header for you — containing the Base64-encoded string. Base Auth 5. Copy the generated header or the encoded string.\ You can use it directly in your API requests or store the Base64 value in your code securely. *** ## Step 3 — Follow best practices * **Generate once, reuse:** Create your token once and reuse it for multiple requests. * **Store securely:** Keep it in a `.env` file or secret manager. * **Avoid hardcoding:** Never paste credentials directly into your source code. * **Always use HTTPS:** This encrypts your traffic and protects sensitive data. * **Rotate regularly:** Update your credentials periodically.\ When you do, generate a new token. *** ## Troubleshooting * Check that your header is formatted correctly: * Encode exactly `username:password` — no extra spaces or characters. * Use HTTPS in all requests; avoid using insecure HTTP. * Test your setup with tools like **Postman** or **cURL** to confirm it works. * If your credentials are compromised, **change your password**, regenerate your API key, and update the token. *** *** ## FAQs No. The Geonode API documentation automatically generates the Base64-encoded string for you. Not at the moment — all proxy API requests require Basic Authorization. Currently, Geonode only supports Basic Authorization. Check future API updates for new methods. No. Base64 is not encryption — it’s only encoding. Always use HTTPS to keep credentials safe. Rotate them regularly or immediately if you suspect compromise. # Proxy Configuration (/docs/proxies/getting-started/setup_and_configuration/advance-configuration/proxy-configuration) *** ## What is the Endpoint Generator The Endpoint Generator helps you create customized proxy lists based on your preferences.\ You can select **IP type**, **host**, **location**, **protocol**, and **session type**, then export endpoints in `.txt` format. *** ## Step 1 — Access the Dashboard 1. Log in to your Geonode dashboard. 2. Scroll down to **Proxy Configuration**. Accessing the User Dashboard *** ## Step 2 — Choose the Endpoint Format There are **six formats** available for generating endpoints, depending on your authentication and connection preferences. Choosing an Endpoint Format → [Understanding Different Proxy Endpoint Formats and Their Uses](/docs/proxies/getting-started/knowledge-base/endpoint-formats) *** ## Step 3 — Set the Endpoint Count Specify how many endpoints you want to generate. Choosing an Endpoint Count *** ## Step 4 — Configure Proxy Parameters There are **seven options** available for customizing proxy endpoints. *** ### 1. IP Type Select the type of IPs you need.\ Three types are available: * **Residential** * **Datacenter** * **Mixed** Selecting IP Type → [Learn about Residential, Datacenter, and Mixed IPs and their best use cases](/docs/proxies/getting-started/knowledge-base/ip-type) *** ### 2. Gateway (Host) Choose the appropriate host for your location.\ Three options are available: | Location | Host Address | | ------------- | ------------------------------------ | | France | `proxy.geonode.io` | | United States | `us.premium-residential.geonode.com` | | Singapore | `sg.premium-residential.geonode.com` | Selecting a host automatically updates the corresponding address. Selecting Gateway → [Learn about Geonode’s proxy gateways and how they work](/docs/proxies/getting-started/knowledge-base/gateway) *** ### 3. Geo-Targeting (Country, State, City) Select your desired **country**, **state**, or **city**.\ By default, the system uses **Any**, meaning proxies can come from any location. **Example:** * **Host:** Singapore * **Target Country:** China This setup provides faster connectivity to nearby Chinese servers. Selecting Target Location → [What is Geo-Targeting](/docs/proxies/getting-started/knowledge-base/geo-targeting) *** ### 4. Protocol Choose how your proxy handles connections: 1. **HTTP/HTTPS** — standard web traffic. 2. **SOCKS5** — flexible and secure (if supported). Selecting Protocol → [Understanding Protocol Type for Proxy Configuration](/docs/proxies/getting-started/knowledge-base/protocol-type) When you change the protocol, port numbers in generated endpoints update automatically. → [Learn how ports work in Geonode’s proxy configuration](/docs/proxies/getting-started/knowledge-base/proxy-usage) *** ### 5. Session Type Define how IPs are handled within a session: 1. **Rotating Session** — IP changes periodically (best for scraping and automation). 2. **Sticky Session** — IP stays the same for the entire session (best for logins or persistent tasks). Selecting Session Type → [A Guide to Rotating and Sticky Sessions](/docs/proxies/getting-started/knowledge-base/session-type) *** ### 6. Rotating Interval (Sticky Sessions Only) If you use Sticky Sessions, you can set a rotation interval to control how often IPs refresh — in minutes or hours. Rotating Interval This keeps your connection stable while maintaining periodic IP rotation for security and reliability. *** ## Step 5 — Select the Output Format After completing the configuration, you can: * **Copy** endpoints directly for immediate use, or * **Download** them as a `.txt` file for later. Select the Output format *** ✅ **You’re all set!**\ You’ve successfully created a ready-to-use list of proxy endpoints.\ Your setup is now fully configured and ready to integrate into your applications. # Whitelist IP (/docs/proxies/getting-started/setup_and_configuration/advance-configuration/whitelist-ip) import VerifyProxyConnectionComponent from "../../../../../snippets/verify-proxy-connection-component.mdx"; import SupportParagraph from "../../../../../snippets/support-paragraph.mdx"; Whitelisting your IP lets you access proxy services without re-entering credentials each time. ## Key limits and requirements * Authentication: Basic Auth (username:password encoded in Base64). * IP limit: up to 150 whitelisted IPs per user. * Batch size: up to 10 IPs per request (add/remove). * Validation: the API rejects invalid or duplicate IPs. * Rate limit: up to 100 requests per minute. ## Add an IP to the whitelist You can use either the dashboard (recommended for non-developers) or the API. ### Using the dashboard #### Add IP addresses 1. Open **Whitelisted IPs** in settings. Navigate to IP Whitelist Settings 2. Enter the IP to whitelist. For a quick setup, click **Detect My IP**.\ Optionally add a description (e.g., Home, Office, Server). Click **Add**. Add an IP Address 3. Verify that the IP appears in the table with correct details. Verify the IP #### Remove an IP 1. Find the IP in the whitelist table. 2. Click **Delete** next to the IP. 3. Confirm the deletion and check the notification. Verify Deletion ### Using the API Available endpoints: * [Retrieve whitelisted IPs](/docs/proxies/api-reference/whitelisting-ip/get) `GET` * [Add whitelisted IPs](/docs/proxies/api-reference/whitelisting-ip/post) `POST` * [Update IP description](/docs/proxies/api-reference/whitelisting-ip/put) `PUT` * [Remove whitelisted IPs](/docs/proxies/api-reference/whitelisting-ip/delete) `DELETE` For request/response formats and examples, see the Geonode API docs:\ [Geonode API Documentation](/docs/proxies/api-reference/whitelisting-ip/whitelisting-ips) *** *** ## FAQs Basic Authentication ensures only authorized users can modify the whitelist. Encode your username and password in Base64 and include it in the request. Yes. Use the Geonode API to add, update, and remove IPs programmatically. Requests that push the total over 150 IPs are rejected. Remove unused entries first. Delete the incorrect IP from the dashboard or via the API, then add the correct one. Whitelisting lets trusted devices connect without re-entering credentials, simplifying access for known locations. *** # Selenium (/docs/proxies/getting-started/setup_and_configuration/automation-frameworks/selenium) import ProxyInfoFromGeonode from "../../../../../snippets/get-proxy-info-from-geonode.mdx"; import VerifyProxyConnectionComponent from "../../../../../snippets/verify-proxy-connection-component.mdx"; import SupportParagraph from "../../../../../snippets/support-paragraph.mdx"; *** This guide will help you integrate the Geonode API with Selenium to manage proxies effectively while running headless browsers. **We will be using Python for this guide.** *** *** ## Prerequisites * Python installed on your system * Geonode API credentials (username and password) * ChromeDriver installed (compatible with your Chrome version) ## Steps: Setting Up a Proxy in Brave Follow these steps to configure a proxy in Selenium: ### Step 1: Set Up a Virtual Environment Creating a virtual environment helps isolate dependencies and avoid conflicts. ``` # Install virtualenv if not already installed pip install virtualenv # Create a virtual environment python -m venv .venv # Activate the virtual environment ## On Windows .venv\Scripts\activate ## On macOS/Linux source .venv/bin/activate ``` ### Step 2: Install Required Libraries ``` pip install selenium python-dotenv selenium-wire ``` * **Selenium:** For browser automation * **python-dotenv:** For managing environment variables * **selenium-wire:** To handle proxy authentication, as Selenium doesn't provide it natively ### Step 3: Configure Geonode Proxy Endpoint With the help of the **Endpoint Generator**, you can easily generate the proxy with specific configurations such as: * Target country * Port * Session persistence * And many more endpoint generator Refer to the guide **[How to Use the Endpoint Generator](/docs/proxies/getting-started/knowledge-base/geo-targeting)** to generate your endpoints. ### Step 4: Setting up environment variables 1. Create a `.env` file in your project directory. 2. Add your credentials: ``` GEONODE_USERNAME=your_geonode_username GEONODE_PASSWORD=your_geonode_password GEONODE_HOST=proxy.geonode.io GEONODE_PORT=9000 GEONODE_DNS=your_geonode_dns ``` Never upload your `.env` file to the internet. ## Code Implementation ### I. Import ``` import os from dotenv import load_dotenv from seleniumwire import webdriver from selenium.webdriver.chrome.options import Options ``` ### ii. Load environment variables ``` load_dotenv() proxy_host = os.getenv('GEONODE_HOST') proxy_port = os.getenv('GEONODE_PORT') username = os.getenv('GEONODE_USERNAME') password = os.getenv('GEONODE_PASSWORD') GEONODE_DNS = os.getenv('GEONODE_PROXY') ``` ### iii. Configure Proxy options ``` proxy_options = { 'proxy': { 'http': f'http://{username}:{password}@{proxy_host}:{proxy_port}', 'https': f'https://{username}:{password}@{proxy_host}:{proxy_port}', } } ``` ### iv. Configure Browser Options ``` chrome_options = Options() chrome_options.add_argument('--disable-gpu') chrome_options.add_argument('--start-maximized') chrome_options.add_argument('--ignore-certificate-errors') ``` ### v. Initialize Browser ``` browser = webdriver.Chrome(seleniumwire_options=proxy_options, options=chrome_options) ``` ### vi. Open IP address website to check ``` urlToGet = "https://ip-api.com/" browser.get(urlToGet) ``` **Output:** selenium-ip-address ### vii. Keep the browser open ``` input("Press Enter to close the browser...") ``` ### viii. Quit the browser ``` browser.quit() ``` ### xi. Full Code ``` import os from dotenv import load_dotenv from seleniumwire import webdriver from selenium.webdriver.chrome.options import Options load_dotenv() proxy_host = os.getenv('GEONODE_HOST') proxy_port = os.getenv('GEONODE_PORT') username = os.getenv('GEONODE_USERNAME') password = os.getenv('GEONODE_PASSWORD') GEONODE_DNS = os.getenv('GEONODE_PROXY') proxy_options = { 'proxy': { 'http': f'http://{username}:{password}@{proxy_host}:{proxy_port}', 'https': f'https://{username}:{password}@{proxy_host}:{proxy_port}', } } chrome_options = Options() chrome_options.add_argument('--disable-gpu') chrome_options.add_argument('--start-maximized') chrome_options.add_argument('--ignore-certificate-errors') browser = webdriver.Chrome(seleniumwire_options=proxy_options, options=chrome_options) urlToGet = "https://ip-api.com/" browser.get(urlToGet) input("Press Enter to close the browser...") browser.quit() ``` ### x. Folder Stucture ``` Project Root ├── .venv/ ├── .env └── app.py ``` ## Source Code You can find the full source code for this script on GitHub at the following link: [https://github.com/geonodecom/proxy-testing-toolkit/tree/automation-framework/selenium](https://github.com/geonodecom/proxy-testing-toolkit/tree/automation-framework/selenium) *** ## Use Cases of Integrating Geonode with Selenium Integrating Geonode with Selenium can be beneficial for a wide range of applications, including: 1. **Web Scraping:** Collect data from websites while maintaining anonymity to avoid IP bans. 2. **Ad Verification:** Test and verify advertisements across different geographies to ensure proper delivery. 3. **Price Monitoring:** Track pricing changes on e-commerce platforms without getting blocked. 4. **SEO Monitoring:** Monitor search engine results and competitor websites without affecting personalized search results. 5. **Market Research:** Gather data from various sources to analyze trends and competitor performance. 6. **Social Media Automation:** Manage multiple social media accounts while avoiding detection. 7. **Fraud Detection:** Simulate real-world traffic for security testing and fraud detection systems. Explore more use cases here *[https://geonode.com/use-cases](https://geonode.com/use-cases)* *** ## Troubleshooting Tips * **Timeout Errors:** Check if the proxy is active or switch to another proxy. * **Authentication Issues:** Double-check your Geonode API credentials. * **Incompatible ChromeDriver:** Ensure ChromeDriver matches your Chrome version. *** ## FAQs The authentication details are embedded in the proxy URL, so no pop-up should appear. {" "} Yes, but you need to adjust the proxy settings using Firefox profiles. Verify the proxy server details and ensure Geonode proxies are correctly configured. # macOS (/docs/proxies/getting-started/setup_and_configuration/desktop-OS/macOS) import BrowsersFaqs from "../../../../../snippets/browsers-faqs.mdx"; import ProxyInfoFromGeonode from "../../../../../snippets/get-proxy-info-from-geonode.mdx"; import VerifyProxyConnectionComponent from "../../../../../snippets/verify-proxy-connection-component.mdx"; import SupportParagraph from "../../../../../snippets/support-paragraph.mdx"; import PopupAuth from "../../../../../snippets/popup-auth.mdx"; *** *** ## Step-by-step setup for macOS Follow these steps to configure your proxy manually. ### Step 1 — Open System Preferences 1. Click the **Apple menu** icon in the top-left corner. 2. Select **System Preferences** from the dropdown. System Settings *** ### Step 2 — Open Network Settings 1. In **System Preferences**, click **Network**. 2. Choose your active Wi-Fi network and click **Details** (or the “i” icon). Other Networks *** ### Step 3 — Select Your Active Network Make sure the selected network is the one you’re currently connected to. *** ### Step 4 — Open the Proxies Tab 1. Click **Advanced** in the Network window. 2. Open the **Proxies** tab. Configure Network *** ### Step 5 — Configure Proxy Settings Add Details 1. In the **Proxies** tab, you’ll see several proxy types: * Web Proxy (HTTP) * Secure Web Proxy (HTTPS) * SOCKS Proxy 2. Check the box next to the proxy type you’re setting up (for example, **Web Proxy (HTTP)**). 3. Enter your **Proxy Server** address and **Port** — both can be copied from your Geonode Dashboard (click the copy icon next to *Host*). *** ### Step 6 — Apply and Save Changes 1. Click **OK** to close Advanced settings. 2. Click **Apply** in the main Network window to save your configuration. *** *** *** # Windows (/docs/proxies/getting-started/setup_and_configuration/desktop-OS/windows) import BrowsersFaqs from "../../../../../snippets/browsers-faqs.mdx"; import ProxyInfoFromGeonode from "../../../../../snippets/get-proxy-info-from-geonode.mdx"; import VerifyProxyConnectionComponent from "../../../../../snippets/verify-proxy-connection-component.mdx"; import SupportParagraph from "../../../../../snippets/support-paragraph.mdx"; import PopupAuth from "../../../../../snippets/popup-auth.mdx"; *** *** ## Step-by-step setup for Windows 10/11 Follow these steps to configure your proxy manually: ### Step 1 — Open Proxy Settings 1. Click the **Windows Search Bar** and type:\ `"Change proxy settings"` Open Proxy Settings in Windows 2. In the settings window, scroll to **Manual Proxy Setup**. Manual Proxy Setup *** ### Step 2 — Enter Proxy Details 1. A configuration window will open: Enter Proxy Details 2. Enter the **Proxy IP** and **Port** you copied from your Geonode Dashboard. Enter the Proxy IP and Port 3. Click **Save** to apply the settings. *** *** *** # AdsPower (/docs/proxies/getting-started/setup_and_configuration/browsers/adspower) import { Steps } from "fumadocs-ui/components/steps"; import ProxyInfoFromGeonode from "../../../../../snippets/get-proxy-info-from-geonode.mdx"; import SupportParagraph from "../../../../../snippets/support-paragraph.mdx"; import BrowsersFaqs from "../../../../../snippets/browsers-faqs.mdx"; This guide explains how to configure a Geonode proxy in the **AdsPower** browser. ## Setting Up a Proxy in AdsPower ### Step 1: Install AdsPower 1. Go to the [AdsPower Download Page](https://www.adspower.com/download). 2. Download the version for your operating system. 3. Install AdsPower and create an account. *** ### Step 2: Create a New Profile 1. Open AdsPower and click **New Profile**.\ New Profile Button 2. Fill in the required details: * Profile name * Operating system * Browser version 3. (Optional) Randomize the fingerprint for extra security. 4. Review your browser details on the right side.\ Add Profile Details *** ### Step 3: Add a Proxy 1. Go to the **Proxy** tab.\ Proxy Management Page 2. Select the connection type — for this guide, choose **HTTPS**.\ Connection Type 3. Enter your proxy details: * Proxy (IP:Port) * Username * Password Proxy Detail in AdsPower 4. Set **Geonode** as the main proxy provider. *** ### Step 4: Create and Launch the Profile 1. Click **Create Profile** to save your settings. 2. Once created, the profile will appear in the list. * If a location appears, the proxy is active.\ Check Proxy 3. Click **Start** to launch the browser with your configured proxy.\ Browser Launched Your **Geonode** proxy is now successfully configured in AdsPower. # Brave (/docs/proxies/getting-started/setup_and_configuration/browsers/brave) import ProxyOs from "../../../../../snippets/proxy-os.mdx"; import BrowsersFaqs from "../../../../../snippets/browsers-faqs.mdx"; import ProxyInfoFromGeonode from "../../../../../snippets/get-proxy-info-from-geonode.mdx"; import VerifyProxyConnectionComponent from "../../../../../snippets/verify-proxy-connection-component.mdx"; import SupportParagraph from "../../../../../snippets/support-paragraph.mdx"; This guide will help you configure a Geonode proxy in the Brave browser. ## Setting Up a Proxy in Brave ### Open Brave Settings 1. Click the three-dot menu in the top-right corner of Brave. 2. Select **Settings** from the dropdown menu. Brave Settings *** ### Access System Proxy Settings 1. In the sidebar, click **System**. 2. Then click **Open your computer’s proxy settings**. System Settings *** ### Configure the Proxy on Your Operating System Brave will now open your system proxy configuration window.\ Follow the appropriate setup guide for your OS below: Once configured, Brave will automatically route its traffic through the assigned proxy. # Chrome (/docs/proxies/getting-started/setup_and_configuration/browsers/chrome) import ProxyOs from "../../../../../snippets/proxy-os.mdx"; import BrowsersFaqs from "../../../../../snippets/browsers-faqs.mdx"; import ProxyInfoFromGeonode from "../../../../../snippets/get-proxy-info-from-geonode.mdx"; import VerifyProxyConnectionComponent from "../../../../../snippets/verify-proxy-connection-component.mdx"; import SupportParagraph from "../../../../../snippets/support-paragraph.mdx"; This guide will help you configure a Geonode proxy in Chrome using two methods: * **Using the Geonode Proxy Manager Extension (recommended)** * **Manual setup through Chrome system settings** ## Method 1: Using the Geonode Proxy Manager Extension The easiest and most flexible way to configure a proxy in Chrome is by using the **Geonode Proxy Manager** extension.\ It lets you switch proxies quickly without changing system-wide settings. ### Install and Configure the Extension 1. Install **Geonode Proxy Manager** from the [Chrome Web Store](https://chromewebstore.google.com/detail/geonode-proxy-manager/ippioaknloonmaibmmepkemhmhinohge). 2. Follow this guide to complete the setup:\ [How to Use the Geonode Chrome Extension for Proxy Management](/docs/proxies/getting-started/setup_and_configuration/extensions/geonode-proxy-manager). Once installed, Chrome will route all traffic through the selected proxy. ## Method 2: Manual Proxy Setup in Chrome If you prefer to configure the proxy manually, follow these steps: ### Open Chrome Settings 1. Click the three-dot menu in the top-right corner of Chrome. 2. Select **Settings** from the dropdown menu. Chrome Settings *** ### Access System Proxy Settings 1. In the left sidebar, click **System**. 2. Then click **Open your computer’s proxy settings**. System Settings *** ### Configure the Proxy on Your Operating System Chrome will now open your system proxy configuration window.\ Follow the appropriate guide for your OS below: Once configured, Chrome will automatically route its traffic through the assigned proxy. # ClonBrowser (/docs/proxies/getting-started/setup_and_configuration/browsers/clonbrowser) import BrowsersFaqs from "../../../../../snippets/browsers-faqs.mdx"; import ProxyInfoFromGeonode from "../../../../../snippets/get-proxy-info-from-geonode.mdx"; import SupportParagraph from "../../../../../snippets/support-paragraph.mdx"; This guide will help you configure a Geonode proxy in ClonBrowser. ## Setting Up a Proxy in ClonBrowser ### Install ClonBrowser 1. Go to the [ClonBrowser Download Page](https://www.clonbrowser.com/download). 2. Download the version for your operating system. 3. Install the browser and create an account. *** ### Create a New Profile 1. Open ClonBrowser and click **Add Profile**.\ New Profile Button 2. Fill in the required details: * Profile name * Operating system 3. (Optional) Randomize the fingerprint for extra security. 4. Review your browser details on the right panel.\ Add Profile Details *** ### Add Proxy Details 1. Open the **Proxy** tab. 2. Choose a connection type — for this guide, select **HTTP**. 3. Enter your proxy information: * Proxy (IP:Port) * Username * Password 4. Click **Create Profile** to save your configuration. Proxy Details *** ### Test the Proxy Connection 1. Click **Connect Test** to verify the connection. 2. If successful, the proxy will appear as active. Check Proxy *** ### Launch the Profile 1. Once the profile is created, it will appear in your list. * If the proxy shows a location, it means it’s active.\ Profile Created 2. Click **Start** to open the browser with your configured proxy.\ Browser Launched Your Geonode proxy is now successfully configured in ClonBrowser. # Dolphin Anty (/docs/proxies/getting-started/setup_and_configuration/browsers/dolphin-anty) import BrowsersFaqs from "../../../../../snippets/browsers-faqs.mdx"; import ProxyInfoFromGeonode from "../../../../../snippets/get-proxy-info-from-geonode.mdx"; import SupportParagraph from "../../../../../snippets/support-paragraph.mdx"; This guide will help you configure a Geonode proxy in the Dolphin Anty browser. ## Setting Up a Proxy in Dolphin Anty ### Install Dolphin Anty 1. Go to the [Dolphin Anty Download Page](https://dolphin-anty.com/download/). 2. Download the version for your operating system. 3. Install the browser and create an account. *** ### Create a New Profile 1. Open Dolphin Anty and click **Add Profile**.\ New Profile Button 2. Fill in the required details: * Profile name * Operating system 3. (Optional) Randomize the fingerprint for extra security. 4. Review your browser details on the right.\ Add Profile Details *** ### Add Proxy Details 1. Open the **Proxy** tab. 2. Choose a connection type — for this guide, select **HTTP**. 3. Enter your proxy information: * Proxy (IP:Port) * Username * Password 4. Click **Create Profile** to save the configuration. 5. A green check mark indicates the proxy is connected. Proxy Details *** ### Launch and Verify the Profile 1. Once created, your profile will appear in the list. * If the proxy shows a location, it’s active.\ Profile Created 2. Click **Start** to launch the browser with the configured proxy. 3. Visit [`http://ip-api.com/json`](http://ip-api.com/json) to confirm your IP. Browser Launched Your Geonode proxy is now successfully configured in Dolphin Anty. # Edge (/docs/proxies/getting-started/setup_and_configuration/browsers/edge) import BrowsersFaqs from "../../../../../snippets/browsers-faqs.mdx"; import ProxyInfoFromGeonode from "../../../../../snippets/get-proxy-info-from-geonode.mdx"; import VerifyProxyConnectionComponent from "../../../../../snippets/verify-proxy-connection-component.mdx"; import ProxyOs from "../../../../../snippets/proxy-os.mdx"; import SupportParagraph from "../../../../../snippets/support-paragraph.mdx"; This guide will help you configure a Geonode proxy in the Microsoft Edge browser. ## Setting Up a Proxy in Edge ### Open Edge Settings 1. Click the three-dot menu in the top-right corner of Edge. 2. Select **Settings** from the dropdown menu. Edge Settings *** ### Access System Proxy Settings 1. In the left sidebar, click **System**. 2. Then click **Open your computer’s proxy settings**. System Settings *** ### Configure the Proxy on Your Operating System Edge will now open your system proxy configuration window.\ Follow the appropriate setup guide for your OS below: Once configured, Edge will automatically route its traffic through the assigned proxy. # Firefox (/docs/proxies/getting-started/setup_and_configuration/browsers/firefox) import BrowsersFaqs from "../../../../../snippets/browsers-faqs.mdx"; import ProxyInfoFromGeonode from "../../../../../snippets/get-proxy-info-from-geonode.mdx"; import VerifyProxyConnectionComponent from "../../../../../snippets/verify-proxy-connection-component.mdx"; import SupportParagraph from "../../../../../snippets/support-paragraph.mdx"; This guide will help you configure a Geonode proxy in Firefox for secure and anonymous browsing. ## Setting Up a Proxy in Firefox ### Open Firefox Settings 1. Click the **three-line menu (☰)** in the top-right corner of Firefox. 2. Select **Settings** from the dropdown menu. Firefox Settings *** ### Access Network Settings 1. Scroll down to **Network Settings**. 2. Click **Settings** to open the proxy configuration panel. Network Settings *** ### Configure Proxy Settings 1. In the **Connection Settings** popup, enter the following details: * **Manual proxy configuration**: select this option. * **HTTP Proxy**: enter your proxy IP. * **Port**: enter your proxy port. * **SOCKS Proxy** (optional): use SOCKS5 if applicable. Connection Settings 2. Click **OK** to save changes. Firefox will now route all traffic through the configured proxy. # GeeLark (/docs/proxies/getting-started/setup_and_configuration/browsers/geelark) import BrowsersFaqs from "../../../../../snippets/browsers-faqs.mdx"; import ProxyInfoFromGeonode from "../../../../../snippets/get-proxy-info-from-geonode.mdx"; import SupportParagraph from "../../../../../snippets/support-paragraph.mdx"; This guide will help you configure a Geonode proxy in the GeeLark browser for both **mobile** and **desktop** profiles. You’ll learn how to: * Add and manage Geonode proxies in GeeLark * Create mobile and desktop profiles * Set up device environments * Launch virtual profiles and test proxy connections ## Setting Up a Proxy in GeeLark ### Install GeeLark 1. Go to the [GeeLark Download Page](https://www.geelark.com/). 2. Download the version for your operating system. 3. Install the software and create an account. *** ### For Mobile Profiles #### Create a New Profile 1. Open GeeLark and click **New Profile**.\ New Profile Button 2. Enter the required details: * Profile name * Operating system 3. Review your browser details on the right side.\ Add Profile Details *** #### Add a Proxy You can add proxies in two ways: ##### Option A: Add Proxies First (Recommended) 1. Go to the **Proxies** tab from the sidebar.\ Proxy Section 2. Click **Add Proxy**.\ Add Proxy Section 3. Enter your proxy in one of the following formats: ``` proxy.geonode.io:9000:geonode_username:password ``` You can also use: ``` username:password@host:port ``` Or ``` http://username:password@proxy.geonode.io:9000 ``` 4. Select: * **Type:** HTTP * **Proxy group:** e.g., “Geonode Proxies” * **IP Query Channel:** use `ip-api` for geolocation checks 5. Add one proxy per line (up to 100). 6. Test them using **Proxy Tests** to confirm connectivity — green icons indicate success.\ Batch Add Screenshot ##### Option B: Add Proxy During Profile Creation 1. While creating a profile, go to the **Proxy** tab. 2. Choose the connection type (e.g., HTTP). 3. Enter: * Proxy (IP:Port) * Username * Password 4. Set IP Query Channel (e.g., `ip-api`) and click **Check Proxy**.\ Proxy Management Page *** #### Configure Profile Settings 1. Under **Profile Settings**, set: * Profile name * Operating system (Android or iOS) * Group, tags, remark (optional)\ Profile Settings 2. Add the proxy: * **Custom Proxy:** enter host, port, username, password * **Saved Proxy:** choose from a pre-added list\ Check Proxy Saved Proxy Selection *** #### Configure Device Information Under **Device Information**, you can simulate mobile hardware and network behavior: * Charging Method: pay per minute or monthly * Android version: e.g., 12–15 * Network: Wi-Fi or Cellular * Phone number: auto or custom * Area, Device Brand, Language: auto or manual Full Device Settings *** #### Create and Launch the Profile 1. Click **Create** to save the profile. 2. The profile appears in your list — showing OS, proxy region, tags, etc. 3. Click the **Action (▶️)** button to launch. 4. A new window opens — visit [ip-api.com](https://ip-api.com) to verify your IP and proxy location.\ Browser Launched *** ### For Desktop Profiles #### Create a New Desktop Profile 1. In GeeLark, click **New Profile**, then switch to the desktop icon (🖥️). 2. Under **Profile Settings**, configure: * Profile name * Group / tags / remark * Operating system (Windows or macOS) * Browser (e.g., Kiwi) * User-Agent (auto or custom) * Optional cookies for session import\ Desktop Profile Form *** #### Set Proxy for Desktop Profile You can either: * Use a **Custom Proxy**, or * Select a **Saved Proxy** (from the “Geonode Proxies” group). *** #### Optional: Account Settings Configure: * Platform credentials (if needed) * Startup behavior (e.g., reopen tabs, open custom page)\ Desktop Account *** #### Optional: Advanced Settings Fine-tune fingerprinting and device behavior: * Time zone, language, geolocation * WebRTC, Canvas, WebGL, AudioContext * Resolution, fonts, storage, noise controls * Device hardware and network simulation\ Advanced Settings Overview *** #### Review Device Information The right-hand panel summarizes fingerprint and environment settings: * Browser, OS, User-Agent * Time zone, WebRTC, language, Canvas, WebGL * Fonts, resolution, and more You can click **Generate New Fingerprint** to randomize your identity.\ Desktop Device Information *** #### Launch the Desktop Profile 1. Click **Create** to save. 2. The new profile appears in your list. 3. Click the **Action (▶️)** button to launch. 4. Visit [ip-api.com](https://ip-api.com) to confirm your Geonode proxy is active.\ Desktop Browser Launched # Ghost Browser (/docs/proxies/getting-started/setup_and_configuration/browsers/ghost) import BrowsersFaqs from "../../../../../snippets/browsers-faqs.mdx"; import ProxyOs from "../../../../../snippets/proxy-os.mdx"; import ProxyInfoFromGeonode from "../../../../../snippets/get-proxy-info-from-geonode.mdx"; import SupportParagraph from "../../../../../snippets/support-paragraph.mdx"; This guide will help you configure a Geonode proxy in Ghost Browser. ## Setting Up a Proxy in Ghost Browser ### Install Ghost Browser 1. Go to the [Ghost Browser Download Page](https://ghostbrowser.com/download/). 2. Download the version for your operating system. 3. Install the software. *** ### Add a Proxy 1. Open **Proxy Control** and click **Add/Edit Proxies**.\ Add/Edit proxies 2. In the new window, select **Add a Single Proxy**.\ Single Proxy Window 3. Enter your proxy details: * Proxy (IP:Port) * Username * Password\ Proxy Detail Fields *** ### Test the Proxy Connection 1. Confirm that your proxy appears in the **Proxy Management Table**.\ Proxy Management Table 2. Click the **Test Proxy** tab. 3. Enter a website to verify your connection.\ Check Proxy 4. If successful, a confirmation popup will appear.\ Connection Popup Your Geonode proxy is now successfully configured in Ghost Browser. # GoLogin (/docs/proxies/getting-started/setup_and_configuration/browsers/gologin) import BrowsersFaqs from "../../../../../snippets/browsers-faqs.mdx"; import ProxyInfoFromGeonode from "../../../../../snippets/get-proxy-info-from-geonode.mdx"; import SupportParagraph from "../../../../../snippets/support-paragraph.mdx"; This guide will help you configure a Geonode proxy in the GoLogin browser. ## Setting Up a Proxy in GoLogin ### Install GoLogin 1. Go to the [GoLogin Download Page](https://gologin.com/download-started/). 2. Download the version for your operating system. 3. Install the software and create an account. *** ### Create a New Profile 1. Open GoLogin and click **Add Profile**.\ New Profile Button 2. Enter the required details: * Profile name * Operating system 3. (Optional) Randomize the fingerprint for extra security. 4. Review your browser details on the right.\ Add Profile Details *** ### Add Proxy Details 1. Go to the **Proxy** tab. 2. Choose the connection type — for this guide, select **HTTP**. 3. Enter your proxy credentials: * Proxy (IP:Port) * Username * Password 4. Click **Create Profile** to save the settings.\ Proxy Details *** ### Test the Proxy Connection 1. Click **Check Proxy** to verify the connection. 2. If successful, the proxy will show as active.\ Check Proxy *** ### Launch the Profile 1. Once created, the profile will appear in your list. * If the proxy displays a location, it means it’s active.\ Profile Created 2. Click **Start** to launch the browser with your configured proxy. 3. Visit [`http://ip-api.com/json`](http://ip-api.com/json) to verify your IP and proxy location.\ Browser Launched Your Geonode proxy is now successfully configured in GoLogin. # Incognito Mode (/docs/proxies/getting-started/setup_and_configuration/browsers/incognito-mode) import ProxyOs from "../../../../../snippets/proxy-os.mdx"; import BrowsersFaqs from "../../../../../snippets/browsers-faqs.mdx"; import ProxyInfoFromGeonode from "../../../../../snippets/get-proxy-info-from-geonode.mdx"; import VerifyProxyConnectionComponent from "../../../../../snippets/verify-proxy-connection-component.mdx"; import SupportParagraph from "../../../../../snippets/support-paragraph.mdx"; This guide will help you configure a Geonode proxy in **Chrome’s Incognito Mode**. ## Setting Up a Proxy in Incognito Mode ### Open Chrome Settings 1. Click the **three-dot menu** in the top-right corner of Chrome. 2. Select **Settings** from the dropdown menu. Chrome Settings *** ### Access System Proxy Settings 1. In the left sidebar, click **System**. 2. Then click **Open your computer’s proxy settings**. System Settings *** ### Configure the Proxy on Your Operating System Chrome will now open your system’s proxy configuration window.\ Follow the appropriate setup guide for your OS below: Once configured, Chrome will automatically route all traffic—including Incognito Mode—through the assigned proxy. # Incogniton (/docs/proxies/getting-started/setup_and_configuration/browsers/incogniton) import BrowsersFaqs from "../../../../../snippets/browsers-faqs.mdx"; import ProxyInfoFromGeonode from "../../../../../snippets/get-proxy-info-from-geonode.mdx"; import SupportParagraph from "../../../../../snippets/support-paragraph.mdx"; This guide will help you configure a Geonode proxy in the Incogniton browser. ## Setting Up a Proxy in Incogniton ### Install Incogniton 1. Go to the [Incogniton Download Page](https://incogniton.com/download-incogniton/). 2. Download the version for your operating system. 3. Install the software and create an account. *** ### Create a New Profile 1. Open Incogniton and click **New Profile**.\ New Profile Button 2. Fill in the required details: * Profile name * Operating system * Browser version 3. (Optional) Randomize the fingerprint for added security. 4. Review your browser details on the right panel.\ Add Profile Details *** ### Add Proxy Settings 1. Go to the **Proxy Management** page.\ Proxy Management Page 2. Choose the connection type — for this guide, select **HTTP**.\ Connection Type 3. Enter your proxy credentials: * Proxy (IP:Port) * Username * Password\ Proxy Detail Fields *** ### Check the Proxy Connection 1. Click **Check Proxy** to verify the connection.\ Check Proxy 2. If successful, the proxy will show as active. *** ### Create and Launch the Profile 1. Click **Create Profile** to save your settings.\ Create Profile 2. Once created, you’ll see the profile in your list. * If the proxy shows a green tick, it means it’s active.\ Profile Created 3. Click **Start** to launch the browser with your configured proxy.\ Browser Launched Your Geonode proxy is now successfully configured in Incogniton. # MoreLogin (/docs/proxies/getting-started/setup_and_configuration/browsers/morelogin) import BrowsersFaqs from "../../../../../snippets/browsers-faqs.mdx"; import ProxyInfoFromGeonode from "../../../../../snippets/get-proxy-info-from-geonode.mdx"; import SupportParagraph from "../../../../../snippets/support-paragraph.mdx"; This guide will help you configure a Geonode proxy in the MoreLogin browser. ## Setting Up a Proxy in MoreLogin ### Install MoreLogin 1. Go to the [MoreLogin Download Page](https://www.morelogin.com/). 2. Download the version for your operating system. 3. Install the software and create an account. *** ### Create a New Profile 1. Open MoreLogin and click **Add Profile**.\ New Profile Button 2. Enter the required details: * Profile name * Operating system 3. (Optional) Randomize the fingerprint for added security. 4. Review your browser details on the right side.\ Add Profile Details *** ### Add Proxy Details 1. Go to the **Proxy** tab. 2. Choose the connection type — for this guide, select **HTTP**. 3. Enter your proxy credentials: * Proxy (IP:Port) * Username * Password 4. Click **Create Profile** to save the configuration.\ Proxy Details *** ### Test the Proxy Connection 1. Click **Proxy Detection** to verify the connection. 2. If successful, the proxy will show as active.\ Check Proxy *** ### Launch the Profile 1. Click **Create Profile** to save your settings. 2. Once created, your profile will appear in the list. * If the proxy shows a location, it means it’s active.\ Profile Created 3. Click **Start** to launch the browser with your configured proxy.\ Browser Launched Your Geonode proxy is now successfully configured in MoreLogin. # MultiLogin (/docs/proxies/getting-started/setup_and_configuration/browsers/multilogin) import BrowsersFaqs from "../../../../../snippets/browsers-faqs.mdx"; import ProxyInfoFromGeonode from "../../../../../snippets/get-proxy-info-from-geonode.mdx"; import SupportParagraph from "../../../../../snippets/support-paragraph.mdx"; This guide will help you configure a Geonode proxy in the MultiLogin browser. ## Setting Up a Proxy in MultiLogin ### Install MultiLogin 1. Go to the [MultiLogin Download Page](https://multilogin.com/). 2. Download the version for your operating system. 3. Install the software and create an account. *** ### Create a New Profile 1. Open MultiLogin and click **Add Profile**.\ New Profile Button 2. Enter the required details: * Profile name * Operating system 3. (Optional) Randomize the fingerprint for extra security. 4. Review your browser details on the right side.\ Add Profile Details *** ### Add Proxy Details 1. Open the **Proxy** tab. 2. Choose the connection type — for this guide, select **HTTP**. 3. Enter your proxy credentials: * Proxy (IP:Port) * Username * Password 4. Click **Create Profile** to save the configuration.\ Proxy Details *** ### Test the Proxy Connection 1. Click **Proxy Detection** to verify the connection. 2. If successful, the proxy will show as active.\ Check Proxy *** ### Launch the Profile 1. Click **Create Profile** to save your settings. 2. Once created, your profile will appear in the list. * If the proxy shows a location, it means it’s active.\ Profile Created 3. Click **Start** to launch the browser with your configured proxy.\ Browser Launched Your Geonode proxy is now successfully configured in MultiLogin. # Octo Browser (/docs/proxies/getting-started/setup_and_configuration/browsers/octo) import BrowsersFaqs from "../../../../../snippets/browsers-faqs.mdx"; import ProxyInfoFromGeonode from "../../../../../snippets/get-proxy-info-from-geonode.mdx"; import SupportParagraph from "../../../../../snippets/support-paragraph.mdx"; This guide will help you configure a Geonode proxy in the Octo Browser. ## Setting Up a Proxy in Octo Browser ### Install Octo Browser 1. Go to the [Octo Browser Download Page](https://octobrowser.org/download/). 2. Download the version for your operating system. 3. Install the software and create an account. *** ### Create a New Profile 1. Open Octo Browser and click **Add Profile**.\ New Profile Button 2. Enter the required details: * Profile name * Operating system 3. (Optional) Randomize the fingerprint for extra security. 4. Review your browser details.\ Add Profile Details *** ### Add Proxy Details 1. Click the **Proxy** button.\ Proxy 2. Choose a connection type — for this guide, select **HTTP**. 3. Enter your proxy credentials: * Proxy (IP:Port) * Username * Password\ Proxy Details 4. Click **Check Proxy Connection** to test it. 5. Click **Confirm** to save your proxy settings. *** ### Launch the Profile 1. Once created, the profile will appear in your list. * If the proxy shows a location, it means it’s active.\ Profile Created 2. Click **Start** to launch the browser with your configured proxy. 3. Visit [`http://ip-api.com`](http://ip-api.com) to verify your IP and proxy location.\ Browser Launched Your Geonode proxy is now successfully configured in Octo Browser. # Safari (/docs/proxies/getting-started/setup_and_configuration/browsers/safari) import BrowsersFaqs from "../../../../../snippets/browsers-faqs.mdx"; import ProxyInfoFromGeonode from "../../../../../snippets/get-proxy-info-from-geonode.mdx"; import VerifyProxyConnectionComponent from "../../../../../snippets/verify-proxy-connection-component.mdx"; import SupportParagraph from "../../../../../snippets/support-paragraph.mdx"; This guide will help you configure a Geonode proxy in Safari. ## Setting Up a Proxy in Safari ### Open Safari Settings 1. In the top menu bar, click **Safari → Settings** (or **Preferences** on older macOS versions). 2. Select the **Advanced** tab. 3. Click **Change Settings** next to **Proxies**. Safari Settings *** ### Configure the Proxy on macOS 1. The **Network** window will open in System Preferences. 2. Select your active network connection (Wi-Fi or Ethernet). 3. Click **Advanced → Proxies**. 4. Choose the protocol you want to configure (e.g., HTTP or HTTPS). 5. Enter your Geonode proxy credentials: * **Proxy server:** IP and Port * **Username / Password** if required System Settings 6. Click **OK**, then **Apply** to save your settings. *** ### Verify the Proxy Connection Once configured, Safari will automatically route traffic through your Geonode proxy.\ You can verify that it’s working correctly below. # FoxyProxy (/docs/proxies/getting-started/setup_and_configuration/extensions/foxyproxy) import ProxyInfoFromGeonode from "../../../../../snippets/get-proxy-info-from-geonode.mdx"; import VerifyProxyConnectionComponent from "../../../../../snippets/verify-proxy-connection-component.mdx"; import SupportParagraph from "../../../../../snippets/support-paragraph.mdx"; import ExtensionFAQs from "../../../../../snippets/extensions-faqs.mdx"; This guide explains how to install, configure, and use **Geonode proxies** in the **FoxyProxy** browser extension. *** ## What Is FoxyProxy FoxyProxy is a browser extension that lets you easily manage and switch between multiple proxy configurations.\ Instead of manually entering proxy details each time, you can save and control proxies directly within your browser. *** *** ## Steps to Set Up Geonode Proxy in FoxyProxy Follow these steps to configure your Geonode proxy using the FoxyProxy extension. *** ### Step 1: Install the FoxyProxy Extension 1. Open the [FoxyProxy extension page](https://chromewebstore.google.com/detail/foxyproxy/gcknhkkoolaabfmlnjonogaaifnjlfnp) in the Chrome Web Store. 2. Click **Add to Chrome** to install the extension. Add to Chrome 3. Confirm the installation by clicking **Add Extension** when prompted. Accept and add the extension *** ### Step 2: Pin the Extension for Quick Access After installation, pin the extension to your Chrome toolbar for faster access. 1. Click the **Extensions** icon (puzzle piece) in Chrome. 2. Find **FoxyProxy** and click the **Pin** icon. Pin the Extension *** ### Step 3: Open FoxyProxy Settings 1. Click the **FoxyProxy** icon in your Chrome toolbar. 2. Select **Options** to open the settings panel. Add New Proxy 3. You’ll be redirected to the **Proxy Manager** page.\ Click on the **Proxies** tab. Geonode proxy manager page *** ### Step 4: Add a New Proxy 1. Click **Add New Proxy**. 2. Fill in your proxy details in the popup window: * **Name**: Any label for easy identification * **Host**: Your proxy server (e.g., `proxy.geonode.io`) * **Port**: Usually `9000` * **Username** and **Password**: Your Geonode credentials 3. Click **Add Proxy** to save the configuration. Add Proxy 4. Return to the **Proxy Manager** to confirm that your new proxy appears in the list. Proxy Added *** ### Step 5: Connect to the Proxy 1. Click the **FoxyProxy** icon in the Chrome toolbar. 2. Select the proxy you added from the list. Extension Showing Proxy Added 3. Click **Connect** to activate the proxy.\ Once connected, all your browser traffic will route through the selected proxy. *** *** *** # Geonode Proxy Manager (/docs/proxies/getting-started/setup_and_configuration/extensions/geonode-proxy-manager) import ProxyInfoFromGeonode from "../../../../../snippets/get-proxy-info-from-geonode.mdx"; import VerifyProxyConnectionComponent from "../../../../../snippets/verify-proxy-connection-component.mdx"; import SupportParagraph from "../../../../../snippets/support-paragraph.mdx"; This guide explains how to install, configure, and use the **Geonode Proxy Manager** Chrome extension. *** ## What Is Geonode Proxy Manager Geonode Proxy Manager is a Chrome extension that simplifies proxy management.\ You can add, save, and switch between multiple proxies directly from your browser without manually entering details each time. *** *** ## Setting Up a Proxy in Geonode Proxy Manager Follow these steps to configure your proxy using the Geonode Proxy Manager extension. *** ### Step 1: Install the Geonode Proxy Manager Extension 1. Open the [Geonode Proxy Manager extension page](https://chromewebstore.google.com/detail/geonode-proxy-manager/ippioaknloonmaibmmepkemhmhinohge). 2. Click **Add to Chrome** to install the extension. Add to Chrome 3. Confirm installation by selecting **Add Extension** when prompted. Accept and add the extension *** ### Step 2: Pin the Extension for Quick Access After installation, pin the extension to your Chrome toolbar for easy access. 1. Click the **Extensions** icon (puzzle piece) in Chrome. 2. Find **Geonode Proxy Manager** and click the **Pin** icon. Pin the Extension *** ### Step 3: Open the Extension 1. Click the **Geonode Proxy Manager** icon in your Chrome toolbar. 2. If no proxies have been added yet, click **Add New Proxy**. Add New Proxy 3. You’ll be redirected to the Proxy Manager page. Geonode proxy manager page *** ### Step 4: Add a New Proxy 1. Click **Add New Proxy**. 2. Enter your proxy details in the popup: * **Name** — any recognizable label for this proxy * **Host** — e.g., `proxy.geonode.io` * **Port** — usually `9000` * **Username** and **Password** — your Geonode credentials Enter Proxy Details 3. Click **Add Proxy** to save your configuration. Add Proxy 4. Once saved, go back to the Proxy Manager to confirm that your new proxy appears in the list. Proxy Added *** ### Step 5: Connect to the Proxy 1. Click the **Geonode Proxy Manager** icon again in your Chrome toolbar. 2. Select the proxy you want to use from the list. Extension Showing Proxy Added 3. Click **Connect** to activate the proxy.\ Once connected, all browser traffic will be routed through the selected proxy. Proxy Connected *** *** *** ## FAQs Yes, you can add multiple proxies and switch between them by selecting the desired one from the list and clicking **Connect**. It saves time by storing proxy details for quick switching, enhances privacy by routing browser traffic through different servers, and supports multiple proxies for varied use cases. Check your IP address using an online tool like [IP API](https://ip-api.com/) or follow this guide: [Verify Proxy Connection](/docs/proxies/getting-started/setup_and_configuration/verify-proxy-connection). Make sure you entered the correct proxy details (host, port, username, password), try reconnecting, and check that your proxy plan is active in the Geonode Dashboard. Yes, but some sites may block proxies. If that happens, switch to a different server or location. Yes, but using a VPN and a proxy simultaneously may cause slower speeds or connection conflicts. Yes, it’s free to install, but you’ll need an active Geonode Proxy Plan to use proxy services. # Android (/docs/proxies/getting-started/setup_and_configuration/mobile-OS/android) import ProxyInfoFromGeonode from "../../../../../snippets/get-proxy-info-from-geonode.mdx"; import VerifyProxyConnectionComponent from "../../../../../snippets/verify-proxy-connection-component.mdx"; import BrowsersFaqs from "../../../../../snippets/browsers-faqs.mdx"; import SupportParagraph from "../../../../../snippets/support-paragraph.mdx"; import PopupAuth from "../../../../../snippets/popup-auth.mdx"; Android settings may look slightly different depending on the device model and Android version, but the overall process remains the same. *** *** ## Step-by-step setup for Android ### Step 1 — Open Settings 1. Tap the **Settings** app (gear icon). Open Settings *** ### Step 2 — Go to Wi-Fi Settings 1. Scroll down and tap **Network & Internet** or **Wi-Fi**, depending on your device. 2. Tap and hold the Wi-Fi network you’re connected to. 3. Select **Modify network** or tap the gear (⚙) or “i” icon next to the Wi-Fi name. All setting Ensure Wi-Fi is turned on and connected to the network you want to configure the proxy for. *** ### Step 3 — Access Wi-Fi List You will now see a list of available Wi-Fi networks. All Wi-Fi list *** ### Step 4 — Find Proxy Settings 1. Scroll down until you find **Proxy settings**. 2. Tap **Proxy** to expand available options. You’ll also see your current IP address — write it down for future reference. Wi-Fi setting *** ### Step 5 — Choose Manual Configuration Select **Manual** from the proxy settings. Manual Proxy Configuration *** ### Step 6 — Enter Proxy Details 1. Enter the **Proxy IP Address** and **Port** from your Geonode Dashboard. Credentials *** ### Step 7 — Save and Exit 1. Scroll down and tap **Save** to apply the settings. 2. If your device doesn’t have a Save button, simply exit — the configuration will be applied automatically. *** *** *** # iOS (/docs/proxies/getting-started/setup_and_configuration/mobile-OS/ios) import ProxyInfoFromGeonode from "../../../../../snippets/get-proxy-info-from-geonode.mdx"; import VerifyProxyConnectionComponent from "../../../../../snippets/verify-proxy-connection-component.mdx"; import BrowsersFaqs from "../../../../../snippets/browsers-faqs.mdx"; import SupportParagraph from "../../../../../snippets/support-paragraph.mdx"; import PopupAuth from "../../../../../snippets/popup-auth.mdx"; *** *** ## Step-by-step setup for iOS Follow these steps to configure your proxy manually: ### Step 1 — Open iPhone Settings 1. Open the **Settings** app. 2. Tap **Wi-Fi** to view available networks. Open Mobile Settings *** ### Step 2 — Access Wi-Fi Network Settings 1. Connect to the Wi-Fi network you want to configure. 2. Tap the **“i” icon** next to the connected network. i icon *** ### Step 3 — Navigate to Proxy Settings 1. Scroll down to the **HTTP Proxy** section. Proxy Configuration 2. You will see three options: * **Off** — disables proxy * **Manual** — enter proxy details manually * **Automatic** — configure using a PAC file Select **Manual** for your Geonode proxy setup. Proxy Manual Configuration *** ### Step 4 — Enter Proxy Details 1. Under **Manual Proxy Configuration**, fill in: * **Server:** Proxy IP address * **Port:** Port number from your Geonode Dashboard 2. If authentication is required, enable **Authentication** and enter your **username** and **password**. Enter Proxy Details 3. Tap **Save** to apply the settings. *** *** ***