Premip
← All solutions

Web data collection

A dependable proxy layer for data collection

Collect public web data at the locations and scale your workflow demands, using the crawler and scheduler you already operate.

Who this is for

Data engineers and analysts who run recurring jobs against public catalogs, listings, or content and need fewer blocks plus consistent geography.

Make recurring jobs easier to operate

Match proxy type, location, and session behavior to each target so collection stays in your stack instead of a closed scraping platform.

Operational gains

Predictable access

Reuse the same host, port, and user for every run. Change location or rotate IP when the job requires it.

The right session model

Sticky sessions for multi-page flows. Fresh IPs when each request should look independent.

Market expansion without new code

Point the same collector at another dashboard location when you add a country.

Collection controls

Rotation when you need a new IP

Rotate from the dashboard. Credentials do not change.

Geographic targeting

Collect market-specific pages, prices, and availability from the location you select.

Standard endpoints

HTTP and SOCKS5 work with browsers, headless tools, and scheduled scripts.

Jobs this setup supports

  • Catalog watches

    Pull public product or listing pages on a schedule.

  • Competitor and market research

    Capture pages that differ by region.

  • Internal datasets

    Feed warehouses and notebooks with repeatable fetches.

Residential or mobile

Use mobile when the target is sensitive to datacenter traffic. Use residential for general public sites. Use datacenter when the target is stable and you need straightforward throughput.

Sessions

Sticky for pagination and carts; rotate between scheduled runs if you want a new IP each cycle.

Location

Set the dashboard location to the market the dataset should represent before the first production run.

Before you start

  • A written target list: sites, paths, request rate, and required countries.
  • An active Premip plan with those locations.
  • A scheduler (cron, Airflow, or your current runner).
  • Logging so you can tell a block page from a real payload.

How to run a collection job

Pilot one site and one location, then clone the job for additional markets.

  1. Step 1

    Map the target

    Document the site, the fields you need, polite request rates, and whether pages vary by country. Note any cookies or headers your client already sends.

  2. Step 2

    Sign up and provision a proxy

    Choose a plan, complete payment, and open the proxy in the dashboard.

  3. Step 3

    Select location and proxy type

    Match the product (mobile, residential, or datacenter) and location to the trust and geography the site requires.

  4. Step 4

    Copy credentials into the job config

    Set HTTP_PROXY and HTTPS_PROXY, or the equivalent in your client, to http://USERNAME:PASSWORD@HOST:PORT. Keep secrets out of the repository.

  5. Step 5

    Prove a single request

    Fetch one URL through the proxy and save the body. Confirm you received the expected language, currency, or layout for that location.

  6. Step 6

    Choose sticky vs rotate

    If the job walks listing → detail pages, keep the session. If each URL is independent, rotate between batches from the dashboard.

  7. Step 7

    Add retries and block detection

    Treat unexpected status codes, tiny bodies, and known block templates as failures. Back off, then retry or rotate after the current batch finishes.

  8. Step 8

    Schedule and expand

    Put the job on your scheduler. Add locations only after the first market stays clean for several runs.

First request through the proxy
Replace USERNAME, PASSWORD, HOST, and PORT with values from your dashboard.
curl -x http://USERNAME:PASSWORD@HOST:PORT https://example.com

Data collection questions

What can I use the proxies with?
Any client that supports HTTP or SOCKS5: scripts, browsers, crawlers, and most collection tools.
Can I target a specific market?
Yes. Location targeting depends on the proxy product and plan you select in the dashboard.
Do I need Premip to host the crawler?
No. Run collection on your machines or cloud jobs. Premip only terminates the proxy connection.
How do I add another country later?
Provision or switch the dashboard location, keep the same collector code, and tag stored records with the new market.

Put your collector on a real location

Get credentials from the dashboard and run the first URL through the proxy before you scale the site list.