Web data collection
A dependable proxy layer for data collection
Collect public web data at the locations and scale your workflow demands, using the crawler and scheduler you already operate.
Who this is for
Data engineers and analysts who run recurring jobs against public catalogs, listings, or content and need fewer blocks plus consistent geography.
Make recurring jobs easier to operate
Match proxy type, location, and session behavior to each target so collection stays in your stack instead of a closed scraping platform.
Operational gains
Predictable access
Reuse the same host, port, and user for every run. Change location or rotate IP when the job requires it.
The right session model
Sticky sessions for multi-page flows. Fresh IPs when each request should look independent.
Market expansion without new code
Point the same collector at another dashboard location when you add a country.
Collection controls
Rotation when you need a new IP
Rotate from the dashboard. Credentials do not change.
Geographic targeting
Collect market-specific pages, prices, and availability from the location you select.
Standard endpoints
HTTP and SOCKS5 work with browsers, headless tools, and scheduled scripts.
Jobs this setup supports
Catalog watches
Pull public product or listing pages on a schedule.
Competitor and market research
Capture pages that differ by region.
Internal datasets
Feed warehouses and notebooks with repeatable fetches.
Recommended setup
Residential or mobile
Use mobile when the target is sensitive to datacenter traffic. Use residential for general public sites. Use datacenter when the target is stable and you need straightforward throughput.
Sessions
Sticky for pagination and carts; rotate between scheduled runs if you want a new IP each cycle.
Location
Set the dashboard location to the market the dataset should represent before the first production run.
Before you start
- A written target list: sites, paths, request rate, and required countries.
- An active Premip plan with those locations.
- A scheduler (cron, Airflow, or your current runner).
- Logging so you can tell a block page from a real payload.
How to run a collection job
Pilot one site and one location, then clone the job for additional markets.
Step 1
Map the target
Document the site, the fields you need, polite request rates, and whether pages vary by country. Note any cookies or headers your client already sends.
Step 2
Sign up and provision a proxy
Choose a plan, complete payment, and open the proxy in the dashboard.
Step 3
Select location and proxy type
Match the product (mobile, residential, or datacenter) and location to the trust and geography the site requires.
Step 4
Copy credentials into the job config
Set HTTP_PROXY and HTTPS_PROXY, or the equivalent in your client, to http://USERNAME:PASSWORD@HOST:PORT. Keep secrets out of the repository.
Step 5
Prove a single request
Fetch one URL through the proxy and save the body. Confirm you received the expected language, currency, or layout for that location.
Step 6
Choose sticky vs rotate
If the job walks listing → detail pages, keep the session. If each URL is independent, rotate between batches from the dashboard.
Step 7
Add retries and block detection
Treat unexpected status codes, tiny bodies, and known block templates as failures. Back off, then retry or rotate after the current batch finishes.
Step 8
Schedule and expand
Put the job on your scheduler. Add locations only after the first market stays clean for several runs.
curl -x http://USERNAME:PASSWORD@HOST:PORT https://example.comData collection questions
- What can I use the proxies with?
- Any client that supports HTTP or SOCKS5: scripts, browsers, crawlers, and most collection tools.
- Can I target a specific market?
- Yes. Location targeting depends on the proxy product and plan you select in the dashboard.
- Do I need Premip to host the crawler?
- No. Run collection on your machines or cloud jobs. Premip only terminates the proxy connection.
- How do I add another country later?
- Provision or switch the dashboard location, keep the same collector code, and tag stored records with the new market.
Message us
Reach support on Telegram, email, phone, or the contact form. Include your account email and, if you have one, your order ID.