AI data collection
A practical data access layer for AI teams
Connect research, evaluation, and product pipelines to the public sources they depend on, with geography and session controls you can document.
Who this is for
Applied AI, data, and platform teams collecting public web sources for features, evals, or monitoring — not only foundation-model training.
Keep proxy access out of the model repo
Treat Premip as a documented access layer: locations, credentials, and session rules living next to the job that fetches sources.
Pipeline benefits
Sources from the right place
Fetch the regional version of a page when the product is location-specific.
Validation across markets
Compare how content appears before you ship an AI feature.
Clear ownership
Models and prompts stay in your repos. Credentials stay in your secret store.
Developer-ready controls
Source coverage
Reach public URLs from dashboard locations your product cares about.
Fresh or stable access
Rotate between jobs; stay sticky during a multi-page fetch.
API-friendly credentials
HTTP proxy URL works in workers, notebooks, and CI smoke tests.
AI product workflows
Research sources
Gather public documents for analysis.
Regional validation
Check outputs against in-market page content.
Retrieval refresh
Re-fetch sources on a schedule for indexes and monitors.
Recommended setup
Residential or mobile
Use residential or mobile when sources differ for ordinary users. Datacenter can work for simple public documentation sites.
Sessions
Sticky per crawl task. Rotate between independent pipeline runs.
Location
Align with the user market the AI feature serves.
Before you start
- A source list and the markets that matter to the product.
- A worker or notebook that can set an HTTP proxy.
- Secret storage for credentials.
- A Premip plan with those locations.
How to connect an AI data job
Wire credentials once, validate one source, then add the job to your pipeline runner.
Step 1
Identify the source set
List URLs or query patterns, required markets, and how fresh the data must be.
Step 2
Create the proxy
Sign up, pick a product, and open the dashboard credential fields.
Step 3
Store secrets correctly
Put USERNAME, PASSWORD, HOST, and PORT in your secret manager. Reference them from the job — never from git.
Step 4
Set the location
Choose the market the feature serves. If you support many countries, parameterize location per run.
Step 5
Fetch a canary URL
Run the worker once. Assert status, encoding, and that the body matches the regional expectation.
Step 6
Configure session policy
Sticky for multi-page source assembly. Rotate after the job if the next run should use a new IP.
Step 7
Review data quality
Sample records for language and block pages. Keep location on every stored document.
Step 8
Schedule in your orchestrator
Add the job beside other pipeline steps. Tune concurrency after canaries stay clean.
curl -x http://USERNAME:PASSWORD@HOST:PORT https://example.comAI data questions
- Is this intended for model training only?
- No. It can support research, evaluation, retrieval, monitoring, and other workflows on public web sources.
- Can developers automate access?
- Yes. Proxy credentials work with automated tools that support proxy connections.
- How is this different from Data for AI and LLMs?
- Data for AI and LLMs focuses on model datasets and evals. This page covers broader AI product pipelines and validation.
- Where are language examples?
- The documentation page includes Python, Node.js, Go, and other clients.
Message us
Reach support on Telegram, email, phone, or the contact form. Include your account email and, if you have one, your order ID.