Premip
← All solutions

AI data collection

A practical data access layer for AI teams

Connect research, evaluation, and product pipelines to the public sources they depend on, with geography and session controls you can document.

Who this is for

Applied AI, data, and platform teams collecting public web sources for features, evals, or monitoring — not only foundation-model training.

Keep proxy access out of the model repo

Treat Premip as a documented access layer: locations, credentials, and session rules living next to the job that fetches sources.

Pipeline benefits

Sources from the right place

Fetch the regional version of a page when the product is location-specific.

Validation across markets

Compare how content appears before you ship an AI feature.

Clear ownership

Models and prompts stay in your repos. Credentials stay in your secret store.

Developer-ready controls

Source coverage

Reach public URLs from dashboard locations your product cares about.

Fresh or stable access

Rotate between jobs; stay sticky during a multi-page fetch.

API-friendly credentials

HTTP proxy URL works in workers, notebooks, and CI smoke tests.

AI product workflows

  • Research sources

    Gather public documents for analysis.

  • Regional validation

    Check outputs against in-market page content.

  • Retrieval refresh

    Re-fetch sources on a schedule for indexes and monitors.

Residential or mobile

Use residential or mobile when sources differ for ordinary users. Datacenter can work for simple public documentation sites.

Sessions

Sticky per crawl task. Rotate between independent pipeline runs.

Location

Align with the user market the AI feature serves.

Before you start

  • A source list and the markets that matter to the product.
  • A worker or notebook that can set an HTTP proxy.
  • Secret storage for credentials.
  • A Premip plan with those locations.

How to connect an AI data job

Wire credentials once, validate one source, then add the job to your pipeline runner.

  1. Step 1

    Identify the source set

    List URLs or query patterns, required markets, and how fresh the data must be.

  2. Step 2

    Create the proxy

    Sign up, pick a product, and open the dashboard credential fields.

  3. Step 3

    Store secrets correctly

    Put USERNAME, PASSWORD, HOST, and PORT in your secret manager. Reference them from the job — never from git.

  4. Step 4

    Set the location

    Choose the market the feature serves. If you support many countries, parameterize location per run.

  5. Step 5

    Fetch a canary URL

    Run the worker once. Assert status, encoding, and that the body matches the regional expectation.

  6. Step 6

    Configure session policy

    Sticky for multi-page source assembly. Rotate after the job if the next run should use a new IP.

  7. Step 7

    Review data quality

    Sample records for language and block pages. Keep location on every stored document.

  8. Step 8

    Schedule in your orchestrator

    Add the job beside other pipeline steps. Tune concurrency after canaries stay clean.

Canary request from a worker image
Replace USERNAME, PASSWORD, HOST, and PORT with values from your dashboard.
curl -x http://USERNAME:PASSWORD@HOST:PORT https://example.com

AI data questions

Is this intended for model training only?
No. It can support research, evaluation, retrieval, monitoring, and other workflows on public web sources.
Can developers automate access?
Yes. Proxy credentials work with automated tools that support proxy connections.
How is this different from Data for AI and LLMs?
Data for AI and LLMs focuses on model datasets and evals. This page covers broader AI product pipelines and validation.
Where are language examples?
The documentation page includes Python, Node.js, Go, and other clients.

Connect the first AI fetch job

Create credentials, store them in secrets, and pass the canary URL before you schedule the full source list.