# Migrate from Coresignal

> How do I translate my Coresignal queries and fields to BetterJobs?

Source: https://docs.betterjobs.cc/guides/migrate-from-coresignal/

Coresignal splits job retrieval in two: a search that finds matches, then a collect call per job. Its Base API returns one row per source and flags duplicates. BetterJobs does all of it in one request: search, fetch, merge and dedup, billed per unique job.

Coresignal is also one of the six providers behind BetterJobs, so its postings can still reach you through the waterfall.

### [Coresignal](https://docs.betterjobs.cc/providers/coresignal.md)

`coresignal`

Large historical job-posting dataset plus company and employee records.

- Records

  Job postings, company records, employee records

- Postings

  475M+ job postings, 70M+ active

- History

  Since August 2020

- Recheck

  Active postings rechecked within 24h

- Credits

  Search free; collect 1 credit per job, 20 per company or employee record

Growth includes 2 partner providers; Pro and above include all six. GET /v1/providers shows what your plan enables.

Facts per Coresignal docs.

## What changes, in one table

| Topic                        | Coresignal (per Coresignal docs)                          | BetterJobs                                                                                                           |
| ---------------------------- | --------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------- |
| Flow                         | Search (free), then collect each job                      | One `POST /v1/jobs/search` returns full jobs                                                                         |
| Free count before you pay    | Search is free                                            | `dry_run: true` returns a free estimate                                                                              |
| Job price                    | 1 credit per job collected                                | 1 credit per unique job returned                                                                                     |
| Duplicates                   | One row per source with an `isDuplicate` flag; you filter | One canonical job per opening; every source in `sources[]`; duplicates free                                          |
| Query language               | Elasticsearch DSL or flat filters                         | One flat `filters` object with suffix grammar. No DSL. See [Filters](https://docs.betterjobs.cc/platform/filters.md) |
| History                      | Since August 2020                                         | `posted_within_days` up to 365, `include_closed: true` for closed jobs                                               |
| Large pulls                  | Bulk JSONL, Parquet or CSV to S3, GCS, Azure or Snowflake | Async searches, up to 10,000 jobs each. No bulk file delivery in v1 preview                                          |
| Company and employee records | 20 credits each                                           | `GET /v1/companies/{domain}`: 1 credit, hiring profile only. No employee records                                     |

> Not a replacement for Coresignal's datasets
>
> If you rely on Coresignal’s multi-year history, employee records or bulk file delivery to a warehouse, BetterJobs v1 preview does not cover that. Keep Coresignal for those jobs. See [When not to use BetterJobs](https://docs.betterjobs.cc/resources/when-not-to-use.md).

## Translate a request

Before: search, collect each hit, drop duplicate rows. `coresignal_search` and `coresignal_collect` stand for your existing wrappers around Coresignal’s endpoints.

```python
# Before (Coresignal): 1 search + N collect calls + your own dedup
ids = coresignal_search(query)                       # free per Coresignal docs
rows = [coresignal_collect(job_id) for job_id in ids] # 1 credit per job collected
jobs = [r for r in rows if not r["isDuplicate"]]      # one row per source, so filter
```

After: one request returns merged, deduplicated jobs. Start with a free estimate, then fetch with a cap.

**curl**

```bash
# Free estimate (replaces the free search step)
curl https://api.betterjobs.cc/v1/jobs/search \
  -H "Authorization: Bearer $BETTERJOBS_API_KEY" \
  -H "BetterJobs-Version: 2026-10-01" \
  -H "Content-Type: application/json" \
  -d '{
    "filters": {
      "title_or": ["Data Engineer", "Analytics Engineer"],
      "country_code_or": ["DE", "NL"],
      "posted_within_days": 30
    },
    "waterfall": { "strategy": "max_coverage" },
    "limit": 100,
    "dry_run": true
  }'


# Fetch (replaces search + collect + dedup). Drop dry_run, add a cap.
curl https://api.betterjobs.cc/v1/jobs/search \
  -H "Authorization: Bearer $BETTERJOBS_API_KEY" \
  -H "BetterJobs-Version: 2026-10-01" \
  -H "Content-Type: application/json" \
  -d '{
    "filters": {
      "title_or": ["Data Engineer", "Analytics Engineer"],
      "country_code_or": ["DE", "NL"],
      "posted_within_days": 30
    },
    "waterfall": { "strategy": "max_coverage", "max_credits": 100 },
    "limit": 100
  }'
```

**Python**

```python
import os


import requests


API = "https://api.betterjobs.cc/v1"
HEADERS = {
    "Authorization": f"Bearer {os.environ['BETTERJOBS_API_KEY']}",
    "BetterJobs-Version": "2026-10-01",
}
BODY = {
    "filters": {
        "title_or": ["Data Engineer", "Analytics Engineer"],
        "country_code_or": ["DE", "NL"],
        "posted_within_days": 30,
    },
    "waterfall": {"strategy": "max_coverage"},
    "limit": 100,
}


estimate = requests.post(f"{API}/jobs/search", headers=HEADERS, json={**BODY, "dry_run": True}, timeout=30)
estimate.raise_for_status()
print(estimate.json()["estimate"]["credits_range"])  # free


body = {**BODY, "waterfall": {**BODY["waterfall"], "max_credits": 100}}
resp = requests.post(f"{API}/jobs/search", headers=HEADERS, json=body, timeout=30)
resp.raise_for_status()
jobs = resp.json()["data"]  # already merged and deduplicated
```

**TypeScript**

```ts
const API = 'https://api.betterjobs.cc/v1';
const headers = {
  Authorization: `Bearer ${process.env.BETTERJOBS_API_KEY}`,
  'BetterJobs-Version': '2026-10-01',
  'Content-Type': 'application/json',
};
const body = {
  filters: {
    title_or: ['Data Engineer', 'Analytics Engineer'],
    country_code_or: ['DE', 'NL'],
    posted_within_days: 30,
  },
  waterfall: { strategy: 'max_coverage' },
  limit: 100,
};


async function search(payload: object) {
  const res = await fetch(`${API}/jobs/search`, { method: 'POST', headers, body: JSON.stringify(payload) });
  if (!res.ok) throw new Error(`${res.status} ${await res.text()}`);
  return res.json();
}


const { estimate } = await search({ ...body, dry_run: true }); // free
console.log(estimate.credits_range);


const { data: jobs } = await search({ ...body, waterfall: { ...body.waterfall, max_credits: 100 } }); // merged, deduplicated
```

For pulls above 100 jobs, use `POST /v1/searches` (up to 10,000 jobs) and read the result pages for free. See [Sync patterns](https://docs.betterjobs.cc/guides/sync-patterns.md#initial-backfill).

### From Elasticsearch DSL to filters

BetterJobs has no query DSL. Translate the clauses you use into the flat filters below; anything else, filter client-side.

| What your Coresignal query does             | BetterJobs filter                                             |
| ------------------------------------------- | ------------------------------------------------------------- |
| Match any of several titles                 | `title_or`                                                    |
| Exclude titles                              | `title_not`                                                   |
| Restrict to countries                       | `country_code_or` (ISO 3166-1 alpha-2)                        |
| Restrict to companies                       | `company_domain_or`                                           |
| Restrict by recency                         | `posted_within_days` (1 to 365, counted from `first_seen_at`) |
| Restrict by seniority or employment type    | `seniority_or`, `employment_type_or`                          |
| Include expired or closed postings          | `include_closed: true`                                        |
| Nested boolean logic, scoring, aggregations | No equivalent. Run several searches or post-filter            |

Different filters combine with AND; values inside one `_or` list combine with OR.

## Map the fields

We do not publish one-to-one field names for Coresignal yet: none are confirmed against Coresignal’s current schema. Map by meaning using the [field dictionary](https://docs.betterjobs.cc/data/field-dictionary.md), which lists every BetterJobs field with its type, derivation and what `null` means. The fields you will use most:

| You need            | BetterJobs field                                                                   |
| ------------------- | ---------------------------------------------------------------------------------- |
| Stable job key      | `id` (canonical, `job_...`)                                                        |
| The source’s own id | `sources[].provider_job_id` where `sources[].provider` is `coresignal`             |
| Duplicate handling  | Not needed: duplicates are merged. `sources[]` lists every source that saw the job |
| Still live?         | `status`, `closed_reason`, `last_verified_at`                                      |
| When it appeared    | `posted_at` (employer date, may be `null`), `first_seen_at` (earliest sighting)    |

## Billing differences

1. **No collect step.** Coresignal charges 1 credit per job collected, per its docs. BetterJobs charges 1 credit per unique job returned by the search itself.

2. **Duplicates are free.** With one row per source, the same opening can arrive more than once. BetterJobs merges them into one canonical job and charges once. `metadata.duplicates_merged` shows how many records were folded.

3. **Re-reads are free.** A job you already paid for returns free in any later search or `GET /v1/jobs/{id}`. The ledger shows it as `already_paid`.

4. **Company data costs less and covers less.** Coresignal company records cost 20 credits each, per its docs. A BetterJobs company profile costs 1 credit and covers hiring only.

Coresignal is a partner provider: Growth includes two partner providers, Pro and above include all six. `GET /v1/providers` shows whether your plan enables `coresignal`. On Scale you can bring your own provider keys. Plans are on [Credits and billing](https://docs.betterjobs.cc/concepts/credits-and-billing.md).

## What changes in your code

1. **Delete the collect loop.** Search returns full jobs. No per-id fetch.

2. **Delete the dedup filter.** No `isDuplicate` check. Key rows on the canonical `id`.

3. **Rewrite DSL queries as `filters`.** Use the table above. Unknown filter names return `400 unknown_filter`.

4. **Replace bulk deliveries with async searches.** One async search per slice (for example per country), up to 10,000 jobs each, upserted on `id`.

5. **Read provenance from `sources[]`.** Each entry has `provider`, `provider_job_id`, `url`, `first_seen_at`, `last_seen_at` and the `fields` it contributed.

## Pitfalls

> null is not false
>
> `location.remote: null` and `is_hiring.value: null` mean unknown. Do not map them to `false` when you port code that expected booleans.

- **`posted_within_days` caps at 365.** For older history, keep Coresignal or your existing archive.
- **Partial results are billed only for what returns.** A provider timeout gives `200` with `metadata.status: partial`. Re-run later; already-paid jobs are free.
- **Set `max_credits`.** Coresignal’s free search let you look before paying. Do the same with `dry_run`, then cap the real request.

## Next

- [Coresignal provider page](https://docs.betterjobs.cc/providers/coresignal.md)
- [Canonical jobs](https://docs.betterjobs.cc/concepts/canonical-jobs.md): how merge and dedup work.
- [Async searches](https://docs.betterjobs.cc/platform/async-searches.md) and [Sync patterns](https://docs.betterjobs.cc/guides/sync-patterns.md)
